Welcome to scTHREAD

Frozen release · 453 runs · 2026-08-03 Live catalog · rolling Manuscript figures and tables cite the frozen release. Release details

Statistics

453

Sequencing runs

923,389

Cells

>200,000

Observed isoforms


Human and mouse long-read transcriptomes

453 runs

Evidence layers

Gene expression
Available
Isoform
Available
Poly(A)
Available
ASE
Available
Junction
Available

Each result reports the runs and biological units that contribute to that evidence layer.

Global cell map

scTHREAD portal cell atlas — 141,313 cells across 21 datasets, colored by cell type

Cell-resolved portal atlas: 141,313 cells across 21 datasets (123 runs, 90 cell types). UMAP is exploratory and does not imply trajectory.

Compare balanced, all-cell and density atlas views →

Transcript & junction browser interactive

Loading the selected data scope…
How these counts differ
Choose a gene to load expression, DIU, APA, ASE and junction evidence.
Scroll over the track to zoom · drag horizontally to pan · select an arc or table row for details. Batch junction lookup
JunctionSpanMoleculesReadsRunsStudies
Signal view

Cell-type annotation

Expression

Loading signal…

Live data scope: results report their contributing runs and biological units by evidence layer. Region queries are aggregated on demand and bounded to a 5 Mb window.

Search

Find a dataset or gene.

Choose a species first. Then open a dataset directly or search by biological context, gene, accession or coordinate.

Species
Browse a dataset

Choose by biological source

Labels describe tissue, disease or cell line; accessions remain available in the results.

Loading the catalog scope…

Search the database

Enter what you know

Choose a dataset above, or try one of the example searches.
Novel models

Search run-local novel transcript models.

Inspect NIC/NNC model IDs observed in the release's discovered-transcript layer. Repeated IDs across runs remain separate observations; no cross-run equivalence is assumed.

Export first 500 matches

Enter a model, run or dataset to inspect the run-local index.

No novel-model query requested.
Analyze

Compare cell types.

Compare the same feature between two cell types across selected studies. Study comparability is checked when you select cell types, before running the comparison. Per-study direction agreement is descriptive; limited or unavailable evidence remains visible.

1
2
3
Studies

Study totals and current-feature coverage are shown separately. Runs are sequencing files, not independent biological sources. Paired-source counts depend on both selected cell types and appear in the results.

4
Cell types
Select cell types to preview study comparability.

Finding eligible comparisons for MACF1…

Comparison

Choose two cell types

The result will show which transcript feature differs and how many paired samples support the comparison.

Across-study overview · isoforms × studies

One row per selected isoform and one column per evaluable selected study. The difference view keeps the A/B pairing gate of the comparison above; the descriptive view browses one cell type without any pairing gate. Cells show their actual supporting sources; low support is marked limited, and unavailable is never plotted as zero. Studies that cannot evaluate this query fold into an explicit list below the matrix. The isoform selection is shared with the UMAP workspace.

Observed · descriptive
Overview mode
Isoforms (rows, shared with UMAP workspace)

Run a comparison or start descriptive browsing to list available isoforms.

Share and export

The share link restores gene, species, studies, cell types, isoforms, measures, scale and pages.

Choose studies and cell types above, then run the comparison (difference mode) or switch to descriptive mode to browse one cell type without the pairing gate.

Isoforms on UMAP · multi-study workspace

One row per selected study: its cell-type reference map, then one panel per selected isoform. Panels within a row share cells, coordinates, zoom and hover; rows are independent studies and are never geometrically aligned or linked. All expression panels in the workspace share one color scale.

Observed · exploratory
Isoforms (columns — paged 3 at a time, shared with the overview matrix)

Loading isoforms…

Displayed quantity

Usage divides by the complete per-cell gene transcript denominator (reference transcripts of the same gene, same cell); it is never approximated from the selected isoforms alone. Cells whose denominator is zero or whose source version cannot be verified stay gray — unavailable, not 0%.

Shared color scale
Display

The workspace follows the gene, species and study selection above.

One shared numeric range for every panel in the current comparison. Gray marks unavailable measurements (never plotted as zero); near-white marks observed zeros.

Junctions

Check a splice junction.

See whether it was observed, where it recurs and whether it uses canonical splice sites.

Junction lookup

Paste one or more GRCh38 junctions, one per line: chr1:+:39282374-39283189

Database records
1 junction

Ready

Results will show observation, recurrence, splice motif and supporting molecules.

Optional model estimate

For one junction, estimate pooled usage when measured coverage is incomplete. A higher score means the model ranks it closer to frequently used junctions; it is not a probability and does not compare cell types.

Checking model availability…

No estimate requested

Out-of-domain junctions return no score.

Uniform processing workflow

01

Raw long reads

Human and mouse sequencing runs enter without relying on published aggregates.

02

Platform-aware processing

Protocol-specific barcode and UMI recovery followed by IsoQuant transcript assignment.

03

Cell × molecule

Run-aware barcode joins and per-cell UMI deduplication.

04

Shared evidence

Isoform, PAS, junction and allelic views from one read-facts layer.

Show processing pipeline and decision thresholds

Long-read alignments and IsoQuant assignments are processed uniformly within each run. Gene, transcript, poly(A)-site, junction and allelic evidence use run-aware barcode namespaces; UMI deduplication is applied where molecule tags are available. Every decision below is recorded per run in the frozen processing manifest.

Raw long reads 453 runs · 34 datasets Identity resolution cell barcode in read? barcode-resolved 155 runs · CB tag in read sample/file-resolved 298 runs · one library per file release partition · 155 + 298 = 453 IsoQuant 3.13.1 transcript assignment 376 verified · 77 unverified Run-aware namespaces per-cell UMI dedup where molecule tags are available Evidence layers core layers 453/453 runs DIU 124 · APA 100 junction 100 · ASE 99

Decision rules. A run is barcode-resolved when every read carries a cell barcode (read group tag:CB); otherwise it is one library per file and cell identity comes from the file name. Assignment uses GENCODE v44 on GRCh38 (human) and GENCODE vM33 on GRCm39 (mouse). The 77 runs without a retained IsoQuant log keep an explicit unverified version label; the pipeline-level version is not extrapolated to them. Core layers (gene counts, transcript counts, read assignments, transcript models) are present for all 453 runs; DIU, APA, junction and ASE layers are present where the underlying measurements exist.

Software and reproducibility

IsoQuant 3.13.1 (per-run verification above). Cell-map builds used scanpy 1.10.3, harmonypy 0.2.0, umap-learn 0.5.12 and scikit-learn 1.6.1; integration diagnostics run kNN (k = 15) on the displayed coordinates and flag a batch-concentration ratio above 3. Per-run records: processing manifest API · frozen manifest TSV · embedding provenance TSV.

Show statistical methods and test statistics

On-demand comparisons merge technical runs into registered biological sample or capture units. One independent unit is a study_id + donor_or_source_id pair; technical runs and repeated states are summed before testing. The interface offers only cell-type pairs supported by at least three shared independent units, and a gene is tested only when it has at least 3 donors per cell type and a minimum depth of 20.

Permutation test. For each gene, cell-type labels are permuted 9,999 times at the biological-unit level (seed 20260727). Matched null calibrations (3 seeds × 200 permutations each) yield zero FDR discoveries in every layer. The reported effect is the equal-donor estimand (effect_equal_donor): each donor contributes equally, not each pooled read count.

Multiple-testing correction. Benjamini–Hochberg q-values at a 5% FDR; a gene is significant only with q < 0.05 and an absolute equal-donor effect of at least 0.2. The q-values and significance flags were recomputed independently, without statsmodels; the largest discrepancy across the observed layers is 2.22 × 10−16.

LayerGenes testedSignificant (FDR 5% + effect ≥ 0.2)Raw p < 0.05 fractionNull raw p < 0.05 range (3 seeds)
Isoform usage (DIU)8,0922,0080.4580.026–0.060
Poly(A) site usage (APA)10,5312,5580.5510.027–0.049
Allelic balance (ASE)6,93000.0780.040–0.047

Allelic balance remains descriptive in live views because haplotype labels are not phased consistently across units; the frozen ASE run above is an archived analysis record in which no gene passes the significance rule.

Prediction model

The optional junction-usage model is sequence- and RBP-informed (220 RBP profiles). Held-out validation: pooled Spearman 0.440, cell-type mean 0.112, cis-only 0.082, shuffled-RBP control 0.139. The estimand is a relative pooled donor-anchored usage propensity — not a calibrated PSI, a probability of biological validity, or a cell-type-specific effect; junctions outside the model domain return no score.

Download

Sample manifests

The frozen release 2026-08-03 is the current portal denominator: 453 run records from 34 human and mouse datasets and 923,389 cells. The live catalog remains the rolling view for Search and Browse.

StudyBiological contextRun accessionSpeciesPlatformLibrary kitCellsAssayStatus
Loading run-level metadata…

Frozen ASE analysis table

Archived analysis record. Current Browse views report run-local descriptive allelic-balance observations.

Download TSV

DIU / APA tables

Verified gene-level RNA-processing results.

DIU CSVAPA CSV

Processing manifest

Frozen 453-run layer contract with public availability, byte counts and status fields; internal paths are redacted.

Download TSV
Show column dictionary for downloaded tables
output_tier · isoquant-core-present
All four IsoQuant core output layers (gene counts, transcript counts, read assignments and transcript models) are present for the run. The tier is a file-presence contract, not a downstream quality score; DIU, APA, junction and ASE layers remain layer-specific. The frozen processing manifest carries the same contract as isoquant_core_status and file_contract_status.
read_type · per-file read type still requires verification
In run_assay_resolution.tsv, 42 PacBio runs are recorded as consensus reads supported by model and source records, but no per-file inspection has confirmed the read type; the 12 runs confirmed by FASTQ inspection say so explicitly in the same column.
isoquant_version · unverified
No isoquant.log was retained for the run, so the version cannot be re-verified per run and the pipeline-level version is not extrapolated to it (77 runs).

Identity-resolution terms (barcode-resolved, sample/file-resolved) and count scopes are defined in the About glossary.

About scTHREAD

Loading release metadata from the backend…

Glossary and count scopes

molecule
One observed transcript-assignment event supporting a feature. Where reads carry UMI tags the count is UMI-deduplicated (molecules can be fewer than reads); where they do not, the same fields report IsoQuant read-assignment counts. Each evidence layer labels its own count unit instead of claiming verified UMI molecules everywhere.
run-local model
A novel transcript model discovered by IsoQuant inside one sequencing run (IDs such as transcript….nic/.nnic). It indexes that run's evidence only: it is not harmonized across runs, and its counts are assignment units, not verified UMI molecules.
barcode-resolved / sample-file-resolved
The two identity-resolution routes of the frozen release. Barcode-resolved runs (155) carry a cell barcode in every read (read group tag:CB); sample-file-resolved runs (298) are one library per file, so cell identity comes from the file itself. Together they partition all 453 runs; deposited assay labels remain record-level metadata.
frozen release / live catalog
The frozen release 2026-08-03 is the cited manuscript denominator: 453 run records from 34 human and mouse datasets. The live catalog is the rolling registry behind Search and Browse and can index more records than the frozen release, so one page can legitimately show both scopes.

How the counts differ

Five cell counts appear across the portal. Each measures a different scope, so the numbers are expected to differ; none is an error.

Loading count scopes…

First-user tutorial: trace PTPRC evidence

01

Search

Open PTPRC search results and select the human gene record.

02

Inspect context

On the gene page, compare the analysis-status strip with the shared-coordinate cell map.

03

Select evidence

Choose an isoform, inspect its usage denominator, then open the junction tab and apply the molecule filter.

04

Reuse

Export the visible table or reproduce the overview through the documented API.

Scope rule. Counts in a filtered browser table are not database-wide totals. Each overview and export labels its evidence layer and scope; the frozen manuscript manifest is separate from the rolling live registry.