Model card
The v6 scanning selection model developed for the study. Everything on
this page is read out of the trained checkpoints that ship with this site, or copied from
the study's own recorded results. The model and its training code will be released with the
paper.
What it computes
A transcript carries an ordered list of candidate reading frames, 5′→3′. Each gets two
numbers from sequence: capture pk, the probability a
scanning ribosome initiates here, read from the start codon window alone; and
decay dk, the probability the transcript is degraded
given that translation begins here, read from both windows and the structural block.
Selection is stick-breaking, so a ribosome reaches candidate k only by passing
every earlier one:
P(select k) = pk ∏j<k (1 − pj)
· P(NMD) = ∑k P(select k) dk
The mass left over, where the ribosome passes every candidate, contributes nothing. That
is why the bars in the prediction panel do not sum to one.
Performance
Held-out test AUROC — ± — across
seeds, on — transcripts of which — are
NMD susceptible, on chromosomes 1, 3, 5 and 7 with paralogs excluded. Taken from the
study's input ablation analysis, using the model with every input retained.
Where the signal comes from
Each cell keeps only the named input groups and removes the rest: SEQ is the
four base channels, JUNC the exon-junction channel, and STRUCT the structural block, which
for this variant is the downstream exon-junction count alone. The floor is measured rather
than assumed to be 0.5, because P(NMD) is a product over a candidate queue whose length
correlates with the label, so a model given nothing informative still discriminates.
Verification
The model on this page is a JavaScript reimplementation of the study's own. It was
checked against PyTorch 2.13 running the real checkpoint on identical inputs, stage by
stage: the convolutions, the batch normalization, the binned max-pool, the fully connected
embedding, the per-candidate capture and decay, the stick-breaking, and both output paths.
Worst relative disagreement across — comparisons:
—, which is float32 rounding rather than a difference.
Known limits, from the model's own authors
Three things worth knowing before you trust a number here.
- The AUG window's fill boundary leaks ORF length. The window is filled
only as far as the midpoint between start and stop, so the downstream fill extent is
min(100, ORF length ÷ 2). For any ORF shorter than 200 nt, which is
— of candidates, where the fill stops encodes the
ORF's length exactly with no sequence involved. The study measures this fill saturation
as a ~—× odds marker for "this is the real ORF". It is
reproduced faithfully here because the model was trained with it and it is computable at
prediction time, but it does mean capture should not be read as pure initiation
context.
- The decay head is handed the answer. Its only non-sequence input is
the downstream exon-junction count, which is the premature stop
indicator.
- Every candidate has a start and a stop by construction, so the model
never saw a negative example for either landmark and had no gradient from which to
learn to check them.
Assumptions this page makes
About
A browser-resident implementation of the scanning selection model, so you
can score a transcript that has never been sequenced.
The three tabs
Isoform Scanner takes one transcript at a time. Paste a sequence with its
exon-exon junctions, or look one up from the study or Ensembl, and the model returns P(NMD)
along with the candidate ORFs behind it.
Bulk Scan takes a whole annotation at once. Give it a GTF and its matching
FASTA, and every transcript is scored in your browser and ranked by P(NMD).
Splice Junction Analysis asks what a splice event does to a gene's NMD
fate. Pick a gene, a cell type, and a junction or exon, and the tab splits that gene's
isoforms by whether they have it, then compares how much RNA each group holds with NMD
running against NMD blocked. These numbers are measured rather than predicted, and the model
has no part in this tab.
What NMD is, and what the model predicts
Nonsense-mediated decay is the cell's system for destroying faulty messenger RNAs. The
model answers one question per transcript: would NMD destroy this? Its
output is a probability. It reads only sequence and ORF structure, with no expression data
and no cell type, so it can be applied to transcripts nobody has measured.
Where the training labels came from
Lung cells were treated with an SMG1 inhibitor, which switches NMD off, so transcripts
that accumulate under SMG1i treatment are likely to have been destroyed by NMD. An isoform
counts as NMD susceptible when its mashr local false sign rate is below
0.05 with a positive posterior mean log₂ fold-change in at least one of four cell
types.
Reading the output honestly
- Junctions matter more than sequence. The study's input ablation
analysis puts the junction channel alone at — AUROC and
the four base channels alone at —. If you
paste a sequence without junction positions, you have removed the model's strongest single
input and it will lean toward "escapes NMD". The page says so when that happens.
- Check the split badge on study isoforms. A transcript in the training
split was used to fit the model, so agreement there is not evidence.
- The reference-start rescue cannot fire here. That rescue is the
study's rule that a transcript's annotated start codon is always admitted, even when it
scores below the Kozak floor. The annotation is not public, so a pasted transcript can get
a slightly smaller candidate pool than the study built for the same sequence. It affected
— of — candidates.
- Five members, not one. —
Scope
- — isoforms indexed from the study, across
— genes, of which — are
observed to be NMD susceptible.
- Exon structures come from the study's long-read transcript annotation. The sequences
themselves run to 1.6 GB and do not ship with this site, so an isoform lookup fills in the
junctions and then either reads the sequence from a sequence store this deployment has
been given, or leaves you to supply it.
checking…
- Any transcript in Ensembl release 115 can be scored, by exact gene
symbol or by
ENST/ENSG identifier. Release 115 is the annotation behind
GENCODE v49 on GRCh38.p14, the gene set this site's Kozak floor was calibrated against,
which is why it is pinned rather than following whatever is current. Sequence and exon
structure arrive together, and the page checks the exon lengths sum to the length of the
cDNA before pairing them. Ensembl's API has no fuzzy search, so a partial symbol finds
nothing there while still matching in the study index.
What leaves your machine
The model, its weights and every calculation on this page are local, so a sequence you
paste is never uploaded. That is the point of shipping the model rather than an API. Two things
are outbound by construction, and both only when you use them.
- The Ensembl lookup sends the same term to this site's own server,
which forwards it to
e115.rest.ensembl.org and returns the answer. It goes
the long way round because Ensembl's archived releases send no CORS header, so a browser
cannot call them directly, so pinning release 115 costs one hop. Your search term reaches
Ensembl either way; what differs is that this site's server sees it too.
- The study isoform lookup, on a deployment with a sequence store
configured, sends the isoform ID you selected to this site's own server and gets that
transcript's sequence back. It sends nothing you typed. Where no store is configured it
makes no request at all. checking…
Each is behind its own checkbox above the search box, and an unticked source is never
contacted. Untick them and the page makes no outbound request at all, though a pasted
sequence still scores and everything else on the page still works.
Who made this
This site, meaning the browser port of the model, the isoform index and the Bulk Scan and
Splice Junction Analysis tabs, is the work of Mateo Martinez,
Yul Leshem and Peter Castaldi. The model itself and the
study whose data it is trained and indexed on are cited separately, under
How to cite.
Reusing it
The code behind this site is released under the MIT license, so take it, change it and
build on it. That covers the browser implementation of the model, the four tabs and the
build scripts. It does not cover the trained weights or the study data shipped with
them, which come from the study's own sequencing and training, and whose terms the study
sets. To reuse the weights or the data rather than the code, cite the study and check its
terms first.
Data version
This build: loading….
Methods summary
What happens between pressing the button and seeing a number. Every step
is the study's own, ported rather than reinvented.
Which set of transcripts a number is about
Numbers on this site count four different things, and they are not the same size. The
largest is the indexed universe of —
isoforms, everything the study measured as expressed, which is what the search covers. Of
those, — carry an observed NMD call and made
up the model's own universe, split into training, validation and test sets. The test AUROC
on this page is measured on the — transcripts held out on
chromosomes 1, 3, 5 and 7. A candidate is one reading frame rather than one
transcript, so those counts run much larger, at — across
the model's universe. The Kozak floor was calibrated on a fifth set that is none of these,
the MANE Select transcripts of GENCODE v49.
1 · Enumerate
Every ATG with an in-frame stop codon downstream becomes a candidate. No minimum length, so
an ATG immediately followed by a stop counts. Two ATGs in the same frame sharing one stop
are separate candidates, as are ORFs in different frames. In the study this averaged
—
candidates per transcript before any filtering.
2 · Score initiation context
Each start codon is scored with the Cavener–Ray position weight matrix at labeled
positions −6…−1, +4, +5 relative to the A of the ATG. Positions running off the transcript
are skipped and the remainder summed.
3 · Admit
A candidate is admitted when it scores at or above the MANE calibrated Kozak floor
(—, the 5th percentile of scores at MANE Select start codons,
— of — scored under
GENCODE v49) and starts in the first half of
the transcript. If nothing is admitted, the five highest-scoring candidates are taken
instead. Average admitted pool in the study: —
candidates, median —, maximum
—.
4 · Encode
Two 1000-nt windows per candidate: one anchored on the AUG, 900 bases upstream and 100
downstream; one anchored on the stop codon, 500 either side. There are nine channels per
position: one-hot A/C/G/T, an exon-junction indicator, a 50-base centered rolling GC
fraction, and three reading-frame channels relative to that candidate's own start codon.
Each window is
filled only up to the midpoint between start and stop, and zero-padded beyond without the
anchor moving.
A position whose base is not A, C, G or T leaves the four base channels empty, and only
those. The junction indicator, the three reading-frame channels and the GC average are
written there as they are anywhere else, and the position counts in the GC denominator as a
non-GC base. This is what the training encoder does, so it is what the site does. Leaving
the position blank would read better, but it would move every prediction on a sequence
containing an N away from what the weights were fitted to.
5 · Score and aggregate
Three convolutional encoders, one for capture and one each for the AUG and stop windows on
the decay path, reduce each window to 32 numbers via two convolutions and a maximum within
each of eight bins along the length. Binning rather than taking a global maximum is what
lets position survive, since a global maximum records only that a pattern occurred
somewhere and not how far along it was, and the distance from a stop codon to the next
junction is the whole question here. The structural block enters through a linear layer, the
three embeddings fuse to 64 dimensions, and stick-breaking over the candidate queue produces
the transcript's NMD logit.
Expression · how the productive fraction is calculated
Everything above is the model. This is the expression table the Splice
tab reads, which the model has no part in: it is measured, not predicted.
Two compartments per gene, per cell type
Within each cell type a gene's isoforms are split by NMD status: unproductive ones are
the ones NMD degrades, productive ones are not. The call is the study's own and is made per
cell type, so an isoform is unproductive where mashr gives a local false sign rate below
0.05 with a positive posterior mean log2 fold change, meaning it rose measurably
when NMD was blocked. A gene's unproductive compartment in a sample is the summed expression of
those isoforms, its productive compartment the rest.
Why the SMG1i arm has to be adjusted
CPM divides by the library total. Under SMG1i the unproductive isoforms accumulate
massively, which inflates that denominator and makes every productive isoform read
smaller than it is. That is an artifact of composition rather than a fact about the gene,
and the adjustment below removes it.
The adjustment, in three steps
For a gene g with productive isoforms prod(g) and unproductive isoforms
unprod(g). First, the productive compartment is summed under DMSO:
PgDMSO =
Σi ∈ prod(g) CPMiDMSO
Second, each unproductive isoform is given the proportion it occupied within that gene
under DMSO, and that share is scaled to the gene's productive compartment in the SMG1i
sample s:
sharei = CPMiDMSO ⁄
PgDMSO, i ∈ unprod(g)
CPMi,s = sharei ×
Σj ∈ prod(g) CPMj,sSMG1i
Third, each adjusted library is rescaled back to one million, with the sum in the
denominator running over the reassigned values from the step above rather than the
originals:
CPMi,sadj = CPMi,s × 106 ⁄
Σk CPMk,s
An unproductive isoform whose gene has no positive DMSO productive compartment has no
share to take, and is held at its DMSO mean instead. Summing CPMadj over
prod(g) gives the gene's adjusted productive CPM, in which the unproductive
surge's contribution is gone.
The productive fraction
prodCPM and unprodCPM are the sums of CPM over a gene's productive and unproductive
isoforms and totalCPM is the two together, so the productive fraction is
that gene's productive share of its own output:
f = prodCPM ⁄ totalCPM
What was checked
Control genes, the ones with no unproductive isoform to surge, come out essentially
unchanged between DMSO and adjusted SMG1i, at Spearman ≥
— with a slope near 1. Gene rankings also survive, at
Spearman ≥ — between adjusted and raw fold changes.
Together those say the adjustment removes a compositional artifact rather than reordering
the biology.
How this relates to the share the Splice Junction Analysis tab prints
Two differences, both deliberate, and worth knowing before the tab's number is quoted
against f above.
It is per group, not per gene. That tab splits one gene's isoforms
by whether they carry a junction and computes a productive share within each side, so it
answers "how much of this half of the gene is productive" rather than the gene's
own f.
Its SMG1i arm is hybrid. The shipped per-condition table carries anchored
values on productive rows and observed values on unproductive ones, so it sums to
1.03–1.09 million rather than to one million. The tab divides that column as it stands: an
anchored numerator over a total whose unproductive part is still raw. The DMSO arm has no
such issue, being a proper CPM library summing to one million. Undoing the anchoring on the
productive arm instead, one factor per cell type
(—), moves the SMG1i
share about 0.5–1.3 points lower. The tab says so under its own tables.
What is not done here
The study always admits a transcript's annotated start codon, which it calls the
reference-start rescue. That annotation is not public, so the rescue never fires on this
page. The Kozak
floor is also recorded to four decimal places rather than at full precision; exactly one
achievable score sits inside that rounding interval, so only starts scoring precisely
−1.2507921 could be admitted differently.