Preprint Seeing below the limit of detection — the methods paper behind Span
Research

Real methods. An open benchmark. Honest results.

We'd rather show the evidence than assert the vision. Two pieces of work anchor Span: a transparent detection method and a public benchmark whose central finding is a negative result we report in full.

In ninety seconds
  • 1The method. A censored-Poisson Bayesian latent-growth change-point test on the sequence of ctDNA detection calls. Zero trainable parameters. On synthetic trajectories in the slow regime, at a matched 10% false-alarm rate, it flags 26.8% ±3.9 of progressions at least twelve weeks ahead against 14.0% ±3.0 for a snapshot rule, mean of five seeds. Synthetic, and the advantage closes to nothing when emergence is fast. The paper, and where it stops working
  • 2The benchmark. OncoTraj v1: 813 patients in EGFR-mutant NSCLC on first-line osimertinib, harmonised from two registries and a published trial supplement, with frozen leakage-audited splits. How it is built
  • 3The finding, and it is negative. Re-scored inside a single source, the timing task collapses to chance: random forest falls from 0.656 to a C-index of 0.432 [0.360, 0.514] against a 0.500 floor. The binding constraint is the modality, not the model. All three tasks, with baselines
  • 4What is not here. The Span detector has not been run on OncoTraj and cannot be — it reads a series, and OncoTraj v1 is snapshots. There is no Span-versus-baseline result on real patients anywhere on this site. Why the two are in different tumour types
  • 5What we want from you. One retrospective serial-ctDNA arm, scored against a pre-specified naive rule. The experiment, in full

All intervals are bootstrap 95%. Both preprints are founder-authored and neither has been peer reviewed.

Methods paper

Seeing below the limit of detection.

The Span detector, a censored-Poisson Bayesian latent-growth (CP-BLG) change-point test, reads the rate of ctDNA detections over time to flag resistance before the variant is callable on any single blood test. It has no trainable parameters; the advantage is structural.

Two bar panels from the synthetic evaluation, early sensitivity (progressions caught at least 12 weeks ahead) and overall sensitivity, both at a matched 10% false-alarm rate. Each panel groups three disease regimes, indolent, slow and fast, with a commodity-snapshot bar beside a Span bar and standard-deviation whiskers over 5 seeds. Span leads in every group; on early sensitivity the ratio runs 2.3 times in the indolent regime, 1.9 in the slow, and 1.2 in the fast, where the two bars nearly meet.
Sensitivity across disease regimes at a matched 10% false-alarm rate (synthetic evaluation, mean of 5 seeds). The advantage scales with sub-LoD dwell time: largest in the clinically hard, indolent regime, and close to gone when emergence is fast.
Open benchmark

OncoTraj: an honest floor for what snapshots can predict.

OncoTraj harmonizes two real-world registries and a published trial supplement into one schema with frozen, leakage-audited splits, and ships reproducible baselines. Its purpose is to establish a public floor, and an honest ceiling, for longitudinal resistance prediction in EGFR-mutant NSCLC on first-line osimertinib.

What OncoTraj v1 harmonises: two real-world registries and one published trial supplement. The held-out patient-level test split is 122 patients, drawn across all three.
SourcePatientsAccess
MSK-CHORD672cBioPortal (open)
FLAURA molecular-resistance supplement107Public supplements
AACR Project GENIE BPC (NSCLC)34Registered access
v1 total813
Headline finding

Re-scored inside a single source, the timing task collapses to chance: random forest falls to a C-index of 0.432 [0.360, 0.514] against a 0.500 floor.

The binding constraint is the data modality, single-timepoint tissue/ctDNA snapshots, not the model. On the mixed-source split that same model reads 0.656; the apparent edge is dataset membership standing in for biology, and it does not survive removing it. Serial ctDNA trajectories, not a better algorithm, are the precondition for above-chance performance. That negative result is the contribution: it tells the field where the signal must come from.

The scope of that claim, stated plainly. It is Task B, the timing task, which is the one task with a committed within-source evaluation. For Task A the corresponding figure — a 12-month landmark logistic AUC of 0.596 [0.478, 0.714], an interval that crosses 0.50 — comes from the benchmark's leaderboard prose and is not reproducible from the committed eval reports, so we quote it as prose and do not chart it. Task C has not been re-scored within-source at all.

The three charted results below are read straight from the benchmark's evaluation reports (oncotraj-eval schema 1.1.0) with bootstrap 95% confidence intervals. The Task C table and the Task A within-source AUC above are not: both are transcribed, and are labelled as such where they appear. We do not claim above-chance within-source performance where the data does not support it.

Two diseases, one gap

Why the method is demonstrated in breast and the benchmark is built in lung.

The two pieces of work are in different tumour types, and the reason is the same fact in both directions: no public serial-ctDNA cohort exists for the disease where the clinical case is strongest.

The clinical case is strongest in HR+/HER2− metastatic breast cancer. It is the largest biomarker-defined solid-tumour population, resistance to first-line CDK4/6-inhibitor plus endocrine therapy arrives through competing, identifiable mechanisms — ESR1 in roughly a third of progressors, with PIK3CA, RB1 loss and HER2-activating mutations among the rest — and PADA-1 has already shown that acting on a rising ESR1 ctDNA signal ahead of imaging is worth something. So that is where the method is demonstrated. There is no public serial cohort in it, which is why that demonstration is on synthetic trajectories from an openly released simulator, and why every CP‑BLG number on this site is labelled synthetic, five seeds.

The public data that does exist is single-timepoint, and the deepest of it is EGFR-mutant NSCLC on first-line osimertinib. So that is where the benchmark is built — and what OncoTraj measures there is the ceiling of the modality itself. Its negative result is the argument that the serial cohort has to be collected.

The Span detector has not been run on OncoTraj, and cannot be. The detector reads a series of detection calls over time; OncoTraj v1 contains snapshots. Nothing on this page is a Span-versus-baseline result on real patients, and the missing middle — a retrospective serial-ctDNA arm where the change-point test is scored against a naive rising-value rule — is the experiment we are looking for a partner to run. What that experiment needs, and what comes back

The experiment

The missing middle, written out.

This is the study that turns everything above from a method into a result, or ends it. It is retrospective, it needs no new samples and no protocol change, and it can be specified completely before anyone runs it — which is the only reason it is worth running.

  • 1The data. An existing serial-ctDNA arm with per-variant detection calls and draw dates — a series, not a pair of timepoints — and a progression date to score against. De-identified. Any tumour type where resistance is driven by identifiable variants; HR+/HER2− breast and EGFR-mutant NSCLC are where our two pieces of work already sit.
  • 2The comparison. The change-point test against a naive rising-value rule on the reported variant fraction, at a matched false-alarm rate. Both rules and the false-alarm rate fixed in writing before outcome labels are unblinded.
  • 3The endpoint. Sensitivity and lead time for each rule, per patient, with bootstrap intervals — the same reporting discipline as the baselines on this page. The criterion, from the methods paper: change-point detection should catch more impending progressions, and catch them earlier, than thresholding the reported variant fraction. If it does not, that is the answer.
  • 4What comes back to you. The full analysis and the code that produced it. Our interest is a result that can be published either way; we are not asking anyone to take a private number on trust, which is the same reason the detector and the benchmark are already open.
  • 5What is not settled. Price, licence terms, data-processing agreements and the regulatory route. Span’s software is research use only, is not a medical device and has not been submitted to any regulator. These follow the first result on real serial data.
Early sensitivity
26.8%
Slow regime: progressions flagged ≥12 wk ahead at a matched 10% false-alarm rate, 26.8% ±3.9 against 14.0% ±3.0 for a snapshot rule, mean of 5 seeds. The sparkline is one 90th-percentile case, not the mean. All demo data synthetic.
SOURCES HARMONISED HELD OUT SCORED ON MSK-CHORD 672 patients · registry · cBioPortal FLAURA supplement 107 patients · randomised trial GENIE BPC 34 patients · registry · registered One schema 813 patients total Frozen splits Leakage-audited Reproducible baselines Nothing is re-split per run. 122 patient test split Not every patient carries every label, so each task is scored on its own subset. A · resistance n = 110, landmark B · timing n = 85 mixed, 101 within C · mechanism n = 85 At these sizes the intervals are wide. The benchmark exists to state the floor honestly, not to rank models.
Fig. 1How OncoTraj v1 is built, and why n differs by task in the tables below — including why the within-source timing count is larger than the mixed-source one, not smaller.

Baseline results

Snapshot baselines on the locked test split, scored two ways: once on the mixed-source split, and once within a single source. The gap between those two numbers is the result.

The three charts below are drawn directly from the benchmark's own evaluation reports (oncotraj-eval schema 1.1.0), not transcribed; the Task C table further down is transcribed by hand, and says so. Task A is scored on the 110 landmark-evaluable patients, Tasks B and C on the 85 with an observed event. The within-source analysis re-fits on MSK-CHORD alone and is scored on n = 101 — more patients than the mixed split, not fewer, because the mixed timing evaluation requires an observed event from every source, while the MSK-CHORD refit simply uses that source's own eligible set. Intervals are bootstrap 95%.

The confound, and what survives it

Trained and scored across all three sources, the timing task looks like it is working: random forest reaches a C‑index of 0.656, well clear of the 0.500 floor. Re-fit and scored inside MSK‑CHORD alone, where dataset membership can no longer stand in for biology, the same model falls to 0.432 — below chance, with an interval that contains it.

Nothing about the model changed. What changed is that it lost access to a shortcut: on the mixed split the date-less FLAURA subset carries a constant pseudo-PFS, so knowing which cohort a patient came from is worth more than knowing anything about the patient. Every learned model loses its apparent edge the same way. The majority floor, having never had one, is unmoved.

Task A — will this patient develop resistance? (n = 110, mixed-source)

On the mixed split the three learned models sit above chance and their intervals exclude 0.500. We do not report that as a finding, because the same confound applies: part of what looks like discrimination is a model recognising which dataset a patient came from. The nearest within-source figure is the 0.596 [0.478, 0.714] quoted above, whose interval crosses chance — but that number comes from the benchmark's leaderboard prose and is not reproducible from the committed eval reports, so it is not charted here and should not be treated as a scored result.

Read this chart as an upper bound on what snapshots can do here — not as evidence that they work.

Calibration

Discrimination is not the only thing a clinical model owes you; the probability has to mean something. Here the ranking inverts. The model that learns nothing is the best calibrated on the split (ECE 0.017), and every learned model is worse — XGBoost worst at 0.159, roughly nine times the floor.

A model can rank patients correctly and still be badly wrong about how likely progression is. At this cohort size, neither property is established.

Full numbers, all three tasks

Task A — resistance, 12-month landmark (n = 110, mixed-source)

ModelROC-AUC95% CIAccuracyF1BrierECE
Majority baseline0.5000.500–0.5000.5640.0000.2460.017
Logistic regression0.6800.581–0.7810.6360.6360.2200.071
Random forest0.6780.569–0.7730.6360.6670.2140.041
XGBoost0.6300.525–0.7390.5820.6100.2550.159

Task B — when? (mixed-source n = 85; within-source MSK-CHORD n = 101)

ModelC-index (mixed)C-index (within)95% CI (within)MAE d (within)
Majority baseline0.5000.5000.500–0.500303.0
Logistic regression0.5810.4920.417–0.570301.5
Random forest0.6560.4320.360–0.514292.9
XGBoost0.6430.4410.361–0.520312.3
Cox PH0.5410.458–0.616513.5

Within source, the best mean absolute error is 292.9 days against a 303.0-day majority floor — a ten-day gap on a year-scale prediction. Cox ranks marginally best on C-index and is the worst on absolute error by 200 days.

Task C — which mechanism? (n = 85, mixed-source)

Transcribed, not generated. Unlike Tasks A and B, Task C is not in data/oncotraj.json, is not charted, and has no within-source re-scoring: these four rows are typed out of the benchmark write-up.

ModelAccuracyMacro-F1F1 on EGFR C797S
Majority baseline0.9880.4970.000
Logistic regression0.8350.5170.125
Random forest0.9530.4880.000
XGBoost0.9650.4910.000

This is the row that matters. The highest accuracy on this task — 98.8% — belongs to the model that never predicts the actionable class at all; it answers "other" every time. Only logistic regression scores above zero on EGFR C797S, at 0.125. Accuracy here is an artefact of class imbalance, which is precisely why the benchmark reports macro-F1 beside it, and why tissue NGS under-capturing C797S makes this a structural non-result rather than a modelling failure.

Why this de-risks the thesis

The negative result is the reason Span exists.

If snapshots could predict resistance, you wouldn't need us. OncoTraj shows, rigorously, on real patients, that they can't. The signal lives in the trajectory, which is exactly what the Span detector is built to read.

Publications

Preprints.

The full write-ups behind the method and the benchmark, open to read, run, and check.

Method · Preprint

Seeing below the limit of detection: a censored-Poisson Bayesian latent-growth change-point detector for serial ctDNA in HR+/HER2− metastatic breast cancer

Aarchi Singh Thakur, Abhijoy Sarkar  ·  arXiv:2606.11876  ·  June 2026

Benchmark · Preprint

OncoTraj: a public benchmark for longitudinal resistance prediction in EGFR-mutant non-small-cell lung cancer on osimertinib

Abhijoy Sarkar, Aarchi Singh Thakur  ·  arXiv:2606.11144  ·  June 2026