Preprint Seeing below the limit of detection — the methods paper behind Span

Backed by Entrepreneur First

Resistance shows up in blood before it shows up on a scan.

A non-detect is not a zero. Span AI reads the series of ctDNA draws, not the single value — a sequential change-point test that counts the misses below the limit of detection as evidence, and flags the turn while the variant is still sub-clinical.

Built for drug developers running Phase II–III programmes, and for the cancer centres that hold the serial draws those programmes generate.

A centrifuged blood-collection tube: straw-gold plasma separated above deep red packed cells, with a single droplet suspended above the open rim.
Plasma, separated. Cell-free tumour DNA circulates in the gold layer - the fraction a single draw reads as one number, and Span reads as a series.
813
patients in OncoTraj v1, harmonized from two real-world registries and a published trial supplement
≈151d
median lead of ctDNA over clinically detected relapse in resected NSCLC — TRACERx, Abbosh et al., Nature 2023. Personalised multi-mutation tracking, not an off-the-shelf panel
Zero
trainable parameters in the Span detector: a fixed statistical test, not a fitted model
Where Span sits

We don’t compete with the tests. We sit on top of them.

Commercial liquid-biopsy panels report a molecular result from a single draw, and they do that well. Each report is one point in time. Span is a reader over the series those draws already produce.

The input is a series of ctDNA draws with per-variant detection calls and dates. Serial draws at six-to-eight-week intervals are routine inside trials and inside molecular-residual-disease programmes; they are not yet routine practice outside them, and that gap is the commercial premise, not a solved problem.

The problem

Care is decided from snapshots. Disease moves in trajectories.

A single biopsy, scan, or genomic report captures one moment of a moving disease. Resistance to targeted therapy is near-universal, but it follows predictable biological pathways, and in blood it often rises along a curve long before it crosses the imaging threshold.

Span reads that curve as it forms. Below, on simulated data, you can watch it happen.

tumour DNA in blood visible on imaging Span flags week 13 Scan shows it week 35 ~22 weeks of lead the ≈151-day TRACERx median (Abbosh 2023); the weeks shown are illustrative weeks on therapy →
Fig. 1Illustrative. The subclone climbs in blood long before it reaches the imaging threshold; the bar is drawn to the ≈151-day median lead reported by TRACERx (Abbosh et al., Nature 2023), and the individual weeks are arbitrary.
The measured result

Snapshots look like they work. Take the dataset away and they don't.

Trained and scored across all three sources in our benchmark, a random forest predicts time to resistance at a C‑index of 0.656 — well clear of the 0.500 floor. Re-fit and scored inside a single registry, where knowing which dataset a patient came from can no longer stand in for biology, the same model falls to 0.432, with an interval that contains chance.

The binding constraint is the data modality — single-timepoint snapshots — and not the model. That is the argument for reading a series, and it is the one result on this page that was measured rather than illustrated.

0.656
C-index, random forest, trained and scored across the mixed-source split
0.432
the same model re-fit and scored inside MSK-CHORD alone, 95% CI 0.360–0.514
0.500
the chance floor both are measured against

OncoTraj v1, Task B (time to resistance). Bootstrap 95% intervals. All three tasks, with the majority baseline beside each

How it works
seven draws, one line of therapy

01

Serial draws, not one

Blood is taken repeatedly through treatment. Each draw yields a cell-free DNA fraction, and each one on its own is a noisy single number.

limit of detection hollow = non-detect, and still evidence

02

Detections, including the misses

Below the assay's limit of detection, a variant flickers in and out. Span keeps the non-detects. They are evidence about how much tumour DNA is there, not absence of it.

change point reported with a date, and a false-alarm rate

03

A change point, with a date

Span fits the rising rate of those detections and flags the point where the trajectory turns, while the variant is still sub-clinical.

What it produces

Three things come out, and a date is one of them.

A series of ctDNA draws goes in and a patient trajectory comes out. Two of these are outputs the method has been evaluated on. The third is listed with what backs it, because it is the one still to be demonstrated.

  • 1Estimated time-to-resistance. How long the current line is likely to keep working.
  • 2A calibrated threshold. The alarm fires at a level set for a chosen false-alarm rate, so the call is auditable.
  • 3A ranked mechanism hypothesis — synthetic evidence only. The detector aggregates evidence across competing resistance pathways. That has been demonstrated on synthetic cohorts and never on real serial data. What would establish it is a serial-ctDNA cohort with sequenced progression biopsies to score against; we list it as an output of the method, not as a result.

Roadmap, not built: EHR and imaging ingestion, and ranked next-line options with trial matching.

What we are asking for

One retrospective serial arm. That is the whole ask.

Span is a method waiting for the data it was built to read, and the experiment that would settle it is small, retrospective, and already specified. It needs no new samples, no protocol change, and no patient to be treated on a Span alarm.

  • 1What we need. An existing serial-ctDNA arm — a trial arm, or a molecular-residual-disease programme — with per-variant detection calls and draw dates, a series rather than a pair of timepoints, and a progression date to score against. De-identified, retrospective, already collected.
  • 2What we run. The change-point test against a pre-specified naive rule — thresholding the reported variant fraction — at a matched false-alarm rate. Both rules fixed in writing before anyone sees the outcome labels, which is the only version of this experiment worth running.
  • 3What comes back. Sensitivity and lead time for both rules, per patient, with the analysis code. If the naive rule wins, that is the answer and the thesis is wrong: the criterion is stated in the methods paper precisely so it can fail.
  • 4What we cannot answer yet. Price, licence terms, and the regulatory route are not settled. Span’s software is research use only — not a medical device, not submitted to any regulator, and no pathway selected. A retrospective re-analysis does not need one. Terms follow the first result on real serial data.
What Span does not claim

Span is pre-clinical. There is no cleared assay, no prospective trial, and no patient has been treated on the basis of a Span alarm. The software is research use only and is not a medical device.

The published results are a methods demonstration on synthetic and public longitudinal cohorts, plus an open benchmark whose headline finding is a negative one: re-scored inside a single source, the timing task collapses to chance. The mechanism task has not been re-scored within-source at all.

That floor is the point. It is the thing a longitudinal model has to beat, and it is why we published it before we published anything flattering.

Latest
  1. Seeing below the limit of detection

    Preprint. The censored-Poisson latent-growth change-point detector, arXiv:2606.11876.

  2. OncoTraj: a public benchmark

    Preprint. 813 patients, frozen leakage-audited splits, arXiv:2606.11144.

  3. OncoTraj v1 baselines evaluated

    Four snapshot models across three tasks. None identifies the resistance mechanism.

Predict the trajectory, not the snapshot.

Oncology decides care from snapshots of a moving disease. If we can learn the trajectory itself, we can act before the window closes, for every patient, not just the ones scanned at the right moment.

Aarchi Singh Thakur Co-founder & CEO