Seeing below the limit of detection: a censored-Poisson Bayesian latent-growth change-point detector for serial ctDNA in HR+/HER2− metastatic breast cancer
Aarchi Singh Thakur, Abhijoy Sarkar · arXiv:2606.11876 · June 2026
We'd rather show the evidence than assert the vision. Two pieces of work anchor Span: a transparent detection method and a public benchmark whose central finding is a negative result we report in full.
All intervals are bootstrap 95%. Both preprints are founder-authored and neither has been peer reviewed.
The Span detector, a censored-Poisson Bayesian latent-growth (CP-BLG) change-point test, reads the rate of ctDNA detections over time to flag resistance before the variant is callable on any single blood test. It has no trainable parameters; the advantage is structural.
OncoTraj harmonizes two real-world registries and a published trial supplement into one schema with frozen, leakage-audited splits, and ships reproducible baselines. Its purpose is to establish a public floor, and an honest ceiling, for longitudinal resistance prediction in EGFR-mutant NSCLC on first-line osimertinib.
| Source | Patients | Access |
|---|---|---|
| MSK-CHORD | 672 | cBioPortal (open) |
| FLAURA molecular-resistance supplement | 107 | Public supplements |
| AACR Project GENIE BPC (NSCLC) | 34 | Registered access |
| v1 total | 813 |
The binding constraint is the data modality, single-timepoint tissue/ctDNA snapshots, not the model. On the mixed-source split that same model reads 0.656; the apparent edge is dataset membership standing in for biology, and it does not survive removing it. Serial ctDNA trajectories, not a better algorithm, are the precondition for above-chance performance. That negative result is the contribution: it tells the field where the signal must come from.
The scope of that claim, stated plainly. It is Task B, the timing task, which is the one task with a committed within-source evaluation. For Task A the corresponding figure — a 12-month landmark logistic AUC of 0.596 [0.478, 0.714], an interval that crosses 0.50 — comes from the benchmark's leaderboard prose and is not reproducible from the committed eval reports, so we quote it as prose and do not chart it. Task C has not been re-scored within-source at all.
The three charted results below are read straight from the benchmark's evaluation reports (oncotraj-eval schema 1.1.0) with bootstrap 95% confidence intervals. The Task C table and the Task A within-source AUC above are not: both are transcribed, and are labelled as such where they appear. We do not claim above-chance within-source performance where the data does not support it.
The two pieces of work are in different tumour types, and the reason is the same fact in both directions: no public serial-ctDNA cohort exists for the disease where the clinical case is strongest.
The clinical case is strongest in HR+/HER2− metastatic breast cancer. It is the largest biomarker-defined solid-tumour population, resistance to first-line CDK4/6-inhibitor plus endocrine therapy arrives through competing, identifiable mechanisms — ESR1 in roughly a third of progressors, with PIK3CA, RB1 loss and HER2-activating mutations among the rest — and PADA-1 has already shown that acting on a rising ESR1 ctDNA signal ahead of imaging is worth something. So that is where the method is demonstrated. There is no public serial cohort in it, which is why that demonstration is on synthetic trajectories from an openly released simulator, and why every CP‑BLG number on this site is labelled synthetic, five seeds.
The public data that does exist is single-timepoint, and the deepest of it is EGFR-mutant NSCLC on first-line osimertinib. So that is where the benchmark is built — and what OncoTraj measures there is the ceiling of the modality itself. Its negative result is the argument that the serial cohort has to be collected.
The Span detector has not been run on OncoTraj, and cannot be. The detector reads a series of detection calls over time; OncoTraj v1 contains snapshots. Nothing on this page is a Span-versus-baseline result on real patients, and the missing middle — a retrospective serial-ctDNA arm where the change-point test is scored against a naive rising-value rule — is the experiment we are looking for a partner to run. What that experiment needs, and what comes back
This is the study that turns everything above from a method into a result, or ends it. It is retrospective, it needs no new samples and no protocol change, and it can be specified completely before anyone runs it — which is the only reason it is worth running.
Snapshot baselines on the locked test split, scored two ways: once on the mixed-source split, and once within a single source. The gap between those two numbers is the result.
The three charts below are drawn directly from the benchmark's own evaluation reports (oncotraj-eval schema 1.1.0), not transcribed; the Task C table further down is transcribed by hand, and says so. Task A is scored on the 110 landmark-evaluable patients, Tasks B and C on the 85 with an observed event. The within-source analysis re-fits on MSK-CHORD alone and is scored on n = 101 — more patients than the mixed split, not fewer, because the mixed timing evaluation requires an observed event from every source, while the MSK-CHORD refit simply uses that source's own eligible set. Intervals are bootstrap 95%.
Trained and scored across all three sources, the timing task looks like it is working: random forest reaches a C‑index of 0.656, well clear of the 0.500 floor. Re-fit and scored inside MSK‑CHORD alone, where dataset membership can no longer stand in for biology, the same model falls to 0.432 — below chance, with an interval that contains it.
Nothing about the model changed. What changed is that it lost access to a shortcut: on the mixed split the date-less FLAURA subset carries a constant pseudo-PFS, so knowing which cohort a patient came from is worth more than knowing anything about the patient. Every learned model loses its apparent edge the same way. The majority floor, having never had one, is unmoved.
On the mixed split the three learned models sit above chance and their intervals exclude 0.500. We do not report that as a finding, because the same confound applies: part of what looks like discrimination is a model recognising which dataset a patient came from. The nearest within-source figure is the 0.596 [0.478, 0.714] quoted above, whose interval crosses chance — but that number comes from the benchmark's leaderboard prose and is not reproducible from the committed eval reports, so it is not charted here and should not be treated as a scored result.
Read this chart as an upper bound on what snapshots can do here — not as evidence that they work.
Discrimination is not the only thing a clinical model owes you; the probability has to mean something. Here the ranking inverts. The model that learns nothing is the best calibrated on the split (ECE 0.017), and every learned model is worse — XGBoost worst at 0.159, roughly nine times the floor.
A model can rank patients correctly and still be badly wrong about how likely progression is. At this cohort size, neither property is established.
| Model | ROC-AUC | 95% CI | Accuracy | F1 | Brier | ECE |
|---|---|---|---|---|---|---|
| Majority baseline | 0.500 | 0.500–0.500 | 0.564 | 0.000 | 0.246 | 0.017 |
| Logistic regression | 0.680 | 0.581–0.781 | 0.636 | 0.636 | 0.220 | 0.071 |
| Random forest | 0.678 | 0.569–0.773 | 0.636 | 0.667 | 0.214 | 0.041 |
| XGBoost | 0.630 | 0.525–0.739 | 0.582 | 0.610 | 0.255 | 0.159 |
| Model | C-index (mixed) | C-index (within) | 95% CI (within) | MAE d (within) |
|---|---|---|---|---|
| Majority baseline | 0.500 | 0.500 | 0.500–0.500 | 303.0 |
| Logistic regression | 0.581 | 0.492 | 0.417–0.570 | 301.5 |
| Random forest | 0.656 | 0.432 | 0.360–0.514 | 292.9 |
| XGBoost | 0.643 | 0.441 | 0.361–0.520 | 312.3 |
| Cox PH | — | 0.541 | 0.458–0.616 | 513.5 |
Within source, the best mean absolute error is 292.9 days against a 303.0-day majority floor — a ten-day gap on a year-scale prediction. Cox ranks marginally best on C-index and is the worst on absolute error by 200 days.
Transcribed, not generated. Unlike Tasks A and B, Task C is not in data/oncotraj.json, is not charted, and has no within-source re-scoring: these four rows are typed out of the benchmark write-up.
| Model | Accuracy | Macro-F1 | F1 on EGFR C797S |
|---|---|---|---|
| Majority baseline | 0.988 | 0.497 | 0.000 |
| Logistic regression | 0.835 | 0.517 | 0.125 |
| Random forest | 0.953 | 0.488 | 0.000 |
| XGBoost | 0.965 | 0.491 | 0.000 |
This is the row that matters. The highest accuracy on this task — 98.8% — belongs to the model that never predicts the actionable class at all; it answers "other" every time. Only logistic regression scores above zero on EGFR C797S, at 0.125. Accuracy here is an artefact of class imbalance, which is precisely why the benchmark reports macro-F1 beside it, and why tissue NGS under-capturing C797S makes this a structural non-result rather than a modelling failure.
If snapshots could predict resistance, you wouldn't need us. OncoTraj shows, rigorously, on real patients, that they can't. The signal lives in the trajectory, which is exactly what the Span detector is built to read.
The full write-ups behind the method and the benchmark, open to read, run, and check.
Aarchi Singh Thakur, Abhijoy Sarkar · arXiv:2606.11876 · June 2026
Abhijoy Sarkar, Aarchi Singh Thakur · arXiv:2606.11144 · June 2026