News that knows you · never shipped
A daily edition is mostly the stories you never see.
1,362 candidates became 50. Other feeds show you the survivors.This one shows what was thrown away, and where.
The sieve
One cell per candidate articleReader fixture
Still in the pipeline
1,362
of 1,362100.0%
Stage 01 / 10 · pool
Corpus pool
Nothing removed yet.
Every article in the frozen snapshot, plus the needles planted for this persona.
Where these counts come from
Recorded survivorship for fixture Ray against corpus 2026-09-02, reconstructed from the scorecard’s stage tallies and asserted against three independently recorded facts in funnel.test.ts. Which particular cell a stage took is not recorded anywhere, so cells are scattered by a fixed hash — the quantities are evidence, the positions are not.
Why a ruler exists
“The feed looks better to me” is not evidence.Same corpus, same fixtures, network off — so two runs are comparable, and a fix that made things worse cannot hide.
- Stored runs
- 7imported, not re-executed
- Frozen corpora
- 33,920 articles
- Reader fixtures
- 10adversarial, not users
- Replay spend
- $0cached, network off
Every fixture, no averaging
prod-llm · 2026-09-02- Anna37.5%17 labelled story placements for fixture Anna: 8 delivered, wanted, 4 delivered, unwanted, 5 lost before the scorer. Capped recall 37.5 percent.
- Cold start8.3%23 labelled story placements for fixture Cold start: 9 delivered, wanted, 3 delivered, unwanted, 4 lost at the scorer, 7 lost before the scorer. Capped recall 8.3 percent.
- Daniel20.0%16 labelled story placements for fixture Daniel: 1 delivered, wanted, 1 delivered, unlabelled, 10 delivered, unwanted, 4 lost before the scorer. Capped recall 20.0 percent.
- Frank18.2%21 labelled story placements for fixture Frank: 2 delivered, wanted, 9 delivered, unlabelled, 1 delivered, unwanted, 1 lost at the scorer, 8 lost before the scorer. Capped recall 18.2 percent.
- Lena16.7%27 labelled story placements for fixture Lena: 10 delivered, wanted, 2 delivered, unwanted, 15 lost before the scorer. Capped recall 16.7 percent.
- Maya16.7%25 labelled story placements for fixture Maya: 11 delivered, wanted, 1 delivered, unlabelled, 1 lost at the scorer, 12 lost before the scorer. Capped recall 16.7 percent.
- Paula40.0%15 labelled story placements for fixture Paula: 2 delivered, wanted, 5 delivered, unlabelled, 5 delivered, unwanted, 1 lost at the scorer, 2 lost before the scorer. Capped recall 40.0 percent.
- Ray30.0%19 labelled story placements for fixture Ray: 7 delivered, wanted, 5 delivered, unwanted, 7 lost before the scorer. Capped recall 30.0 percent.
- Tom33.3%22 labelled story placements for fixture Tom: 7 delivered, wanted, 1 delivered, unlabelled, 1 delivered, unwanted, 1 lost at the scorer, 12 lost before the scorer. Capped recall 33.3 percent.
- Will0.0%29 labelled story placements for fixture Will: 5 delivered, wanted, 6 delivered, unlabelled, 1 delivered, unwanted, 1 lost at the scorer, 16 lost before the scorer. Capped recall 0.0 percent.
- Delivered, wanted
- Delivered, unlabelled
- Delivered, unwanted
- Lost at the scorer
- Lost before the scorer
Bar length is labelled story placements; the right column is capped recall at k, mean 22.1%. A fixture at 0.0% received none of the stories its labels said it needed.
An average is a way of not looking at the worst case.The mean is 22.1%. Two fixtures are at zero.
Open the explorerOne recorded batch
40 sent · 254 returned40 articles were sent and 254 verdicts came back. The pairs below show which verdict the scorer attached to which article.
Nothing in the response names an article — position is the only join.
Fixing it made the measured numbers worse. The baseline was not re-recorded.
Read the defect report- Guard fires
- 63one run, ten fixtures
- Worst response
- 201verdicts for 40 articles
- Unwanted rate
- +15.6points, after the fix
- Re-baselined
- Nothe gate still reports red
What this site never claims
No readers
Never shipped. The ten profiles are evaluation fixtures, not people, and reading here creates no data.
Provisional ground truth
Labels are model-written with an agent pass. Human review is outstanding, so absolute values are provisional.
No significance
No confidence intervals; none were computed. “Material” is a fixed ±0.02 the harness’s author chose.
Live mode unbuilt
Signing in is not implemented. Live end-to-end behaviour is unverified.
Unknown stays unknown
What the evidence cannot establish is recorded as unknown — never null, never zero.
$0 of model spend
Costs shown are reconstructed token-equivalents. Actual provider spend is zero.