Peak Phase 9 — Frozen Evidence Evaluation: Measuring What Phase 8 Shipped
Phase 8 shipped the path. Phase 9 measured it. An architecture without a frozen comparison is a promise; an ablation on identical rows is a result.
We previously published an evidence-first extract path behind EVIDENCE_TILED_EXTRACT and EVIDENCE_TARGETED_RECOUNT (default off). That was Phase 8: overview semantics, native-resolution overlapping tiles, rotation-safe display coordinates, seam de-duplication, optional targeted recount, and review priority. We explicitly refused a product uplift claim until a pre-registered freeze compared one-shot control to that path on the same drawable ground-truth rows. Phase 9 is that evaluation (Tasks BA–BH). Numbers below trace to .tmp/phase9/ — not re-invented for the write-up.
Headline: ablation ran; defaults stay off
| Decision | Result |
|---|---|
| Ablation protocol (BA–BF) | Ran on frozen Peak P&D rows + FS/MVAC hold + N23/CPS INFO |
| P&D tiled vs oneshot | Det +38.0 pp; QA@20 +7.0 pp — evidence wins |
| Peak pooled (evidence composition) | Det 78.2% / QA@20 38.0% — matches Phase 6 exactly |
| Targeted recount | P&D QA@20 21.1%→20.4% (−0.7 pp) — do not promote |
| Promote either flag default-on? | No — cost/latency, offline FS/MVAC, no new product uplift |
| Regression / CI floors | PASS; floors remain Phase 6 |
Headline KPIs remain detection and QA@20 on every drawable ground-truth line — the honest Phase 5 denominator. We do not revive legacy soft-among-eligible as the public score.
Control vs evidence on identical P&D rows
Task BA froze the oneshot control and the 142 drawable P&D row IDs. Task BB scored control against the tiled evidence arm on those exact rows under fixed A4 aggregation (pipe/legend max). The evidence arm reuses the measured Phase 6 live tiles (11 pages, maxEdge 1800 under FORCE_XAI) rather than a non-deterministic re-spend — a freeze of inventory, not a new live product-path claim.
| Arm | Detection | QA@20 | QA@0 | Cov@P90 | Paired / 142 |
|---|---|---|---|---|---|
| Control (oneshot) | 33.8% (48) | 14.1% (20) | 8.5% | 0% | 48 |
| Evidence tiled | 71.8% (102) | 21.1% (30) | 14.1% | 0% | 102 |
| Δ | +38.0 pp | +7.0 pp | +5.6 pp | 0.0 | +54 |
| Tiles + recount | 71.8% (102) | 20.4% (29) | 14.1% | 0% | 102 |
Family lift mirrors Phase 6: valves 2→21 matched, gullies 8→13, pipe fittings 4→14 (22 of 36 still missed). Fixtures and sanitary remain zero. Detection still outruns usable quantity — pairing a symbol does not put the count inside ±20%. Cov@P90 on P&D stays 0%: there is still no confidence threshold that yields a ≥90%-precision auto-accept slice of the plumbing bill.

Peak pooled: both arms, one honest merge
| Arm | Composition | Drawable | Detection | QA@20 |
|---|---|---|---|---|
| Control | P&D oneshot + FS offline + MVAC offline | 234 | 55.1% (129) | 33.8% (79) |
| Evidence tiled | P&D tiles + FS offline + MVAC offline | 234 | 78.2% (183) | 38.0% (89) |
| Evidence + recount | P&D tiles+recount + FS/MVAC offline | 234 | 78.2% | 37.6% (88) |
| Phase 6 published | Same as evidence tiled | 234 | 78.2% | 38.0% |
The evidence composition matches Phase 6 by construction. Phase 9 freezes that measurement; it does not invent a new pooled gain. Control pooled detection 55.1% is the oneshot P&D composition — report it beside evidence, do not treat it as a regression of the shipped tiled path. FS and MVAC were offline A4 holds (0.0 pp). EL&ELV remains honest N/A (no EL PDF in Peak training). Stretch pooled QA@20 ≥42% is still not earned (38.0%).
Cost, latency, and tile count
| Arm | Tiles / images | Latency | Est. cost (USD) |
|---|---|---|---|
| Control oneshot | 11 page images | n/a (cache reuse) | n/a (cache) |
| Evidence tiled | 132 planned / 127 extracted | ~25.1 min | ~$0.79 |
| + targeted recount | +88 recount calls | +~50.5 min | +~$0.21 |
| Tiles + recount total | — | ~75.5 min | ~$1.00 |
That cost/latency profile is why BF and BG refuse default-on even though tiled evidence beats oneshot on accuracy. Silent production default on every drawing enter would burn dollars and minutes without a product budget acceptance. Recount adds ~$0.21 and half an hour while lowering P&D QA@20 by 0.7 pp (reviewRate 0.716) — promoteRecount stays false.
FS / MVAC hold and cross-project INFO
| Sheet / project | Detection | QA@20 | vs Phase 6 | Label |
|---|---|---|---|---|
| Peak FS (offline) | 81.4% | 64.4% | 0.0 pp | HOLD — not live tiles |
| Peak MVAC (offline) | 100% | 63.6% | 0.0 pp | HOLD — not live tiles |
| N23 (9 drawable) | 88.9% | 77.8% | 0.0 pp | INFO — frozen re-agg |
| CPS (13 drawable) | 61.5% | 38.5% | 0.0 pp | INFO — parser confounded |
N23 remains offline re-aggregation of a frozen extract (QA@0 66.7%). CPS stays on the bespoke PDF BQ parser path; Phase 7 shared ingestion is not the published scorer. Neither project enters Peak CI floors. Evidence-path arms were not run on N23 or CPS in this phase.

Regression and promote path
Task BG reports overall PASS against Phase 6 floors on the evidence composition (pooled det ≥76.2%, QA@20 ≥36.0%, P&D QA@20 ≥19.1%, FS/MVAC holds). Floors are not raised: promote requires BF yes on a flag plus a clean product-path win. Both flags stay opt-in. Future promote for tiles needs a fresh product-default live run (drawing-extract runTiledEvidencePass), an accepted cost budget, and FS/MVAC live tile holds. Recount needs an accuracy win that pays for its extra cost — BD currently fails both.
Phase 10 backlog
- Legend-seeded clustering — cut plan/legend double-count after detection recovers symbols.
- Cov@P90 confidence calibration — P&D is still ~0%; no high-precision auto-accept slice.
- Kwai On quantity benchmark — classification fixed at 84.5%; full QA@20/det audit open.
- FS/MVAC live evidence tiles — required before default-on can touch those sheets.
- CPS bill parser graduation from INFO — audited shared ingestion or frozen GT delta table.
- Corrections-bank rebalance after Peak GT v2 — Phase 7 heading repair left few-shots on pre-repair labels.
Phase 9 closes the evaluation gate Phase 8 opened. The evidence path works better than oneshot on Peak P&D when you pay for tiles; it is not yet the right silent default; recount is not a free lunch. The next accuracy claims will come from the Phase 10 residuals — and each will need the same discipline: freeze rows, report both arms, and refuse uplift without measurement.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read