Blog
EngineeringAI AccuracyResearchPeakPhase 9Evidence-firstAblation

Peak Phase 9 — Frozen Evidence Evaluation: Measuring What Phase 8 Shipped

Teraquant Team10 min read
Phase 8 shipped the path. Phase 9 measured it. An architecture without a frozen comparison is a promise; an ablation on identical rows is a result.

We previously published an evidence-first extract path behind EVIDENCE_TILED_EXTRACT and EVIDENCE_TARGETED_RECOUNT (default off). That was Phase 8: overview semantics, native-resolution overlapping tiles, rotation-safe display coordinates, seam de-duplication, optional targeted recount, and review priority. We explicitly refused a product uplift claim until a pre-registered freeze compared one-shot control to that path on the same drawable ground-truth rows. Phase 9 is that evaluation (Tasks BA–BH). Numbers below trace to .tmp/phase9/ — not re-invented for the write-up.

Headline: ablation ran; defaults stay off

DecisionResult
Ablation protocol (BA–BF)Ran on frozen Peak P&D rows + FS/MVAC hold + N23/CPS INFO
P&D tiled vs oneshotDet +38.0 pp; QA@20 +7.0 pp — evidence wins
Peak pooled (evidence composition)Det 78.2% / QA@20 38.0% — matches Phase 6 exactly
Targeted recountP&D QA@20 21.1%→20.4% (−0.7 pp) — do not promote
Promote either flag default-on?No — cost/latency, offline FS/MVAC, no new product uplift
Regression / CI floorsPASS; floors remain Phase 6

Headline KPIs remain detection and QA@20 on every drawable ground-truth line — the honest Phase 5 denominator. We do not revive legacy soft-among-eligible as the public score.

Control vs evidence on identical P&D rows

Task BA froze the oneshot control and the 142 drawable P&D row IDs. Task BB scored control against the tiled evidence arm on those exact rows under fixed A4 aggregation (pipe/legend max). The evidence arm reuses the measured Phase 6 live tiles (11 pages, maxEdge 1800 under FORCE_XAI) rather than a non-deterministic re-spend — a freeze of inventory, not a new live product-path claim.

ArmDetectionQA@20QA@0Cov@P90Paired / 142
Control (oneshot)33.8% (48)14.1% (20)8.5%0%48
Evidence tiled71.8% (102)21.1% (30)14.1%0%102
Δ+38.0 pp+7.0 pp+5.6 pp0.0+54
Tiles + recount71.8% (102)20.4% (29)14.1%0%102

Family lift mirrors Phase 6: valves 2→21 matched, gullies 8→13, pipe fittings 4→14 (22 of 36 still missed). Fixtures and sanitary remain zero. Detection still outruns usable quantity — pairing a symbol does not put the count inside ±20%. Cov@P90 on P&D stays 0%: there is still no confidence threshold that yields a ≥90%-precision auto-accept slice of the plumbing bill.

Peak FS page 1 with integrity-safe evidence pins (prior annotate-benchmark path)
Integrity-safe annotation from scripts/annotate-benchmark.ts (no demo fallback). Phase 9 does not invent new synthetic overlays. FS metrics held offline at Phase 6 levels under the BC hold — not a live tile claim on FS.

Peak pooled: both arms, one honest merge

ArmCompositionDrawableDetectionQA@20
ControlP&D oneshot + FS offline + MVAC offline23455.1% (129)33.8% (79)
Evidence tiledP&D tiles + FS offline + MVAC offline23478.2% (183)38.0% (89)
Evidence + recountP&D tiles+recount + FS/MVAC offline23478.2%37.6% (88)
Phase 6 publishedSame as evidence tiled23478.2%38.0%

The evidence composition matches Phase 6 by construction. Phase 9 freezes that measurement; it does not invent a new pooled gain. Control pooled detection 55.1% is the oneshot P&D composition — report it beside evidence, do not treat it as a regression of the shipped tiled path. FS and MVAC were offline A4 holds (0.0 pp). EL&ELV remains honest N/A (no EL PDF in Peak training). Stretch pooled QA@20 ≥42% is still not earned (38.0%).

Cost, latency, and tile count

ArmTiles / imagesLatencyEst. cost (USD)
Control oneshot11 page imagesn/a (cache reuse)n/a (cache)
Evidence tiled132 planned / 127 extracted~25.1 min~$0.79
+ targeted recount+88 recount calls+~50.5 min+~$0.21
Tiles + recount total~75.5 min~$1.00

That cost/latency profile is why BF and BG refuse default-on even though tiled evidence beats oneshot on accuracy. Silent production default on every drawing enter would burn dollars and minutes without a product budget acceptance. Recount adds ~$0.21 and half an hour while lowering P&D QA@20 by 0.7 pp (reviewRate 0.716) — promoteRecount stays false.

FS / MVAC hold and cross-project INFO

Sheet / projectDetectionQA@20vs Phase 6Label
Peak FS (offline)81.4%64.4%0.0 ppHOLD — not live tiles
Peak MVAC (offline)100%63.6%0.0 ppHOLD — not live tiles
N23 (9 drawable)88.9%77.8%0.0 ppINFO — frozen re-agg
CPS (13 drawable)61.5%38.5%0.0 ppINFO — parser confounded

N23 remains offline re-aggregation of a frozen extract (QA@0 66.7%). CPS stays on the bespoke PDF BQ parser path; Phase 7 shared ingestion is not the published scorer. Neither project enters Peak CI floors. Evidence-path arms were not run on N23 or CPS in this phase.

Peak MVAC page 1 with coordinate-consistent evidence and measured runs
Regenerated integrity-safe Peak AC page 1 (prior live cache). MVAC Phase 9 metrics held offline; purple pins and cyan 量度 polylines are real model geometry, not a fabricated demo grid.

Regression and promote path

Task BG reports overall PASS against Phase 6 floors on the evidence composition (pooled det ≥76.2%, QA@20 ≥36.0%, P&D QA@20 ≥19.1%, FS/MVAC holds). Floors are not raised: promote requires BF yes on a flag plus a clean product-path win. Both flags stay opt-in. Future promote for tiles needs a fresh product-default live run (drawing-extract runTiledEvidencePass), an accepted cost budget, and FS/MVAC live tile holds. Recount needs an accuracy win that pays for its extra cost — BD currently fails both.

Phase 10 backlog

  • Legend-seeded clustering — cut plan/legend double-count after detection recovers symbols.
  • Cov@P90 confidence calibration — P&D is still ~0%; no high-precision auto-accept slice.
  • Kwai On quantity benchmark — classification fixed at 84.5%; full QA@20/det audit open.
  • FS/MVAC live evidence tiles — required before default-on can touch those sheets.
  • CPS bill parser graduation from INFO — audited shared ingestion or frozen GT delta table.
  • Corrections-bank rebalance after Peak GT v2 — Phase 7 heading repair left few-shots on pre-repair labels.

Phase 9 closes the evaluation gate Phase 8 opened. The evidence path works better than oneshot on Peak P&D when you pay for tiles; it is not yet the right silent default; recount is not a free lunch. The next accuracy claims will come from the Phase 10 residuals — and each will need the same discipline: freeze rows, report both arms, and refuse uplift without measurement.