Peak Phase 6 — Accuracy Sprint Across Peak, N23, and Central Police Station
Detection without a usable quantity is incomplete. Quantity without detection is fiction. Phase 6 measures both on every drawable ground-truth line — and refuses to revive soft-among-eligible as the headline.
Phase 5 established that the honest Peak end-to-end quantity accuracy within ±20% (QA@20) was 33.8% on 234 drawable SOR lines, with detection stuck at 55.1%. The wall was physical: A0 sheets at ~7021×4967 px were compressed to a single 1800–2560 px JPEG before vision, so fittings and valves disappeared. Phase 6 is the detection-first live sprint: native-resolution overlapping tiles on Peak P&D, a fittings/valve prompt lift, a live corrections-seed A/B, plus two cross-project baselines (N23 offline Phase-5 re-score and Central Police Station first live). Numbers below come from .tmp/peak-phase6/, .tmp/n23-benchmark/phase5-results.json, and .tmp/cps-benchmark/results.json — not re-invented for the write-up.
Headline metrics (QA@20 on all drawable)
| Project | Drawable | Detection | QA@20 | Provenance |
|---|---|---|---|---|
| Peak pooled | 234 | 78.2% (183) | 38.0% (89) | P&D live tiles + FS/MVAC offline A4 |
| Peak P&D | 142 | 71.8% (102) | 21.1% (30) | live tiles, xAI grok-4.5 |
| Peak FS | 59 | 81.4% | 64.4% | Phase 3 cache hold |
| Peak MVAC | 33 | 100% | 63.6% | Phase 3 cache hold |
| N23 | 9 | 88.9% (8) | 77.8% (7) | offline A4 re-agg of frozen extract |
| CPS (INFO) | 13 | 61.5% (8) | 38.5% (5) | first live + PDF BQ |
Targets: pooled detection ≥65% met (78.2%); pooled QA@20 ≥42% missed (38.0%); P&D detection ≥45% met (71.8%); P&D QA@20 ≥22% missed by 0.9pp (21.1%). FS and MVAC held at 0.0pp by construction — they were not re-extracted live. EL&ELV remains honest N/A (no EL PDF in Peak training).
Phase 5 → Phase 6 Peak
| Sheet | Detection P5 → P6 | QA@20 P5 → P6 | Δ det | Δ QA@20 |
|---|---|---|---|---|
| Pooled | 55.1% → 78.2% | 33.8% → 38.0% | +23.1 pp | +4.3 pp |
| P&D | 33.8% → 71.8% | 14.1% → 21.1% | +38.0 pp | +7.0 pp |
| FS | 81.4% → 81.4% | 64.4% → 64.4% | 0.0 | 0.0 |
| MVAC | 100% → 100% | 63.6% → 63.6% | 0.0 | 0.0 |
All measured pooled gain is P&D live tiling. Do not read Phase 6 as a full four-sheet live re-extract.
Detection rose far more than QA@20 because pairing a symbol does not put the quantity inside ±20%. P&D AI-only rows ballooned to 692 under tiles — more detections also mean more unmatched noise and higher review load. Cov@P90 on P&D remains 0%: there is still no confidence threshold that yields a ≥90%-precision auto-accept slice of the plumbing bill. Legacy soft on P&D fell 60.6%→42.6% while detection rose — more soft-eligible pairs enter the denominator. That is why soft is not the headline.
What AA–AC changed
Task AA ran 11 Peak P&D pages as 3×4 overlapping tiles (maxEdge 1800, overlap 0.15, 127/132 tile extracts after five hard socket failures). Control (one-shot Phase 3 cache) vs tiles on the same A4 aggregation:
| Arm | Detection | QA@20 | Paired / 142 |
|---|---|---|---|
| Control (one-shot) | 33.8% | 14.1% | 48 |
| Tiles (live) | 71.8% | 21.1% | 102 |
| Δ | +38.0 pp | +7.0 pp | +54 |
Family lift: valves 2→21 matched (of 23), gullies 8→13 (all), pipe fittings 4→14 (of 36 — 22 still missed), manholes and water tanks to full match. Pumps and provisional lines remain zero. Task AB (PROMPT_VERSION 1.1.8) rewrote fittings/valve COUNT guidance and added a generic-elbow noise filter; offline FS/MVAC held; live smoke showed a small plan-page row lift without elbow spam. Task AC measured the Z4 corrections seed ON vs OFF on Peak P page 1: seed ON did not help (runs 4→2, total m 88→50). Wiring works; efficacy on that short block plan was null/negative — overcount-oriented few-shots made the model more conservative.

N23 Phase-5 re-score
Task AD re-aggregated the frozen seven-page Chinese P&D extract under Phase 5 A4 options (pipeAcrossPages: max, legendOnlyCombine: max). No new live vision. The page-sum ×6 pipe error collapses:
| Metric | Z3 sum pages | A4 max pages |
|---|---|---|
| Detection | 88.9% (8/9) | 88.9% (8/9) |
| QA@20 | 44.4% (4/9) | 77.8% (7/9) |
| QA@0 | 44.4% (4/9) | 66.7% (6/9) |
Copper Ø15: 30 m → 5 m (match SOR 5). uPVC Ø50: 72 m → 12 m (match SOR 12). Residual fails are extract noise (gully 1→16; bottle trap miss), not aggregation. Gate: INFO only — not pooled into Peak CI floors.

Central Police Station first baseline
Task AE is the third project: English Block 4 drainage, DRAWINGS.pdf pages 4–10, heading-aware BQ.pdf (not xlsx). First live baseline — INFO floors only, no CI fail.
| Metric | Value | Notes |
|---|---|---|
| Drawable GT | 13 | 21 BQ rows − 8 provisional Item |
| Detection | 61.5% (8/13) | paired / drawable |
| QA@20 | 38.5% (5/13) | equals QA@0 on this run |
| Cov@P90 | 15.4% | selective coverage |
| Classification | 47.6% | 10 classified / 21 BQ rows |
Stresses PDF BQ structure, three identical bare “150 mm Diameter pipe.” rows under different parent headings, and English drainage vocabulary (EN 877 CI, D.I., stainless gutter). Honest misses include manhole, sump pit, and bare-Ø other pipes; over-counts hit CI runs, rainwater outlets, and gutter length. Sample variance is high at n=13.
What still fails
- P&D pipe fittings: 22 of 36 still gt_only after tiling — largest residual miss mass.
- Det→QA@20 gap: 71.8% detection vs 21.1% usable quantities on P&D.
- AI-only flood (692) raises review load; Cov@P90 stays 0% on P&D.
- Corrections seed not yet a measured quantity win (single-page null/negative).
- FS/MVAC not live-tiled; provider is xAI only from HK — no Gemini parity claim.
Regression and floors
Task AG regression matrix: PASS. Domain 169 tests green; web build OK; all Peak Phase 5 floors cleared with headroom. Floors raised to measured − 2pp for pooled detection (76.2%), pooled QA@20 (36.0%), and P&D QA@20 (19.1%). FS/MVAC floors held. N23 and CPS remain INFO. scripts/accuracy-floors.json records phase: 6.
What comes next
Phase 7 (BOQ ingestion integrity) already completed in parallel — corpus classification and Peak ground-truth repair — and did not change production defaults. Phase 8 ships an opt-in evidence-first path (overview → native tiles → display-coordinate evidence → seam de-dupe → targeted recount → review). Research directions after this sprint:
- Close the P&D det→QA@20 gap with schema-constrained targeted recount and legend-seeded fitting clusters.
- Calibrate quantity confidence so Cov@P90 on P&D is non-zero — selective prediction for QS auto-accept.
- Rebalance the corrections bank (miss-recovery vs overcount lectures) and re-measure on dense floor plans.
- Kwai On quantity baseline (classification already 84.5% held-out).
- Live FS/MVAC tile runs before claiming full multi-sheet live Peak metrics; CPS parser audit before CI floors.
- Phase 8 evaluation: freeze one-shot vs evidence-tiled extract on identical rows; report det, QA@20, cost, latency — no uplift claim without that ablation.
Companion posts: Phase 5 /blog/peak-phase5-measuring-what-we-ship; Phase 8 architecture /blog/phase8-evidence-first-drawing-intelligence; Phase 4 /blog/peak-phase4-cross-project-validation. Full internal report: docs/PEAK_PHASE6_ACCURACY_REPORT.md.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read