Peak Phase 4 — Accuracy Recovery & Cross-Project Validation
ERRATUM (2026-08-06): the N23 overlay images in this post were produced by an ad-hoc script that was never committed, so they cannot be reproduced. They are real renders, but the underlying detection coordinates are largely placeholder grids — evenly spaced round percentages such as 25/45/65 by 28/55 — rather than grounded positions, so the pins do not reliably indicate where an item was found. Detection pin positions from this post should not be relied on. The numeric N23 results (detection 88.9%, soft 50.0%, exact 44.4%) come from the scored benchmark artefacts and are unaffected. A committed, reproducible renderer is described in Peak Phase 5.
SUPERSEDED FOR PRODUCT CLAIMS: After Phase 5, legacy soft among soft-eligible pairs is not the programme headline. The pooled soft 70.7% below (58/82 soft-eligible pairs) is a historical Phase 4 score — Phase 5 re-measured honest Peak QA@20 on all drawable rows as 24.8% (then 33.8% after aggregation fixes), not that soft ratio on 82 of 234 rows. For fixed-denominator Detection and QA@20, read /blog/peak-phase5-measuring-what-we-ship and /blog/peak-phase6-accuracy-sprint. N23 pin-position erratum above remains authoritative.
Peak Phase 1 fixed classification (90.2% of 512 qty>0 SOR rows). Phase 2 measured quantity soft among soft-eligible pairs (~68.6%) and found detection only 23.8% — 402 misses dominated. Phase 3 attacked detection first: weighted detection 54.3%, pooled soft 60.0%, EL honest N/A, FS soft 70.7%, MVAC soft 80.8% / det 100% — but P&D soft collapsed 72.0%→17.4% while detection rose 23.2%→32.4%. Phase 4 is the recovery and generalization cycle: clean false pairs, de-dupe pipe double-counts across drawings, seed few-shot corrections, gate CI, rescore, and run the first non-Peak Chinese P&D project (N23). Numbers below come from verified artefacts under .tmp/peak-phase4/ and .tmp/n23-benchmark/ — not re-invented for the write-up.
Soft recovery without detection theatre: P&D soft 17.4%→53.3% on a cleaned 15-pair set; detection still 46/142 = 32.4%. Soft is not “53% of the whole plumbing SOR.”
Phase 3 → Phase 4 headline table
Overall soft and detection exclude EL&ELV when it is honest N/A (no Peak EL PDFs). Soft % is among soft-eligible pairs (±20%, excluding length_scope). Detection and exact are on drawable ground-truth rows.
| Metric | Phase 3 | Phase 4 | Δ / note |
|---|---|---|---|
| Overall soft (pooled) | 60.0% (54/90) | 70.7% (58/82) | +10.7 pp |
| Overall soft (sheet mean) | 56.3% | 68.3% | +12.0 pp |
| Overall detection (weighted) | 54.3% | 54.3% | Hold (127/234) |
| Overall exact (mean) | 33.9% | 34.6% | +0.7 pp |
| P&D soft | 17.4% (4/23) | 53.3% (8/15) | +35.9 pp; goal ≥50% met |
| P&D detection | 32.4% | 32.4% | Hold |
| FS soft / det | 70.7% / 81.4% | 70.7% / 81.4% | Hold |
| MVAC soft / det | 80.8% / 100% | 80.8% / 100% | Hold |
| EL&ELV | N/A | N/A | Honest — no EL PDFs |
| Z7 regression | FAIL (P&D soft) | PASS (all gates) | vs Phase 3 floors |
Per-sheet scorecard
| Sheet | Soft P3 → P4 | Det P3 → P4 | Exact P3 → P4 | Soft-eligible P4 | Notes |
|---|---|---|---|---|---|
| EL&ELV | N/A | N/A | N/A | — | no_el_drawings |
| P&D | 17.4% → 53.3% | 32.4% → 32.4% | 2.8% → 4.9% | 8/15 | Z1 quality gate + Z2 pipe max |
| FS | 70.7% → 70.7% | 81.4% → 81.4% | 47.5% → 47.5% | 29/41 | Held; 5× floor scale |
| MVAC | 80.8% → 80.8% | 100% → 100% | 51.5% → 51.5% | 21/26 | Best balanced; held |
Phase 4 goals from the Z6 summary all met: P&D soft ≥50%, P&D detection hold ≥32%, FS soft hold, MVAC soft hold, overall pooled soft ≥65% (actual 70.7%), N23 baseline recorded, CI gate pass. Stretch overall soft >80% remains open — detection mass on P&D and length_scope industrial metres still dominate the path.
What Z1–Z5 shipped
- Z1 — Soft quality gate: isSoftQualityPair / softEligible filters description pairs that pass confidence but fail quantity-quality heuristics after Pass 4; reshow-max for sanitary/gully/vent. P&D soft 17.4%→53.3%; eligible 23→15; detection unchanged.
- Z2 — Cross-drawing pipe de-dupe: pipeAcrossDrawings max for P&D (FS/MVAC still sum). CI 100 mm AI qty 2823→1515; top-15 sum |Δ| 10849→7884 (−27.3%). Network junction merge remains partial.
- Z3 — N23 benchmark harness (scripts/benchmark-n23.ts): first Chinese renovation P&D project outside Peak.
- Z4 — Corrections bank seed: 15 curated entries in apps/web/data/corrections-seed.json; ## Prior corrections in extract prompts (default ON; CORRECTIONS_SEED=0 off). Offline rescore does not re-run live vision with seed.
- Z5 — CI floors: scripts/accuracy-floors.json + rescore --ci. Buffered per-sheet min soft and min detection; EL skipped as N/A.
N23 cross-project baseline
N23 is 澳門大學—N23科研大樓六樓應用物理及材料工程研究院辦公室建造工程: a small 6/F renovation BOQ in Chinese, plumbing and drainage only. Seven drawing pages, fourteen SOR lines (nine drawable, five provisional 項 lump-sum). This is the first public cross-project quantity baseline outside Peak training.
| Metric | Value | Denominator |
|---|---|---|
| Detection | 88.9% | 8 / 9 drawable GT paired |
| Soft (±20%) | 50.0% | 4 / 8 soft-eligible |
| Exact | 44.4% | 4 / 9 drawable exact |
| Page-comparable soft | 80.0% | Count items only |
| Classification (drawable) | 100% | 9 / 9 families correct |
Drawable item outcomes: valves (Ø15 angle, Ø25 gate), stainless sink set, and S-type trap matched exactly. Pipe metres overcounted about ×6 (copper Ø15 5→30 m, copper Ø25 10→72 m, uPVC Ø50 12→72 m) — multi-page sum of re-shows and existing services vs renovation BQ scope. Floor drain 1→16; bottle trap missed (gt_only). High ai_only (61) is expected: drawings show building services outside the BQ.
N23 shows detection and Chinese classification transfer; soft fails where Peak already fails — pipe scope across pages, not family naming.
Per-item drawable outcomes (from .tmp/n23-benchmark/results.json). Provisional 項 lines (demolition, connections, paint, T&C) are correctly excluded from the soft denominator — they are not drawable quantity claims.
| ID | Description | SOR | AI | Verdict |
|---|---|---|---|---|
| 2.1.1 | Ø15mm 包膠銅管 | 5 m | 30 m | ai_over |
| 2.1.2 | Ø25mm 包膠銅管 | 10 m | 72 m | ai_over |
| 2.2.1 | Ø15mm 銅角閥 | 2 | 2 | match |
| 2.2.2 | Ø25mm 銅閘閥 | 1 | 1 | match |
| 2.3 | 不鏽鋼洗滌盆 | 1 set | 1 | match |
| 3.1 | Ø50mm UPVC 排水喉 | 12 m | 72 m | ai_over |
| 3.2.1 | S型存水彎 Ø100 | 1 | 1 | match |
| 3.2.2 | 瓶型存水彎 Ø100 | 1 | — | gt_only |
| 3.3 | 防臭地漏 Ø100 | 1 | 16 | ai_over |
Aggregation compressed 104 raw AI items to 69 (10 merges). Classification on drawable rows is 100% (valve, sanitary_fixture, copper_pipe, upvc_pipe, pipe_fitting, gully). The 64.3% rate on all 14 SOR lines includes five provisional 項 — not comparable to Peak’s 90.2% on 512 qty>0 lines.




Failure taxonomy (Phase 4 Peak)
Measurable sheets only (excl. EL N/A audit rows). Same failureType rules as Phase 2/3.
| Sheet | miss | under | over | length_scope | near_miss | match | ai_only |
|---|---|---|---|---|---|---|---|
| P&D | 100 | 11 | 5 | 22 | 1 | 7 | 127 |
| FS | 16 | 2 | 10 | 7 | 1 | 28 | 11 |
| MVAC | 3 | 4 | 1 | 7 | 4 | 17 | 22 |
| Measurable total | 119 | 17 | 16 | 36 | 6 | 52 | 160 |
Versus Phase 3: P&D hard over 9→5 and match 4→7 track soft recovery + max de-dupe. Miss mass (119 measurable) is unchanged — detection is still the long pole. length_scope (36) still holds the industrial |Δ| bulk (uPVC/CI/concrete/GI/refrigerant metres).
Top remaining Peak failures by |Δ|
Excludes EL N/A, exact match, near_miss, and ai_only. Source: .tmp/peak-phase4/combined-results.json top15 (first 8 shown). Phase 3 #1 CI 100 mm |Δ| was 2190 (AI 2823); Phase 4 is 882 (AI 1515) after max-across-drawings.
| # | Sheet | Type | Family | SOR | AI | |Δ| |
|---|---|---|---|---|---|---|
| 1 | P&D | length_scope | upvc_pipe 100mm | 167 | 1296 | 1129 |
| 2 | P&D | length_scope | ci_pipe 100mm | 633 | 1515 | 882 |
| 3 | P&D | miss | copper_pipe 28mm | 740 | — | 740 |
| 4 | P&D | length_scope | concrete_pipe 100mm | 230 | 915 | 685 |
| 5 | MVAC | length_scope | refrigerant 9.52mm | 635 | 80 | 555 |
| 6 | FS | length_scope | gi_pipe 150mm | 200 | 698 | 498 |
| 7 | P&D | length_scope | copper_pipe 22mm | 489 | 0 | 489 |
| 8 | FS | length_scope | gi_pipe 32mm | 750 | 1200 | 450 |
Regression & CI (Task Z7)
- Domain package tests 142/142 PASS; web build PASS; offline Peak classification 90.2% PASS (≥88%).
- Per-sheet floors vs Phase 3 publish: P&D soft ≥17.4% (actual 53.3%), det ≥32.4%; FS soft ≥70.7% / det ≥81.4%; MVAC soft ≥80.8% / det ≥99% (actual 100%). All PASS.
- CI buffered floors (accuracy-floors.json) also PASS. N23 is INFO only — first baseline, no prior floor.
Phase 5 research directions
Phase 4 closed the soft-collapse emergency and proved a second project can run the same compare harness. Stretch soft >80% needs detection mass and scope honesty, not more percentage theatre on tiny pair sets.
- P&D detection 32.4%→≥50%: legend-seeded symbol clustering for fittings, valves, gullies; optional live re-extract with corrections seed (offline rescore cannot measure seed impact).
- Page-role / BOQ-scope filtering: plan vs riser vs schedule vs existing-services pages before length sum — the shared Peak + N23 failure mode.
- Finish network-aware polyline merge at junctions; convert some length_scope into real soft or clearer non-comparable labels.
- P&D soft stretch toward published Phase 2 72% without re-admitting false pairs; do not claim Phase 2 soft restored at 53.3%.
- EL PDF acquisition — keep N/A until schedule-bearing electrical drawings exist.
- Promote N23 from INFO baseline to regression project after a second run with page-role fixes; grow corrections bank from human claim-correct events.
Conclusion
Peak Phase 4 is complete as a research cycle. Soft recovery on P&D (17.4%→53.3%) with held detection and held FS/MVAC raised pooled soft 60.0%→70.7%. Cross-drawing pipe max cut industrial overcount without inventing detections. N23 establishes det 88.9% / soft 50% / exact 44.4% on a Chinese renovation P&D BOQ. Z7 regression and CI gates are green. Classification remains 90.2%. Verification-first means publishing the cleaned soft denominator, the unchanged miss mass, and the Phase 5 path — not a single victory percentage. Companion posts: Phase 3 /blog/peak-phase3-accuracy-deep-dive; Phase 2 /blog/peak-phase2-quantity-audit; Phase 1 /blog/peak-full-sor-audit. Full internal report: docs/PEAK_PHASE4_ACCURACY_REPORT.md.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read