Peak Phase 13 — First-Principles Vector Geometry Did Not Move QA@20
A well-documented null on the idea we most wanted to work is a better outcome than a breakthrough that only exists in prose.
Phases 8–12 asked one question five times: can evidence-first VLM tiling (`EVIDENCE_TILED_EXTRACT` / recount) ship default-on? The answer was always no — cost envelope, live FS/MVAC regression, soft-hold recovery without dollar recovery, or recount cost not justified. Pooled detection climbed into the high seventies; QA@20 stayed flat near 37–38% and never cleared the Phase 6 stretch target of ≥42%. Phase 13 deliberately changed axis. It did not reopen those flags. Numbers below trace to `.tmp/phase13/` — not re-invented for the write-up.
Headline: primary hypothesis not confirmed
| Decision / metric | Result |
|---|---|
| Vector geometry primary arm (BA 142 P&D) | det 71.8% / QA@20 21.1% both arms — Δ 0.0 / 0.0 pp |
| replace_if_confident ablation | REGRESS — det −22.5 pp, QA@20 −6.3 pp |
| Continuity pooled (Phase 12 hold) | 78.6% det / 37.6% QA@20 on 234 drawable — 0.0 pp |
| Promote VECTOR_GEOMETRY_COUNT default-on? | No |
| Promote pipeAcrossPages graph product default? | No (stays max) |
| CORRECTIONS_SEED / FEW_SHOT | Seed keep ON (15→165); few-shot stay opt-in |
| EVIDENCE_* flags | Untouched — still opt-in (Phase 12 closed) |
| Regression / CI floors | PASS; floors remain Phase 6 |
| QA@20 ≥42% stretch | Still unmet (37.6%) |
Headline metrics remain detection and QA@20 on all drawable ground-truth rows — the denominator Phase 5 forced after the gameable “soft among soft-eligible pairs” framing. We do not headline legacy soft.
Why this axis
Session 1 of the accuracy programme called real vector-PDF geometry with legend-seeded template matching the highest-leverage non-VLM fix: Peak drawings are AutoCAD-exported vectors, not pure scans. An early prototype failed two opposite ways — transitive over-merge of whole pages into one cluster, and atomic shredding of multi-path symbols. Phase 10’s “legend clustering” was a text-description similarity proxy over VLM items (+0.7 pp QA@20). That result said nothing about real geometry. Phase 13 built the geometry path for real and measured it on the same BA freeze of 142 drawable P&D rows used since Phase 9.
Track A — vector geometry (null)
Method: PyMuPDF path extract → bounded proximity clustering (legend-derived bbox cap) → composite shape signature match to seeds from prior VLM pins. Counting step API cost $0. Control: flag off. Primary: flag on, mode=supplement (inject geometry, keep VLM). Ablation: replace_if_confident (wipe VLM for high-confidence families).
| Arm | Detection | QA@20 | Δ det / QA@20 |
|---|---|---|---|
| Control (flag off) | 71.8% | 21.1% | — |
| Primary (supplement) | 71.8% | 21.1% | 0.0 / 0.0 pp |
| Ablation (replace) | 49.3% | 14.8% | −22.5 / −6.3 pp |
The two original structural failure modes did not reappear: no page-scale over-merge (largest cluster 65 paths, not ~20k), and multi-path assemblies formed for circle+line valves. What failed is semantic. Peak SOR rows need diameter and subtype; many CAD symbols for 15 mm vs 20 mm valves are the same block insert at different scale. Shape match cannot recover that. Engine raw valve count ~1050 vs SOR valve mass ~227. Family-level QA@20 on the five target discrete families stayed at zero for both arms. Wiping VLM diameter text made pairing worse — which is why replace_if_confident regresses.
Real geometry did not beat the Phase 10 text proxy (+0.7 pp). Zero is zero. Promote gate 1 failed; $0 cost and safe fallback do not rescue a null lift.
Track B — cross-page network length
Phase 12 left per-page merge ready but cross-page connected components false. Phase 13 ships an opt-in aggregation option `pipeAcrossPages: 'graph'` with identity re-show edges (normalized Hausdorff), junction tags, and scale-justified world endpoints (2 mm paper × 1:N). Product default stays `max`.
| Case | Before | After graph |
|---|---|---|
| N23 identity re-show (GT 5 m) | sum abs err 15 | abs err 0 (equals max) |
| Plan + riser tag (expect 28 m) | max 20 undercounts | 28 (beats max) |
| Peak AC multi-page identity | sum 102.2 m | 51.1 m (equals max) |
| N23 upvc50 (GT 12 m) | sum abs err 150 | abs err 114 (residual large) |
| Negative: different shape / far origins | — | 0 edges (no invent) |
Graph equals max on pure re-shows is honesty success — geometry-backed de-duplication — not a Schedule C discrete-count breakthrough. Residual errors remain. No Peak BA det/QA@20 lift is claimed from this track. Promote graph as product default: no.
Track C — corrections bank at real scale
Prior few-shot tests were n=1–2 and contradictory (Phase 4 positive, Phase 6 negative, Phase 12 mixed). Phase 13 mined failure rows from cached benchmarks, filtered honestly (no length_scope structural artefacts, no ai_only, no EL noise), and raised the seed corpus from 15 to 165 entries — target ≥100 met without padding.
| A/B arm (8 Peak pages) | Detection | QA@20 |
|---|---|---|
| Seed OFF | 6.4% | 1.7% |
| Seed ON | 13.7% | 2.1% |
| Δ (estimand) | +7.3 pp | +0.4 pp |
Absolute rates are partial-page extract versus whole-building SOR — do not compare them to Phase 6 full-drawing 78.2% / 38.0%. The estimand is ON minus OFF. Directionally positive detection; QA@20 lift tiny; not claimed significant. That partially resolves the old contradiction: seed helps more often than the single Phase 6 smoke suggested, and still is not a quantity silver bullet. CORRECTIONS_SEED stays default ON. Remote CORRECTIONS_FEW_SHOT stays opt-in.
Track D — first confidence intervals across projects
Every prior public number in this programme was a bare point percentage. Phase 13 pre-registered methods (frozen commit, drawable denominator, Wilson per project, project-level bootstrap for pooled means) and then scored. That methodological upgrade stands even where accuracy did not.
| Project | Detection (Wilson 95%) | QA@20 (Wilson 95%) | Notes |
|---|---|---|---|
| Peak (continuity) | 78.6% [72.9, 83.4] (184/234) | 37.6% [31.6, 44.0] (88/234) | Development reference |
| N23 (held-out) | 88.9% [56.5, 98.0] (8/9) | 77.8% [45.3, 93.7] (7/9) | n=9 — wide CI |
| CPS (held-out) | 61.5% [35.5, 82.3] (8/13) | 38.5% [17.7, 64.5] (5/13) | Parser-confounded; INFO |
| Kwai On | null (blocked) | null (blocked) | Class 84.5% [81.6, 87.0]; estate vs block scope |
| S33 | N/A quantity | N/A quantity | Class 86.9% [83.6, 89.7] |
| Pooled estimand | Point | 95% CI |
|---|---|---|
| Detection — unweighted project bootstrap (Peak+N23+CPS) | 76.4% | [61.5, 88.9] |
| QA@20 — unweighted project bootstrap | 51.3% | [37.6, 77.8] |
| QA@20 row-weighted (transparency; Peak-dominated) | 39.1% | Wilson [33.3, 45.2] |
Kwai On still has no honest quantity score: the bill prices the whole estate while drawings are per block, and no QS block-factor table exists in the training set. Inventing a BL1-only number would confound scope with vision accuracy. Classification 84.5% with a tight Wilson interval is what we can publish. Wide CIs on N23 and CPS are the point of the study — small-n point percentages alone were always overconfident.
Regression and defaults
Task FG: PASS on Phase 6 floors against the Phase 12 continuity arm (78.6% / 37.6%). No new flag promoted → floors stay at phase 6. VECTOR_GEOMETRY_COUNT remains opt-in (default off). Graph length remains an aggregation option, not product max. EVIDENCE_* untouched. CORRECTIONS_SEED already default ON with a larger honest corpus.
Phase 14 backlog
- Hybrid counting: VLM/OCR for diameter text + vector geometry for instance location — do not promote pure shape count until ≥0.5 pp QA@20 on the same row set.
- True legend-table crop / structured legend text — FA seeds still lean on VLM pins because outline legend labels barely extract.
- Graph length product path only after live polylines and residual audit (upvc50-class abs err still large).
- Powered multi-project corrections A/B before any CORRECTIONS_FEW_SHOT default-on.
- Kwai On quantity only after estate-vs-block factors exist — not by forcing a confounded BL1 number.
- Keep CIs on every cross-project claim; leave QA@20 ≥42% marked unmet rather than redefined.
- Do not reopen the EVIDENCE_* cost envelope without a new accepted budget that holds ≥70.5% detection.
Full Harvard-style report: docs/PEAK_PHASE13_ACCURACY_REPORT.md. Artefacts: .tmp/phase13/. Companion phases: promote close (Phase 12), evidence evaluation (Phase 9), measuring what we ship (Phase 5).
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read