Blog
EngineeringAI AccuracyResearchPeakPhase 13Vector GeometryNull ResultConfidence Intervals

Peak Phase 13 — First-Principles Vector Geometry Did Not Move QA@20

Teraquant Team12 min read
A well-documented null on the idea we most wanted to work is a better outcome than a breakthrough that only exists in prose.

Phases 8–12 asked one question five times: can evidence-first VLM tiling (`EVIDENCE_TILED_EXTRACT` / recount) ship default-on? The answer was always no — cost envelope, live FS/MVAC regression, soft-hold recovery without dollar recovery, or recount cost not justified. Pooled detection climbed into the high seventies; QA@20 stayed flat near 37–38% and never cleared the Phase 6 stretch target of ≥42%. Phase 13 deliberately changed axis. It did not reopen those flags. Numbers below trace to `.tmp/phase13/` — not re-invented for the write-up.

Headline: primary hypothesis not confirmed

Decision / metricResult
Vector geometry primary arm (BA 142 P&D)det 71.8% / QA@20 21.1% both arms — Δ 0.0 / 0.0 pp
replace_if_confident ablationREGRESS — det −22.5 pp, QA@20 −6.3 pp
Continuity pooled (Phase 12 hold)78.6% det / 37.6% QA@20 on 234 drawable — 0.0 pp
Promote VECTOR_GEOMETRY_COUNT default-on?No
Promote pipeAcrossPages graph product default?No (stays max)
CORRECTIONS_SEED / FEW_SHOTSeed keep ON (15→165); few-shot stay opt-in
EVIDENCE_* flagsUntouched — still opt-in (Phase 12 closed)
Regression / CI floorsPASS; floors remain Phase 6
QA@20 ≥42% stretchStill unmet (37.6%)

Headline metrics remain detection and QA@20 on all drawable ground-truth rows — the denominator Phase 5 forced after the gameable “soft among soft-eligible pairs” framing. We do not headline legacy soft.

Why this axis

Session 1 of the accuracy programme called real vector-PDF geometry with legend-seeded template matching the highest-leverage non-VLM fix: Peak drawings are AutoCAD-exported vectors, not pure scans. An early prototype failed two opposite ways — transitive over-merge of whole pages into one cluster, and atomic shredding of multi-path symbols. Phase 10’s “legend clustering” was a text-description similarity proxy over VLM items (+0.7 pp QA@20). That result said nothing about real geometry. Phase 13 built the geometry path for real and measured it on the same BA freeze of 142 drawable P&D rows used since Phase 9.

Track A — vector geometry (null)

Method: PyMuPDF path extract → bounded proximity clustering (legend-derived bbox cap) → composite shape signature match to seeds from prior VLM pins. Counting step API cost $0. Control: flag off. Primary: flag on, mode=supplement (inject geometry, keep VLM). Ablation: replace_if_confident (wipe VLM for high-confidence families).

ArmDetectionQA@20Δ det / QA@20
Control (flag off)71.8%21.1%
Primary (supplement)71.8%21.1%0.0 / 0.0 pp
Ablation (replace)49.3%14.8%−22.5 / −6.3 pp

The two original structural failure modes did not reappear: no page-scale over-merge (largest cluster 65 paths, not ~20k), and multi-path assemblies formed for circle+line valves. What failed is semantic. Peak SOR rows need diameter and subtype; many CAD symbols for 15 mm vs 20 mm valves are the same block insert at different scale. Shape match cannot recover that. Engine raw valve count ~1050 vs SOR valve mass ~227. Family-level QA@20 on the five target discrete families stayed at zero for both arms. Wiping VLM diameter text made pairing worse — which is why replace_if_confident regresses.

Real geometry did not beat the Phase 10 text proxy (+0.7 pp). Zero is zero. Promote gate 1 failed; $0 cost and safe fallback do not rescue a null lift.

Track B — cross-page network length

Phase 12 left per-page merge ready but cross-page connected components false. Phase 13 ships an opt-in aggregation option `pipeAcrossPages: 'graph'` with identity re-show edges (normalized Hausdorff), junction tags, and scale-justified world endpoints (2 mm paper × 1:N). Product default stays `max`.

CaseBeforeAfter graph
N23 identity re-show (GT 5 m)sum abs err 15abs err 0 (equals max)
Plan + riser tag (expect 28 m)max 20 undercounts28 (beats max)
Peak AC multi-page identitysum 102.2 m51.1 m (equals max)
N23 upvc50 (GT 12 m)sum abs err 150abs err 114 (residual large)
Negative: different shape / far origins0 edges (no invent)

Graph equals max on pure re-shows is honesty success — geometry-backed de-duplication — not a Schedule C discrete-count breakthrough. Residual errors remain. No Peak BA det/QA@20 lift is claimed from this track. Promote graph as product default: no.

Track C — corrections bank at real scale

Prior few-shot tests were n=1–2 and contradictory (Phase 4 positive, Phase 6 negative, Phase 12 mixed). Phase 13 mined failure rows from cached benchmarks, filtered honestly (no length_scope structural artefacts, no ai_only, no EL noise), and raised the seed corpus from 15 to 165 entries — target ≥100 met without padding.

A/B arm (8 Peak pages)DetectionQA@20
Seed OFF6.4%1.7%
Seed ON13.7%2.1%
Δ (estimand)+7.3 pp+0.4 pp

Absolute rates are partial-page extract versus whole-building SOR — do not compare them to Phase 6 full-drawing 78.2% / 38.0%. The estimand is ON minus OFF. Directionally positive detection; QA@20 lift tiny; not claimed significant. That partially resolves the old contradiction: seed helps more often than the single Phase 6 smoke suggested, and still is not a quantity silver bullet. CORRECTIONS_SEED stays default ON. Remote CORRECTIONS_FEW_SHOT stays opt-in.

Track D — first confidence intervals across projects

Every prior public number in this programme was a bare point percentage. Phase 13 pre-registered methods (frozen commit, drawable denominator, Wilson per project, project-level bootstrap for pooled means) and then scored. That methodological upgrade stands even where accuracy did not.

ProjectDetection (Wilson 95%)QA@20 (Wilson 95%)Notes
Peak (continuity)78.6% [72.9, 83.4] (184/234)37.6% [31.6, 44.0] (88/234)Development reference
N23 (held-out)88.9% [56.5, 98.0] (8/9)77.8% [45.3, 93.7] (7/9)n=9 — wide CI
CPS (held-out)61.5% [35.5, 82.3] (8/13)38.5% [17.7, 64.5] (5/13)Parser-confounded; INFO
Kwai Onnull (blocked)null (blocked)Class 84.5% [81.6, 87.0]; estate vs block scope
S33N/A quantityN/A quantityClass 86.9% [83.6, 89.7]
Pooled estimandPoint95% CI
Detection — unweighted project bootstrap (Peak+N23+CPS)76.4%[61.5, 88.9]
QA@20 — unweighted project bootstrap51.3%[37.6, 77.8]
QA@20 row-weighted (transparency; Peak-dominated)39.1%Wilson [33.3, 45.2]

Kwai On still has no honest quantity score: the bill prices the whole estate while drawings are per block, and no QS block-factor table exists in the training set. Inventing a BL1-only number would confound scope with vision accuracy. Classification 84.5% with a tight Wilson interval is what we can publish. Wide CIs on N23 and CPS are the point of the study — small-n point percentages alone were always overconfident.

Regression and defaults

Task FG: PASS on Phase 6 floors against the Phase 12 continuity arm (78.6% / 37.6%). No new flag promoted → floors stay at phase 6. VECTOR_GEOMETRY_COUNT remains opt-in (default off). Graph length remains an aggregation option, not product max. EVIDENCE_* untouched. CORRECTIONS_SEED already default ON with a larger honest corpus.

Phase 14 backlog

  • Hybrid counting: VLM/OCR for diameter text + vector geometry for instance location — do not promote pure shape count until ≥0.5 pp QA@20 on the same row set.
  • True legend-table crop / structured legend text — FA seeds still lean on VLM pins because outline legend labels barely extract.
  • Graph length product path only after live polylines and residual audit (upvc50-class abs err still large).
  • Powered multi-project corrections A/B before any CORRECTIONS_FEW_SHOT default-on.
  • Kwai On quantity only after estate-vs-block factors exist — not by forcing a confounded BL1 number.
  • Keep CIs on every cross-project claim; leave QA@20 ≥42% marked unmet rather than redefined.
  • Do not reopen the EVIDENCE_* cost envelope without a new accepted budget that holds ≥70.5% detection.

Full Harvard-style report: docs/PEAK_PHASE13_ACCURACY_REPORT.md. Artefacts: .tmp/phase13/. Companion phases: promote close (Phase 12), evidence evaluation (Phase 9), measuring what we ship (Phase 5).