Blog
EngineeringAI AccuracyResearchPeakPhase 6N23CPSTiling

Peak Phase 6 — Accuracy Sprint Across Peak, N23, and Central Police Station

Teraquant Team11 min read
Detection without a usable quantity is incomplete. Quantity without detection is fiction. Phase 6 measures both on every drawable ground-truth line — and refuses to revive soft-among-eligible as the headline.

Phase 5 established that the honest Peak end-to-end quantity accuracy within ±20% (QA@20) was 33.8% on 234 drawable SOR lines, with detection stuck at 55.1%. The wall was physical: A0 sheets at ~7021×4967 px were compressed to a single 1800–2560 px JPEG before vision, so fittings and valves disappeared. Phase 6 is the detection-first live sprint: native-resolution overlapping tiles on Peak P&D, a fittings/valve prompt lift, a live corrections-seed A/B, plus two cross-project baselines (N23 offline Phase-5 re-score and Central Police Station first live). Numbers below come from .tmp/peak-phase6/, .tmp/n23-benchmark/phase5-results.json, and .tmp/cps-benchmark/results.json — not re-invented for the write-up.

Headline metrics (QA@20 on all drawable)

ProjectDrawableDetectionQA@20Provenance
Peak pooled23478.2% (183)38.0% (89)P&D live tiles + FS/MVAC offline A4
Peak P&D14271.8% (102)21.1% (30)live tiles, xAI grok-4.5
Peak FS5981.4%64.4%Phase 3 cache hold
Peak MVAC33100%63.6%Phase 3 cache hold
N23988.9% (8)77.8% (7)offline A4 re-agg of frozen extract
CPS (INFO)1361.5% (8)38.5% (5)first live + PDF BQ

Targets: pooled detection ≥65% met (78.2%); pooled QA@20 ≥42% missed (38.0%); P&D detection ≥45% met (71.8%); P&D QA@20 ≥22% missed by 0.9pp (21.1%). FS and MVAC held at 0.0pp by construction — they were not re-extracted live. EL&ELV remains honest N/A (no EL PDF in Peak training).

Phase 5 → Phase 6 Peak

SheetDetection P5 → P6QA@20 P5 → P6Δ detΔ QA@20
Pooled55.1% → 78.2%33.8% → 38.0%+23.1 pp+4.3 pp
P&D33.8% → 71.8%14.1% → 21.1%+38.0 pp+7.0 pp
FS81.4% → 81.4%64.4% → 64.4%0.00.0
MVAC100% → 100%63.6% → 63.6%0.00.0
All measured pooled gain is P&D live tiling. Do not read Phase 6 as a full four-sheet live re-extract.

Detection rose far more than QA@20 because pairing a symbol does not put the quantity inside ±20%. P&D AI-only rows ballooned to 692 under tiles — more detections also mean more unmatched noise and higher review load. Cov@P90 on P&D remains 0%: there is still no confidence threshold that yields a ≥90%-precision auto-accept slice of the plumbing bill. Legacy soft on P&D fell 60.6%→42.6% while detection rose — more soft-eligible pairs enter the denominator. That is why soft is not the headline.

What AA–AC changed

Task AA ran 11 Peak P&D pages as 3×4 overlapping tiles (maxEdge 1800, overlap 0.15, 127/132 tile extracts after five hard socket failures). Control (one-shot Phase 3 cache) vs tiles on the same A4 aggregation:

ArmDetectionQA@20Paired / 142
Control (one-shot)33.8%14.1%48
Tiles (live)71.8%21.1%102
Δ+38.0 pp+7.0 pp+54

Family lift: valves 2→21 matched (of 23), gullies 8→13 (all), pipe fittings 4→14 (of 36 — 22 still missed), manholes and water tanks to full match. Pumps and provisional lines remain zero. Task AB (PROMPT_VERSION 1.1.8) rewrote fittings/valve COUNT guidance and added a generic-elbow noise filter; offline FS/MVAC held; live smoke showed a small plan-page row lift without elbow spam. Task AC measured the Z4 corrections seed ON vs OFF on Peak P page 1: seed ON did not help (runs 4→2, total m 88→50). Wiring works; efficacy on that short block plan was null/negative — overcount-oriented few-shots made the model more conservative.

Peak FS page 1 with coordinate-consistent evidence pins from the integrity-safe annotator
Integrity-safe annotation (scripts/annotate-benchmark.ts path — no demo fallback). Purple markers are coordinate-consistent model evidence. Aggregate SOR verdicts are not painted as per-pin match/fail. FS metrics held at Phase 5 A4 levels under the Phase 6 merge (offline Phase 3 extract).

N23 Phase-5 re-score

Task AD re-aggregated the frozen seven-page Chinese P&D extract under Phase 5 A4 options (pipeAcrossPages: max, legendOnlyCombine: max). No new live vision. The page-sum ×6 pipe error collapses:

MetricZ3 sum pagesA4 max pages
Detection88.9% (8/9)88.9% (8/9)
QA@2044.4% (4/9)77.8% (7/9)
QA@044.4% (4/9)66.7% (6/9)

Copper Ø15: 30 m → 5 m (match SOR 5). uPVC Ø50: 72 m → 12 m (match SOR 12). Residual fails are extract noise (gully 1→16; bottle trap miss), not aggregation. Gate: INFO only — not pooled into Peak CI floors.

N23 drawing page 1 annotated detection pins
N23 page 1 from the Phase 4 annotated set (integrity-safe path). Phase 6 Task AD does not re-extract vision; it re-scores the frozen extract under A4 aggregation — QA@20 77.8% on 9 drawable rows.

Central Police Station first baseline

Task AE is the third project: English Block 4 drainage, DRAWINGS.pdf pages 4–10, heading-aware BQ.pdf (not xlsx). First live baseline — INFO floors only, no CI fail.

MetricValueNotes
Drawable GT1321 BQ rows − 8 provisional Item
Detection61.5% (8/13)paired / drawable
QA@2038.5% (5/13)equals QA@0 on this run
Cov@P9015.4%selective coverage
Classification47.6%10 classified / 21 BQ rows

Stresses PDF BQ structure, three identical bare “150 mm Diameter pipe.” rows under different parent headings, and English drainage vocabulary (EN 877 CI, D.I., stainless gutter). Honest misses include manhole, sump pit, and bare-Ø other pipes; over-counts hit CI runs, rainwater outlets, and gutter length. Sample variance is high at n=13.

What still fails

  • P&D pipe fittings: 22 of 36 still gt_only after tiling — largest residual miss mass.
  • Det→QA@20 gap: 71.8% detection vs 21.1% usable quantities on P&D.
  • AI-only flood (692) raises review load; Cov@P90 stays 0% on P&D.
  • Corrections seed not yet a measured quantity win (single-page null/negative).
  • FS/MVAC not live-tiled; provider is xAI only from HK — no Gemini parity claim.

Regression and floors

Task AG regression matrix: PASS. Domain 169 tests green; web build OK; all Peak Phase 5 floors cleared with headroom. Floors raised to measured − 2pp for pooled detection (76.2%), pooled QA@20 (36.0%), and P&D QA@20 (19.1%). FS/MVAC floors held. N23 and CPS remain INFO. scripts/accuracy-floors.json records phase: 6.

What comes next

Phase 7 (BOQ ingestion integrity) already completed in parallel — corpus classification and Peak ground-truth repair — and did not change production defaults. Phase 8 ships an opt-in evidence-first path (overview → native tiles → display-coordinate evidence → seam de-dupe → targeted recount → review). Research directions after this sprint:

  • Close the P&D det→QA@20 gap with schema-constrained targeted recount and legend-seeded fitting clusters.
  • Calibrate quantity confidence so Cov@P90 on P&D is non-zero — selective prediction for QS auto-accept.
  • Rebalance the corrections bank (miss-recovery vs overcount lectures) and re-measure on dense floor plans.
  • Kwai On quantity baseline (classification already 84.5% held-out).
  • Live FS/MVAC tile runs before claiming full multi-sheet live Peak metrics; CPS parser audit before CI floors.
  • Phase 8 evaluation: freeze one-shot vs evidence-tiled extract on identical rows; report det, QA@20, cost, latency — no uplift claim without that ablation.

Companion posts: Phase 5 /blog/peak-phase5-measuring-what-we-ship; Phase 8 architecture /blog/phase8-evidence-first-drawing-intelligence; Phase 4 /blog/peak-phase4-cross-project-validation. Full internal report: docs/PEAK_PHASE6_ACCURACY_REPORT.md.