Peak Phase 14 — What Is Knowable from These Drawings
Phase 14 is different in kind from everything before it. Thirteen phases optimised perception — better extract, tiles, recount, vector geometry, corrections at scale. Six of them in a row left pooled quantity accuracy within ±20% (QA@20) stuck at 37.6–38.0% of drawable Schedule C rows. This phase forbade itself from trying to move that number. Its success criterion is not a higher percentage. It is a defensible answer to: where does the remaining accuracy live, how much is reachable from these drawings, and what is it worth?
37.6% of drawable rows — which is 45.1% of the pessimistic answerable ceiling (37.6% of the optimistic ceiling) — and an unknown share of the bill’s value, because Peak-SOR has zero priced rows.
No accuracy improvement was attempted, and none occurred. Continuity is exactly the Phase 13 hold: 78.6% detection / 37.6% QA@20 on 234 drawable rows. The non-interference proof (Task GG) reads NON_INTERFERENCE_PROVEN. That control is not a footnote — it is what makes every other number in this phase trustworthy.
Why stop optimising
The raw verdict wall had been visible since Phase 2. On Peak P&D alone (n=273): 127 AI-only, 100 ground-truth-only, 22 length-scope, 11 under, 9 over, 4 exact match. Unpaired: 227 of 273 — 83.2%. Only 24 rows ever reach a quantity comparison (8.8%). That figure is P&D-dominant: FS unpaired 36.0%, MVAC 38.5%. EL&ELV is 100% unpaired because Peak training has no electrical PDFs — honest N/A since Phase 3, not a matcher failure. Phase 13’s vector-geometry arm confirmed the semantic side: wiping diameter text collapsed pairing 102→70 and detection −22.5 points even after the geometry engine fixed its own clustering bugs. Objects are often found. Identity to the right bill line is not.
Five margins had never been measured: the answerable ceiling; a staged causal error budget; money-weighted accuracy (unit rates were parsed in ingestion and reached zero comparison rows); Standard Method of Measurement semantics in code; and whether held-out projects would re-rank the same way. Phase 14 measures the first three and forensically classifies the unpaired wall. It does not implement SMM, and it does not re-score held-outs under new denominators.
The answerable ceiling
Every drawable Peak row on sheets with drawings (234 = P&D 142 + FS 59 + MVAC 33; EL excluded as N/A) was classified A (fully answerable from available drawings), B (partial — e.g. page scope vs building bill), or C (not answerable). Counts on the published denominator: 195 A, 39 B, 0 C.
| Ceiling assumption | QA@20 ceiling | Current 37.6% as share of ceiling |
|---|---|---|
| Optimistic (A+B) | 100% | 37.6% |
| Pessimistic / mid (A only) | 83.3% | 45.1% |
The Phase 6 stretch target — QA@20 ≥42% — sits inside both ceilings. It was never an unanswerable target. It was an unmet performance target. The residual gap is mostly pipeline: pairing, quantity fidelity, and scope — not “the drawings cannot support the bill.” QS adjudication of a 56-row sample is still pending; until those verdicts fold in, the ceiling remains rule-based with a sensitivity envelope of 76.5–100% if sample classes flip.
Value-weighted vs row-weighted — the commercial gap
For a quantity surveyor, the most useful thing this phase produced may be the negative result on money. Peak-SOR’s Unit Rate and Total columns are formula shells over empty inputs. Parsed unit rates and amounts on every drawable comparison row: null. Value-weighted detection, QA@20, and Cov@P90 are therefore unavailable — not zero usefulness, unknown. Thirteen phases of “accuracy” never scored bill value because the training money columns are empty shells. We refuse to impute rates.
| Weighting | Detection | QA@20 | Denominator |
|---|---|---|---|
| Row (continuity) | 78.6% | 37.6% | 234 drawable rows |
| Bill value | — | — | 0/234 priced (blocked) |
| Length (metres) | 95.4% | 36.4% | 11,421 GT metres (64 m-rows) |
Length is the commercial proxy when rates are absent. Detection jumps +16.8 points when weighted by metres — large pipe runs usually pair — but QA@20 stays almost flat (−1.2 pp). Reading without euphemism: the system finds the runs and still mis-measures a large share of metre mass (about 7,260 of 11,421 metres outside ±20%). Top length-at-risk rows are building-scope undercounts: 28 mm copper 740 m → AI 12 m; 9.52 mm refrigerant 635 → 80; slim duct cover 800 → 585.8. Row-weighting was not wrong as a scientific control. It was misleading as a standalone commercial KPI: a twelve-count valve and an 800 m duct run share equal weight.
Selective prediction (Cov@P90 — largest share acceptable at ≥90% precision) is 23.5% row-weighted. Value-weighted Cov@P90 is undefined without rates. Length-weighted Cov@P90 is 0% under current uncalibrated confidence scores. Review economics remain an open product metric once prices exist — not a substitute primary this phase.
The error budget — two instruments, two rankings
A staged oracle on frozen extracts injects perfect information one stage at a time — pairing rescue, perfect counts on real pairs, scope scaling, unit semantics — without editing the matcher. These are diagnostic upper bounds, not shippable accuracy. On the BA freeze of 142 drawable P&D rows (tiles control baseline 71.8% det / 21.1% QA@20):
| Stage (isolated) | Δ QA@20 pp | Δ detection pp |
|---|---|---|
| Counting (perfect qty on real pairs) | +50.7 | 0 |
| Scope (⊂ counting) | +14.1 | 0 |
| Pairing (keep AI quantities) | +8.5 | +26.1 |
| Unit semantics | 0 | 0 |
Do not sum those points — stages overlap, and pairing × counting is superadditive. Joint oracle of all stages reaches 97.9% QA@20 on BA-142 and 94% pooled; residual perception miss under this decomposition is only 2.1–6 points. The largest error term is reducible in principle, not an irreducible floor.
Independently, correspondence forensics classified every unpaired Peak row. Excluding EL’s 266 no-drawing lines, the top classes are heading-context lost (79 rows, 28.8%), description vocabulary (54, 19.7%), scope coverage (44, 16.1%), AI noise (41, 15.0%), and subtype/diameter (31, 11.3%). Correspondence classes together are 66.8% of non-EL unpaired rows. Bidirectional forensic pairing on P&D finds greedy one-to-one overlap of only 33% of ground-truth-only lines (any-peer 52%) — many bare “15 mm dia.” bill lines share one unsized AI “Gate valve.”
GC’s largest QA@20 lever (counting on already-paired rows) is invisible to GE. GE’s largest non-EL class (heading context lost) is folded into GC’s pairing arm. Both can be true. Averaging them into one ranking would be a methodological error.
Binding constraint is metric-dependent: quantity accuracy on current pairs → counting fidelity; detection and entry into the quantity test → correspondence. Unit semantics are agreed negligible on Peak under current extracts (0 pp oracle; 7 forensic rows).
Integrity: the number that did not move
Every measurement instrument sat next to the scoring path — optional money fields on comparison rows, optional weights for selective prediction, new offline scripts. Task GG’s job was to prove none of that perturbed production continuity. Required and measured: pooled 78.6% / 37.6%, sheet holds exact, CI floors still phase 6, EVIDENCE_* and VECTOR_GEOMETRY_COUNT still opt-in promote-no. Offline re-score of frozen extracts after the domain patches still hits the tiles control exactly (71.8% / 21.1% on BA-142). Without that control, a “ceiling” could be an artefact of a silent scorer change. With it, Phase 14 describes the same system Phases 6–13 published.
No flag was promoted. No metric moved. That was the intent. A measurement phase that improves the score is not a measurement phase.
What we were doing wrong
For six phases after Phase 6, flat QA@20 was treated as a perception problem. Perception instruments moved detection (oneshot ~33.8% → ~72% P&D under tiles) and then quantity plateaued. The programme never budgeted quantity error on already-paired rows, never measured how much of the bill is answerable from the drawings, and never weighted a metre of copper against a valve count. Optimising only perception or only ontology each leaves the other mass untouched. Optimising “the number” without an error budget is how six flat phases happen.
Phase 14 does not recommend abandoning QA@20. The largest residual terms are reducible under the oracle decomposition. It recommends re-targeting which failure mode to attack, and adding commercial parallel metrics once rates exist.
Phase 15 target
Primary target, chosen by recoverable metre-mass and diagnostic QA@20 points because bill value is null: linear quantity fidelity — length measurement, building/page scope scaling, and diameter-aware aggregation on unit=m work already paired. Measured size: counting oracle +50.7 pp QA@20 isolated on BA-142 (diagnostic); length-weighted QA@20 36.4% at 95.4% detection on 11,421 m. Metre mass is found and mismeasured.
Correspondence repair (heading context, vocabulary, diameter) is the explicit secondary track — mandatory if length work plateaus while P&D discrete QA@20 stays near 20%. Falsification is pre-registered: if a length/scope intervention fails to raise length-weighted QA@20 by ≥5.0 pp on the same 64 m-rows while holding detection ≥90%, the primary target is falsified for metre economics; residual mass on unpaired valves then pivots the programme to ontology.
Full methods, tables, limitations, and artefact paths: the Phase 14 report. Public research index: Peak study card on /projects.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read