Peak Phase 2 Quantity Accuracy Audit: Soft ±20% Across EL&ELV, P&D, FS, and MVAC
Phase 1 of the Peak full-SOR audit answered a classification question: can every qty>0 Schedule C line get a takeoff family? The answer was yes for 90.2% of 512 rows (EL&ELV 85.7%, P&D 97.3%, FS 92.2%, MVAC 91.7%). That is necessary — without a family, quantity verification has nothing to attach to — but it is not sufficient. Phase 2 asks the harder product question: when the model looks at Peak drawings, do the extracted quantities match Peak-SOR rows within an honest tolerance? This post reports the live multi-sheet quantity audit (Tasks L–O, merged as Task P). Numbers come from verified artefacts — not re-run for this write-up.
SUPERSEDED FOR PRODUCT CLAIMS: After Phase 5, legacy soft among soft-eligible pairs is not the programme headline. Soft figures below (~68.6% overall, and per-sheet soft %) are historical Phase 2 scores on a filtered paired set — not fixed-denominator accuracy on every drawable SOR line. For Detection and QA@20 on all drawable ground truth, read /blog/peak-phase5-measuring-what-we-ship and /blog/peak-phase6-accuracy-sprint.
Soft accuracy is measured among paired, soft-eligible rows after length_scope (and similar structural) exclusions — not on every SOR line. We do not claim “every Schedule C line is 70% accurate.”
Metric definitions
Phase 1 and Phase 2 use different denominators on purpose. Classification asks whether a family was assigned. Detection, exact, and soft ask whether AI quantities line up with ground truth after live extract and aggregation.
| Metric | Definition | Denominator / notes |
|---|---|---|
| Classification % | Non-provisional family assigned (not other) | All qty>0 SOR rows (Phase 1: 512) |
| Detection % | Paired drawable GT (match / under / over / length_scope) | Drawable GT rows (Phase 2: 463) |
| Exact % | AI qty equals SOR qty | Drawable GT rows |
| Soft % (±20%) | Within ±20% on soft-eligible pairs | Excludes length_scope from soft denominator |
| length_scope | Partial-page vs whole-building linear (or count) structural gap | Excluded from soft %; still counted in detection |
| count_scope | Dense count items not page-comparable to building stock | Same honesty rule as length_scope where applied |
Per-sheet results
Drawable item counts are Phase 2 denominators (provisional rows stripped). Classification % is carried forward from Phase 1 family assignment. Soft target was ≥70% among soft-eligible pairs; detection targets were ≥70% per sheet (MVAC ≥85%).
| Sheet | Items (drawable) | Classification % | Detection % | Exact % | Soft % (±20%) | Notes |
|---|---|---|---|---|---|---|
| EL&ELV | 229 | 85.7% | 0.0% | 0.0% | 0.0% | No Peak EL drawings paired — honest ceiling |
| P&D | 142 | 97.3% | 23.2% | 9.9% | 72.0% | Soft ≥70% among paired non–length_scope |
| FS | 59 | 92.2% | 76.3% | 16.9% | 60.9% | Plateau; dense sprinklers / fittings undercount |
| MVAC | 33 | 91.7% | 97.0% | 48.5% | 73.9% | Soft ≥70%; length_scope for whole-building pipes |
| Overall | 463 | 90.2% | 23.8% | 8.6% | ~68.6% | Soft = soft passes / soft-eligible pairs |
Scorecard vs target: P&D and MVAC pass soft ≥70%; FS passes detection only (76.3%) but soft 60.9% misses the bar; EL&ELV fails both because pairing is zero. Overall soft ~68.6% sits just under the 70% Phase 2 bar — driven by EL total miss and FS plateau, not by MVAC or P&D pair quality alone.
Methodology
The Phase 2 pipeline reuses the Peak full-SOR harness with quantity-focused outputs. Live vision extract runs on Peak training drawings for the discipline under test; AI claims are aggregated, then compared to Peak-SOR ground truth.
- Live extract: Gemini / xAI vision on Peak drawing pages (discipline-routed: P+D → P&D, FS → FS, AC → MVAC).
- Aggregation: aggregateAiItems() — equipment max-dedup, pipe length sum, count-item sum, legend vs plan heuristics, P+D multi-drawing rules.
- Comparison: compareAiToGroundTruth() — match / ai_under / ai_over / length_scope / gt_only / ai_only; soft ±20% on soft-eligible pairs.
- Artefacts: per-sheet JSON/CSV under .tmp/peak-phase2/; combined-results.json + overall-summary.json; Task P report docs/PEAK_PHASE2_QTY_REPORT.md.
- Gates: --min-soft=N exit code when any sheet falls below threshold; Phase 2 did not invent hard-coded SOR answers into classifiers.
Per-discipline findings
EL&ELV — critical pairing failure
Soft, exact, and detection are all 0.0% on 229 drawable GT lines. That is not a near-miss story: zero GT rows paired with AI. All 266 non-provisional-ish EL lines in the comparison are miss (gt_only). Top miss families: luminaire (70), distribution_board (43), switch_isolator (27), electrical_cable (27), conduit (18). Absolute quantity gaps are dominated by whole-building cable and conduit metres (tens of thousands of metres with AI = null). AI-only rows (49) are dominated by plumbing families (water_tank, valve, sanitary_fixture, pump) — a strong signal that the EL run did not attach electrical claims to EL SOR lines (wrong drawing set, wrong scope, or broken family routing). Honest ceiling until Peak EL drawings are in the extract path: 0% quantity metrics, while classification remains 85.7%.
P&D — soft OK, detection weak
Soft 72.0% clears ≥70% on the paired soft-eligible set; detection is only 23.2% (113 misses of 142 drawable). Misses concentrate in pipe_fitting, valve, bare-diameter children, and some concrete/uPVC/copper pipe lines. Overs on gully diameters (e.g. 100 mm: SOR 8 vs AI 79) suggest over-aggregation or wrong symbol→family mapping. length_scope on CI / uPVC pipe metres is expected under partial-page geometric takeoff and is correctly excluded from soft. Page-comparable soft sits at 66.7%; building-scope soft at 33.3% — another reminder that product KPIs should prefer page-comparable soft, not whole-building linear forced into the same bucket.
FS — detection OK, soft short
Detection 76.3% meets the ≥70% detection target; soft 60.9% misses ≥70%. Seventeen length_scope rows (GI pipe, some sprinkler/count scope) are honest exclusions; remaining under/over on fittings, valves, hydrants, and instruments still hurt soft. Large AI-only other (37 of 53 ai_only) indicates noisy FS extract labels that never match SOR wording. Dense sprinkler and fitting undercounts are the plateau: more pairing alone will not push soft to 80% without cleaner subtype matching (gate ≠ check; diameter-only children) and less double-count across pages.
MVAC — best balanced sheet
Soft 73.9%, detection 97.0%, exact 48.5% — best exact rate of the four, and the only sheet that clears both soft ≥70% and a high detection bar. Misses are mostly provisional plus one controller line. Hard unders on indoor FCU kW buckets (e.g. 9.0 kW: SOR 41 vs AI 14; 5.6 kW: 16 vs 7) reflect partial AC drawing coverage vs whole-building Schedule C stock. Refrigerant and condensate metres correctly often land as length_scope (building-scope linear on only two AC pages). Relative to Task I’s earlier ~50–55% soft estimate, Phase 2 soft 73.9% after aggregation and scope honesty is the fairer operational read.
Failure taxonomy
Every non-match row in the combined results carries a failureType. Dominant mode is miss (gt_only): 402 SOR rows never paired with AI — 266 of them on EL&ELV alone. Hard quantity errors are secondary.
| Type | Meaning | Total |
|---|---|---|
| miss | SOR row not detected by AI (gt_only) | 402 |
| under | AI qty < SOR by >20% (not soft pass) | 16 |
| over | AI qty > SOR by >20% (not soft pass) | 10 |
| length_scope | Structural partial-page / building-scope linear (or metre undercount) | 35 |
| near_miss | Within ±20% but not exact (soft pass) | 9 |
| match | Exact quantity match | 40 |
| ai_only | AI line with no SOR pair (tracked separately) | 137 |
| Sheet | miss | under | over | length_scope | near_miss | match | ai_only |
|---|---|---|---|---|---|---|---|
| EL&ELV | 266 | 0 | 0 | 0 | 0 | 0 | 49 |
| P&D | 113 | 2 | 4 | 9 | 4 | 14 | 19 |
| FS | 19 | 10 | 4 | 17 | 4 | 10 | 53 |
| MVAC | 4 | 4 | 2 | 9 | 1 | 16 | 16 |
| Total | 402 | 16 | 10 | 35 | 9 | 40 | 137 |
Top failures by absolute quantity gap
Ordered by |AI − SOR| (misses use |SOR| when AI is null). Industrial impact order: EL cable/conduit whole-building misses dominate the top of the list.
| # | Sheet | Failure | Family | Description | SOR qty | AI qty | |Δ| |
|---|---|---|---|---|---|---|---|
| 1 | EL&ELV | miss | electrical_cable | 2.5mm2 | 35760 | — | 35760 |
| 2 | EL&ELV | miss | electrical_cable | 4mm2 | 15100 | — | 15100 |
| 3 | EL&ELV | miss | conduit | 25mm dia. | 10720 | — | 10720 |
| 4 | EL&ELV | miss | electrical_cable | 4mm2 | 9000 | — | 9000 |
| 5 | EL&ELV | miss | conduit | 20mm dia. | 8940 | — | 8940 |
| 6 | EL&ELV | miss | conduit | 25mm dia. | 8576 | — | 8576 |
| 7 | EL&ELV | miss | conduit | 25 mm dia. | 8100 | — | 8100 |
| 8 | EL&ELV | miss | conduit | 20mm dia. | 4768 | — | 4768 |
| 9 | EL&ELV | miss | electrical_cable | 2.5mm2 | 4500 | — | 4500 |
| 10 | EL&ELV | miss | conduit | 20 mm dia. | 4050 | — | 4050 |
Notable paired hard failures outside the absolute-Δ top 10 (useful once detection exists): MVAC indoor FCU 9.0 kW under (41 vs 14), FCU 5.6 kW under (16 vs 7), controller wiring under (79 vs 41); P&D gully 100 mm over (8 vs 79); P&D pipe_fitting 15 mm under (36 vs 16); FS pipe_fitting 100 mm under (12 vs 6).
What we shipped in code
- length_scope (and related scope honesty) so partial-page linear items do not silently fail soft accuracy.
- Matching passes that prefer family + size/material context over bare diameter tokens alone.
- aggregateAiItems() fixes — equipment max-dedup, pipe sum, count sum, legend vs plan, P+D multi-drawing rules.
- Benchmark plumbing: per-item JSON/CSV under .tmp/peak-phase2/, retry on live API, --min-soft exit gate.
- Task P merge report: combined taxonomy, overall-summary.json, docs/PEAK_PHASE2_QTY_REPORT.md.
Next research directions (Phase 3)
Phase 2 proved classification ≠ quantity. Soft >80% overall needs both higher pairing and cleaner qty on pairs. Suggested workstreams:
- EL drawings required — unblock EL&ELV extract→aggregate pairing; stop cross-discipline plumbing pollution on EL runs; count families (luminaires, DBs, switches) first for quick soft wins.
- Valve / fitting symbol clustering — subtype matching (gate ≠ check) and diameter-only children with parent-heading context on live extract.
- Full-building coverage metadata — report page-comparable soft as the primary product KPI; show building-scope linear separately.
- Network-aware length summation — merge connected polylines at junctions for true multi-page metres when coverage exists.
- Soft >80% on page-comparable items — FS soft 61% → 80%+, P&D detection 23% → 50%+ with stable pair quality, EL pairing ≥50% of drawable with soft ≥70% on pairs.
Conclusion
Peak Phase 2 quantity audit is complete. Classification stayed strong (~90%), but quantity soft accuracy among soft-eligible pairs is ~68.6% overall — just under the 70% bar. Two sheets clear soft ≥70% (P&D 72.0%, MVAC 73.9%); FS plateaus at 60.9% with good detection; EL&ELV is an honest 0% until electrical drawings pair. Dominant failure is miss (402), not near-miss (9). Soft % never meant “every SOR line is 70% accurate” — it is the share of paired, soft-eligible comparisons within ±20% after length_scope exclusions. Verification-first means publishing that distinction, the top industrial misses, and the Phase 3 path to soft >80% on page-comparable items. Companion classification story: /blog/peak-full-sor-audit. Phase 3 detection-first results (weighted detection 54.3%, pooled soft 60.0%, EL N/A, P&D soft regression): /blog/peak-phase3-accuracy-deep-dive. Full internal report: docs/PEAK_PHASE2_QTY_REPORT.md.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read