Blog
EngineeringAI AccuracyResearchPeakQuantityVerification

Peak Phase 2 Quantity Accuracy Audit: Soft ±20% Across EL&ELV, P&D, FS, and MVAC

Teraquant Team16 min read

Phase 1 of the Peak full-SOR audit answered a classification question: can every qty>0 Schedule C line get a takeoff family? The answer was yes for 90.2% of 512 rows (EL&ELV 85.7%, P&D 97.3%, FS 92.2%, MVAC 91.7%). That is necessary — without a family, quantity verification has nothing to attach to — but it is not sufficient. Phase 2 asks the harder product question: when the model looks at Peak drawings, do the extracted quantities match Peak-SOR rows within an honest tolerance? This post reports the live multi-sheet quantity audit (Tasks L–O, merged as Task P). Numbers come from verified artefacts — not re-run for this write-up.

SUPERSEDED FOR PRODUCT CLAIMS: After Phase 5, legacy soft among soft-eligible pairs is not the programme headline. Soft figures below (~68.6% overall, and per-sheet soft %) are historical Phase 2 scores on a filtered paired set — not fixed-denominator accuracy on every drawable SOR line. For Detection and QA@20 on all drawable ground truth, read /blog/peak-phase5-measuring-what-we-ship and /blog/peak-phase6-accuracy-sprint.
Soft accuracy is measured among paired, soft-eligible rows after length_scope (and similar structural) exclusions — not on every SOR line. We do not claim “every Schedule C line is 70% accurate.”

Metric definitions

Phase 1 and Phase 2 use different denominators on purpose. Classification asks whether a family was assigned. Detection, exact, and soft ask whether AI quantities line up with ground truth after live extract and aggregation.

MetricDefinitionDenominator / notes
Classification %Non-provisional family assigned (not other)All qty>0 SOR rows (Phase 1: 512)
Detection %Paired drawable GT (match / under / over / length_scope)Drawable GT rows (Phase 2: 463)
Exact %AI qty equals SOR qtyDrawable GT rows
Soft % (±20%)Within ±20% on soft-eligible pairsExcludes length_scope from soft denominator
length_scopePartial-page vs whole-building linear (or count) structural gapExcluded from soft %; still counted in detection
count_scopeDense count items not page-comparable to building stockSame honesty rule as length_scope where applied

Per-sheet results

Drawable item counts are Phase 2 denominators (provisional rows stripped). Classification % is carried forward from Phase 1 family assignment. Soft target was ≥70% among soft-eligible pairs; detection targets were ≥70% per sheet (MVAC ≥85%).

SheetItems (drawable)Classification %Detection %Exact %Soft % (±20%)Notes
EL&ELV22985.7%0.0%0.0%0.0%No Peak EL drawings paired — honest ceiling
P&D14297.3%23.2%9.9%72.0%Soft ≥70% among paired non–length_scope
FS5992.2%76.3%16.9%60.9%Plateau; dense sprinklers / fittings undercount
MVAC3391.7%97.0%48.5%73.9%Soft ≥70%; length_scope for whole-building pipes
Overall46390.2%23.8%8.6%~68.6%Soft = soft passes / soft-eligible pairs

Scorecard vs target: P&D and MVAC pass soft ≥70%; FS passes detection only (76.3%) but soft 60.9% misses the bar; EL&ELV fails both because pairing is zero. Overall soft ~68.6% sits just under the 70% Phase 2 bar — driven by EL total miss and FS plateau, not by MVAC or P&D pair quality alone.

Methodology

The Phase 2 pipeline reuses the Peak full-SOR harness with quantity-focused outputs. Live vision extract runs on Peak training drawings for the discipline under test; AI claims are aggregated, then compared to Peak-SOR ground truth.

  • Live extract: Gemini / xAI vision on Peak drawing pages (discipline-routed: P+D → P&D, FS → FS, AC → MVAC).
  • Aggregation: aggregateAiItems() — equipment max-dedup, pipe length sum, count-item sum, legend vs plan heuristics, P+D multi-drawing rules.
  • Comparison: compareAiToGroundTruth() — match / ai_under / ai_over / length_scope / gt_only / ai_only; soft ±20% on soft-eligible pairs.
  • Artefacts: per-sheet JSON/CSV under .tmp/peak-phase2/; combined-results.json + overall-summary.json; Task P report docs/PEAK_PHASE2_QTY_REPORT.md.
  • Gates: --min-soft=N exit code when any sheet falls below threshold; Phase 2 did not invent hard-coded SOR answers into classifiers.

Per-discipline findings

EL&ELV — critical pairing failure

Soft, exact, and detection are all 0.0% on 229 drawable GT lines. That is not a near-miss story: zero GT rows paired with AI. All 266 non-provisional-ish EL lines in the comparison are miss (gt_only). Top miss families: luminaire (70), distribution_board (43), switch_isolator (27), electrical_cable (27), conduit (18). Absolute quantity gaps are dominated by whole-building cable and conduit metres (tens of thousands of metres with AI = null). AI-only rows (49) are dominated by plumbing families (water_tank, valve, sanitary_fixture, pump) — a strong signal that the EL run did not attach electrical claims to EL SOR lines (wrong drawing set, wrong scope, or broken family routing). Honest ceiling until Peak EL drawings are in the extract path: 0% quantity metrics, while classification remains 85.7%.

P&D — soft OK, detection weak

Soft 72.0% clears ≥70% on the paired soft-eligible set; detection is only 23.2% (113 misses of 142 drawable). Misses concentrate in pipe_fitting, valve, bare-diameter children, and some concrete/uPVC/copper pipe lines. Overs on gully diameters (e.g. 100 mm: SOR 8 vs AI 79) suggest over-aggregation or wrong symbol→family mapping. length_scope on CI / uPVC pipe metres is expected under partial-page geometric takeoff and is correctly excluded from soft. Page-comparable soft sits at 66.7%; building-scope soft at 33.3% — another reminder that product KPIs should prefer page-comparable soft, not whole-building linear forced into the same bucket.

FS — detection OK, soft short

Detection 76.3% meets the ≥70% detection target; soft 60.9% misses ≥70%. Seventeen length_scope rows (GI pipe, some sprinkler/count scope) are honest exclusions; remaining under/over on fittings, valves, hydrants, and instruments still hurt soft. Large AI-only other (37 of 53 ai_only) indicates noisy FS extract labels that never match SOR wording. Dense sprinkler and fitting undercounts are the plateau: more pairing alone will not push soft to 80% without cleaner subtype matching (gate ≠ check; diameter-only children) and less double-count across pages.

MVAC — best balanced sheet

Soft 73.9%, detection 97.0%, exact 48.5% — best exact rate of the four, and the only sheet that clears both soft ≥70% and a high detection bar. Misses are mostly provisional plus one controller line. Hard unders on indoor FCU kW buckets (e.g. 9.0 kW: SOR 41 vs AI 14; 5.6 kW: 16 vs 7) reflect partial AC drawing coverage vs whole-building Schedule C stock. Refrigerant and condensate metres correctly often land as length_scope (building-scope linear on only two AC pages). Relative to Task I’s earlier ~50–55% soft estimate, Phase 2 soft 73.9% after aggregation and scope honesty is the fairer operational read.

Failure taxonomy

Every non-match row in the combined results carries a failureType. Dominant mode is miss (gt_only): 402 SOR rows never paired with AI — 266 of them on EL&ELV alone. Hard quantity errors are secondary.

TypeMeaningTotal
missSOR row not detected by AI (gt_only)402
underAI qty < SOR by >20% (not soft pass)16
overAI qty > SOR by >20% (not soft pass)10
length_scopeStructural partial-page / building-scope linear (or metre undercount)35
near_missWithin ±20% but not exact (soft pass)9
matchExact quantity match40
ai_onlyAI line with no SOR pair (tracked separately)137
Sheetmissunderoverlength_scopenear_missmatchai_only
EL&ELV2660000049
P&D11324941419
FS191041741053
MVAC442911616
Total402161035940137

Top failures by absolute quantity gap

Ordered by |AI − SOR| (misses use |SOR| when AI is null). Industrial impact order: EL cable/conduit whole-building misses dominate the top of the list.

#SheetFailureFamilyDescriptionSOR qtyAI qty|Δ|
1EL&ELVmisselectrical_cable2.5mm23576035760
2EL&ELVmisselectrical_cable4mm21510015100
3EL&ELVmissconduit25mm dia.1072010720
4EL&ELVmisselectrical_cable4mm290009000
5EL&ELVmissconduit20mm dia.89408940
6EL&ELVmissconduit25mm dia.85768576
7EL&ELVmissconduit25 mm dia.81008100
8EL&ELVmissconduit20mm dia.47684768
9EL&ELVmisselectrical_cable2.5mm245004500
10EL&ELVmissconduit20 mm dia.40504050

Notable paired hard failures outside the absolute-Δ top 10 (useful once detection exists): MVAC indoor FCU 9.0 kW under (41 vs 14), FCU 5.6 kW under (16 vs 7), controller wiring under (79 vs 41); P&D gully 100 mm over (8 vs 79); P&D pipe_fitting 15 mm under (36 vs 16); FS pipe_fitting 100 mm under (12 vs 6).

What we shipped in code

  • length_scope (and related scope honesty) so partial-page linear items do not silently fail soft accuracy.
  • Matching passes that prefer family + size/material context over bare diameter tokens alone.
  • aggregateAiItems() fixes — equipment max-dedup, pipe sum, count sum, legend vs plan, P+D multi-drawing rules.
  • Benchmark plumbing: per-item JSON/CSV under .tmp/peak-phase2/, retry on live API, --min-soft exit gate.
  • Task P merge report: combined taxonomy, overall-summary.json, docs/PEAK_PHASE2_QTY_REPORT.md.

Next research directions (Phase 3)

Phase 2 proved classification ≠ quantity. Soft >80% overall needs both higher pairing and cleaner qty on pairs. Suggested workstreams:

  • EL drawings required — unblock EL&ELV extract→aggregate pairing; stop cross-discipline plumbing pollution on EL runs; count families (luminaires, DBs, switches) first for quick soft wins.
  • Valve / fitting symbol clustering — subtype matching (gate ≠ check) and diameter-only children with parent-heading context on live extract.
  • Full-building coverage metadata — report page-comparable soft as the primary product KPI; show building-scope linear separately.
  • Network-aware length summation — merge connected polylines at junctions for true multi-page metres when coverage exists.
  • Soft >80% on page-comparable items — FS soft 61% → 80%+, P&D detection 23% → 50%+ with stable pair quality, EL pairing ≥50% of drawable with soft ≥70% on pairs.

Conclusion

Peak Phase 2 quantity audit is complete. Classification stayed strong (~90%), but quantity soft accuracy among soft-eligible pairs is ~68.6% overall — just under the 70% bar. Two sheets clear soft ≥70% (P&D 72.0%, MVAC 73.9%); FS plateaus at 60.9% with good detection; EL&ELV is an honest 0% until electrical drawings pair. Dominant failure is miss (402), not near-miss (9). Soft % never meant “every SOR line is 70% accurate” — it is the share of paired, soft-eligible comparisons within ±20% after length_scope exclusions. Verification-first means publishing that distinction, the top industrial misses, and the Phase 3 path to soft >80% on page-comparable items. Companion classification story: /blog/peak-full-sor-audit. Phase 3 detection-first results (weighted detection 54.3%, pooled soft 60.0%, EL N/A, P&D soft regression): /blog/peak-phase3-accuracy-deep-dive. Full internal report: docs/PEAK_PHASE2_QTY_REPORT.md.