Blog 網誌
Notes on verification-first takeoff
Product updates, engineering notes, and practical guides — available in English and 繁體中文.
- ProductGuidePhase 283D ReconstructionEngineeringResearchMEP TakeoffS33
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min read - ProductGuidePhase 273D ReconstructionEngineeringResearchMEP TakeoffS33
Phase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read - ProductGuidePhase 263D ReconstructionEngineeringResearchMEP TakeoffS33
Phase 26 — Coverage Ladder 3/21; Count CP 2/63, Still Not 80/80
Phase 26 is the execute-close of the coverage-ladder + count-CP pack Phase 25 planned: COVERAGE_LADDER_PASS, three training_data graphs (S33 preserved, plus N23 and Kwai On), fabricatedGeometry false. Count CP is evaluated at 2/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
9 min read - ProductGuidePhase 253D ReconstructionEngineeringResearchMEP TakeoffS33
Phase 25 — Training Coverage Ladder 3/21; Count CP 0/63, Still Not 80/80
Phase 25 earned COVERAGE_LADDER_PASS: three training_data graphs (S33 preserved, plus N23 and Kwai On), fabricatedGeometry false. Count CP is evaluated at 0/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
9 min read - ProductGuidePhase 243D ReconstructionEngineeringResearchMEP Takeoff
Phase 24 — Live Pipe Graph: 21 Centerlines on Estimated G, Still Not 80/80
Phase 24 earned LIVE_PIPE_GRAPH_PASS on live 圖則.pdf page 2: geometric takeoff (xAI fallback after Gemini fetch failed) persisted 4 merged runs, reconstruct lifted 21 pipe_run polylines plus the Phase 23 eight fittings. Estimated G is not a printed 地面層. Signed-in 3D tab was not browser-certified. Not Peak P&D 3D 80/80.
1 min read - ProductGuidePhase 233D ReconstructionEngineeringResearchMEP Takeoff
Phase 23 — Estimated Floor, Scale, and Height So a Live 3D Graph Can Persist
Phase 23 earned LIVE_GRAPH_PASS on live 圖則.pdf by estimating missing parameters, not inventing pipes. Unnamed 供水平面圖 becomes estimated G. Eight world fittings/valves persisted. Zero pipe runs. Lift 8/8 is page-2 located claims only. Not Peak P&D 3D 80/80. The signed-in 3D tab was not browser-certified this wave.
3 min read - ProductGuidePhase 213D ReconstructionEngineeringResearchMEP Takeoff
Phase 21 — Live 3D Unblock Is Shape B: Height + Rebuild, Still an Empty Mesh
Phase 21 earned LIVE_DRAWING_PASS with passKind shape_b_controls on live 圖則.pdf: numeric assumed height and Rebuild shipped; extract no longer returns the banned “page not rendered” banner. The mesh is still empty — floors=0, elements=0, persisted=false; renderedCount was 0 at OI. Residual DISCIPLINE_UNRESOLVED / FLOOR_IDENTITY_AMBIGUOUS / NO_PLAN_PAGES. This is not Shape A, not Peak P&D 3D 80/80, and not a corpus 80–90% claim.
16 min read - ProductGuidePhase 193D ReconstructionAI AccuracyEngineeringResearchMEP Takeoff
Phase 19 — 3D Product Surfaces Shipped; Peak P&D 80/80 Not Earned
Phase 19 shipped the upload generate-3D checkbox (default off), a project /3d page, and the drawing 3D tab. Peak P&D 3D Detection / QA@20 ≥80% is CLAIM_NOT_EARNED: no MK artefact, percents stay null (not 0%), and the Peak 3D view is still empty (FLOOR_IDENTITY_AMBIGUOUS, 0 floors). Phase 20’s S33 offline SUCCESS_MODEL_PASS is a separate training result. No corpus 80–90% claim.
14 min read - ProductGuideEngineeringAI AccuracyResearchPhase 203D ReconstructionMEP TakeoffSchedule CS33
Phase 20 — First Training 3D Model and a Count/Length Ledger, Without an 80–90% Claim
Phase 20 earned SUCCESS_MODEL_PASS on one University of Macau S33 training PDF: sole-storey G, four printed-label fittings, fabricatedGeometry false. The SOR ledger writes countRows = 11 and lengthMetresSum = 0. That is a count ledger plus an empty measured-length section — not a pipe-length win, not Peak P&D 3D 80/80, and not a corpus 80–90% claim. The signed-in Peak 3D tab is still empty.
15 min read - ProductGuideEngineeringAI AccuracyResearchPhase 183D ReconstructionMEP TakeoffSchedule C
Phase 18 — PDF to In-App 3D to SOR, Without a 3D Accuracy Claim
Phase 18 shipped a PDF → in-app 3D → SOR path. Corpus reconstruct is 0 fully rebuilt, 9 partial, 12 blocked out of 21. Every 3D Detection and QA@20 figure is blocked (null, not 0%). Badge: INSUFFICIENT_CROSS_PROJECT_COVERAGE. No 80–90% claim. This post is the investigation journal plus a signed-in how-to for the 3D model tab — including honest empty / partial screenshots.
18 min read - EngineeringAI AccuracyResearchPeakPhase 12Promote CloseEvidence-first
Peak Phase 12 — Promote Close: Soft-Hold Det Recovered, Cost Still Blocks, Flags Stay Off
Phase 11 cleared FS/MVAC NOOP but cost plateaued (DA $2.25 / det 54.9%). Phase 12 tried to close promote: densest-8 fine grid holds det 71.1% (≥70.5%) at $3.69 / 79.7 min — still fails ~$0.79/~25 min; product DC matcher ships (P&D QA@20 26.8%); redesigned recount beats tiles +0.7pp QA@20 but cost not justified. Continuity 78.6% det / 37.6% QA@20. Promote both flags: no. Regression PASS; CI floors Phase 6; rollout playbook deferred.
13 min read - EngineeringAI AccuracyResearchPeakPhase 11Promote RecoveryEvidence-first
Peak Phase 11 — Promote Recovery: Cost Plateau, FS/MVAC NOOP, Flags Still Off
Phase 10 failed promote honestly (cost $4.53 + FS/MVAC live REGRESS). Phase 11 recovered FS/MVAC via designed discipline NOOP_SKIP and cut P&D tile cost ~50% ($4.53→$2.25) — still fails the ~$0.79/~25 min envelope; det collapsed −17.6pp on the cheap arm. Continuity holds 78.6% det / 37.6% QA@20. Offline P&D QA@20 stretch 26.8% (narrative). Promote both flags: no. Regression PASS; CI floors Phase 6.
12 min read - EngineeringAI AccuracyResearch JournalPhase 17Cross-projectCoverageMEP TakeoffQuantity Blind
Phase 17 — Cross-Project Standard: Coverage Before Accuracy Headlines
An investigation journal for the whole training corpus, not a Peak-only score. Phase 17 inventories 21 projects, evaluates classification on 13 (3,502/6,067 = 57.7%), finds only one project with full count evaluation (CPS 5/7 QA@20), and records partial length with 0/100 row QA@20. Research sufficiency remains INSUFFICIENT_CROSS_PROJECT_COVERAGE. No 80–90% claim. JI: PASS_WITH_BLOCKED_LIVE_OVERLAY.
13 min read - EngineeringAI AccuracyResearch PlanPre-registrationPeakPhase 16MEP TakeoffEvidence
Peak Phase 16 — Local Evidence, Attributed Edges, and the Test for 80–90%
A pre-registered investigation plan, not a result. Phase 16 resumes the exact incomplete localized-tile arm from Phase 15, then adds an instance count ledger, short-polyline edge ledger, evidence-scored diameter attribution and drawing-supported typical-floor scope. B0–B8 isolate each causal contribution before a five-project blind test. Only that final arm may earn an 80–90% claim.
16 min read - EngineeringAI AccuracyResearch JournalPeakPhase 15Quantity BlindMEP TakeoffEvidence
Peak Phase 15 — The Quantity-Blind Investigation That Refused a Breakthrough
A detailed investigation journal of the first quantity-blind Peak P&D live baseline. Full-page prompting failed: count detection was 25.0% and QA@20 2.2%; length detection was 88.4% but row QA@20 only 23.3%. A localized 4×3 tile arm reached 18/120 cached tiles before provider saturation and remains unscored. The durable gains were the integrity boundary, 21-project training registry, coordinate-space repair, measurement recipes, evidence ledgers, and a falsifiable route into Phase 16. No 80–90% claim. No promotion.
16 min read - EngineeringAI AccuracyResearchPeakPhase 14MeasurementError BudgetAnswerability
Peak Phase 14 — What Is Knowable from These Drawings
After six flat phases at 37.6–38.0% QA@20, Phase 14 deliberately did not try to raise the number. We measured the answerable ceiling (83.3–100%), the staged error budget (counting +50.7pp diagnostic on already-paired P&D; pairing +26.1pp detection), length-weighted accuracy (36.4% of 11,421 m within ±20%), and found bill-value accuracy is unknown because Peak-SOR has zero priced rows. Continuity held exactly 78.6% / 37.6%. No flag promoted — by intent.
15 min read - EngineeringAI AccuracyResearchPeakPhase 13Vector GeometryNull ResultConfidence Intervals
Peak Phase 13 — First-Principles Vector Geometry Did Not Move QA@20
After five straight “promote: no” decisions on VLM tiling (Phases 8–12), we tested the idea Phase 1 never properly finished: real vector-PDF geometry with legend-seeded template matching. On 142 frozen P&D rows the primary arm was flat — detection 71.8% and QA@20 21.1% with the flag on or off (Δ 0.0 / 0.0 pp). Continuity holds at 78.6% det / 37.6% QA@20. What we did earn: clustering failure modes fixed, cross-page length graph infrastructure, a 15→165 corrections corpus, and the programme’s first confidence intervals across projects. No new flag default-on.
12 min read - EngineeringAI AccuracyResearchPeakPhase 10Promote ReadinessEvidence-first
Peak Phase 10 — Promote Readiness: Earning Default-On (and Why We Did Not)
Phase 9 measured the opt-in evidence path. Phase 10 tried to earn promote: fresh product-path P&D tiles beat oneshot (det 72.5% / QA@20 20.4%) but cost ~$4.53 / ~98 min; live FS/MVAC tiles regress QA@20 (−49pp / −27pp). Legend +0.7pp; Cov@P90 0%→10.6%. Flags stay opt-in. Regression PASS; CI floors Phase 6.
11 min read - EngineeringAI AccuracyResearchPeakPhase 9Evidence-firstAblation
Peak Phase 9 — Frozen Evidence Evaluation: Measuring What Phase 8 Shipped
Phase 8 shipped the opt-in evidence path. Phase 9 ran the pre-registered frozen ablation. On Peak P&D, tiled evidence beats oneshot (det 33.8%→71.8%, QA@20 14.1%→21.1%). Pooled evidence holds Phase 6 at det 78.2% / QA@20 38.0%. Targeted recount does not help. Production defaults stay opt-in. Regression PASS; CI floors unchanged.
10 min read - EngineeringAI AccuracyResearchPhase 8Evidence-firstMEP Takeoff
Phase 8 — Evidence-First Drawing Intelligence: From One-Shot Guesses to Traceable Quantities
The next accuracy frontier is not another matching heuristic. It is an evidence pipeline: understand the sheet globally, inspect small symbols at native resolution, map every observation back into one drawing coordinate system, de-duplicate seams, recount only contested types, and expose uncertainty for review. We publish three-project baselines and the first implementation. Phase 6 live Peak tiling is complete (det 78.2% / QA@20 38.0%) — Phase 8 still claims no separate production uplift until its own ablation runs.
8 min read - EngineeringAI AccuracyResearchPeakPhase 6N23CPSTiling
Peak Phase 6 — Accuracy Sprint Across Peak, N23, and Central Police Station
Phase 5 made the denominator honest: Peak QA@20 was 33.8% on all drawable rows, not legacy soft 70.7%. Phase 6 ran live native tiles on Peak P&D and published three-project numbers. Pooled detection 55.1%→78.2% and QA@20 33.8%→38.0%. P&D detection +38pp (33.8%→71.8%). Stretch QA@20 targets missed. N23 offline A4 QA@20 77.8%. CPS first live baseline det 61.5% / QA@20 38.5% (INFO). Regression PASS.
11 min read - EngineeringAI AccuracyResearchPeakPhase 7Ground TruthIngestion
Peak Phase 7 — The Bill of Quantities We Never Looked At
Six phases measured the drawing side. Phase 7 measured the other input to the same comparison — the bill of quantities — and found Peak's own P&D ground truth was corrupt: 18 rows shared an identical key and were unpairable by any extractor, however good. Fixing it raised corpus-wide classification 19.8%→57.5% and ingestion coverage 8/21→13/21 projects. Then the honest part: repairing Peak's ground truth did NOT improve Peak's score. Ran fully in parallel with Phase 6, touching zero files it owned.
18 min read - EngineeringAI AccuracyResearchPeakPhase 5Kwai OnHeld-out
Peak Phase 5 — Measuring What We Actually Ship
We asked what the denominator of our own headline metric was, and found it was 82 rows out of 234. Honest end-to-end quantity accuracy was 24.8%, not 70.7%. This post reports the re-measurement (which changed the score by +0.0pp), the three aggregation fixes it exposed (24.8% → 33.8%), a held-out project where fixing BOQ parsing was worth +51.6 points, and two silent-fallback bugs — one of which had been publishing fabricated images to this blog for two phases.
18 min read - EngineeringAI AccuracyResearchPeakPhase 4N23Cross-Project
Peak Phase 4 — Accuracy Recovery & Cross-Project Validation
Phase 3 left P&D soft at 17.4% after detection-first pairing. Phase 4 recovers P&D soft to 53.3% (8/15 eligible), holds FS 70.7% / MVAC 80.8%, raises pooled soft 60.0%→70.7%, establishes N23 Chinese P&D baseline (det 88.9% / soft 50% / exact 44.4%), and turns Z7 regression green with CI floors.
15 min read - EngineeringAI AccuracyResearchPeakPhase 3First Principles
Peak Phase 3 Accuracy Deep Dive: Detection-First After the 402-Miss Problem
Phase 1 classification hit 90.2%. Phase 2 soft among soft-eligible pairs was ~68.6% with detection only 23.8% — misses dominated. Phase 3 attacked detection first: weighted detection 54.3% (goal >50% met); pooled soft 60.0% (goal >80% missed). EL is honest N/A; FS soft 70.7% and MVAC soft 80.8% / det 100% improve; P&D soft collapsed 72.0%→17.4% while detection rose 23.2%→32.4%.
12 min read - EngineeringAI AccuracyResearchPeakQuantityVerification
Peak Phase 2 Quantity Accuracy Audit: Soft ±20% Across EL&ELV, P&D, FS, and MVAC
Phase 1 got classification to 90.2%. Phase 2 asks whether AI quantities match Peak-SOR. Soft accuracy among soft-eligible pairs is ~68.6% overall: P&D 72.0% and MVAC 73.9% clear ≥70%; FS plateaus at 60.9%; EL&ELV is 0% — no paired electrical drawings. Detection is only 23.8% because misses dominate (402 gt_only rows).
16 min read - EngineeringAI AccuracyResearchPeakSORMEP
Peak Full-SOR Accuracy Audit: 512 Items × 18 Pages — Classification from 41% to 90%
We audited every qty>0 row in Peak-SOR.xlsx across EL&ELV, P&D, FS, and MVAC against 18 drawing pages. Six agent tasks (F–K) raised overall classification from 41.4% to 90.2%, shipped cross-page aggregation, and achieved 97% live detection on MVAC. Annotated page gallery included.
18 min read - EngineeringAI AccuracyResearchClassificationMEPMacau
Accuracy Research Phase 2: Five Studies That Pushed MEP Classification from 22% to 31% — Distribution Boards, Architectural Finishes, Junction Detection, and Chinese Drawing Verification
We implemented five follow-up studies from our S33 cross-market audit: distribution board classification, architectural finishes as a standalone domain, branching pipe detection, Chinese drawing OCR verification, and a live API benchmark framework. Together they added 11 new takeoff families and improved S33 MEP classification from 22.3% to 31.4%.
12 min read - EngineeringAI AccuracyResearchCross-MarketMacauBS Standards
Cross-Market Verification: When AI Takeoff Meets Macau S33 — A 701-Item Ground Truth Audit
We took our AI takeoff system — built on Hong Kong training data — and tested it against a 701-item Macau University BQ. The results revealed critical blind spots in Chinese-language processing, BS standard classification, and architectural finishes detection that no amount of same-market testing would have uncovered.
11 min read - EngineeringAI AccuracyResearchLength Takeoff
Geometric Length Takeoff: Drawing the Line Between AI Measurement and Ground Truth
We fixed 7 silent bugs in the pipe-length measurement pipeline and validated the results across two real projects — Hong Kong and Macau — uncovering a unit-normalization gap that made the system completely blind to Chinese-language bills of quantities.
10 min read - EngineeringAI AccuracyResearch
Why a Single Accuracy Number Is the Wrong Goal
We went looking for "the fix" that would make AI takeoff extremely accurate. What we found instead — through real testing, not a demo — is more useful: exactly where AI is reliable, where it isn't, and why.
8 min read - ProductGuideQS Workflow
How to Use Teraquant: A Field Guide for QS Teams
A practical walkthrough of the Teraquant workflow — from uploading drawings and a SOR, through AI-assisted takeoff, to a signed-off, exportable quote.
6 min read - ProductEngineeringAI Accuracy
Building Verification-First AI Takeoff: A Development Update
Where Teraquant stands today, and what we learned from testing our AI takeoff pipeline against real drawings and real tender answer keys — not synthetic benchmarks.
6 min read