Blog
EngineeringAI AccuracyResearch PlanPre-registrationPeakPhase 16MEP TakeoffEvidence

Peak Phase 16 — Local Evidence, Attributed Edges, and the Test for 80–90%

Teraquant Team16 min read
STATUS: PLANNED / NOT YET EXECUTED. Every percentage below is either a frozen Phase 15 baseline or a Phase 16 gate. None is a new Phase 16 accuracy result.

Phase 16 is a pre-registration of the next accuracy investigation. It states the hypothesis, treatment sequence, denominators, evidence contracts, stop rules and final claim threshold before the unfinished model calls are resumed. The purpose is simple: if the result later reaches 80–90%, readers should be able to see exactly which intervention caused it. If it fails, we should not be able to rewrite the experiment after seeing the answer.

Why this phase exists

Phase 15 made the first Peak P&D live quantity run genuinely quantity-blind. The model could see identity fields such as description, unit, material, service and diameter, but it could not see ground-truth quantity. Inference was sealed before Schedule C was joined. That integrity change removed a comforting illusion: full-page prompting did not come close to the target.

Frozen Phase 15 A0RowsDetectionQA@20QA@0WAPE
Count9225.0%2.2%0.0%91.4%
Length estimate4388.4%23.3%2.3%76.4%

Count failed because small installed symbols, local tags and diameter text disappeared in the full-page raster, while notes and legends could be mistaken for installed instances. Length failed differently: 88.4% detection showed that the model often recognised the right pipe family, but pipeLengthScan returned estimated runs without measurable polylines. The system named work that existed while assigning the wrong amount, diameter or scope.

Phase 15 then started a native 4×3 overlapping-tile count treatment. The first page completed 12/12 calls, but provider transport saturation stopped the wider run at 18/120 cached non-P2 tiles; P page 2 had another 12 unresolved tiles. At 15% coverage, the treatment was labelled incomplete and deliberately left unscored. Phase 16 begins by finishing that same arm, not by replacing it with a more convenient experiment.

The one hypothesis

Complete native localized evidence, count only positioned instances, measure only unique short edges, attribute every accepted edge to supported diameter and material, and multiply scope only when the drawing proves it. This combination can move fixed-denominator count and length QA@20 into the 80–90% range.

This is not four independent guesses bundled under one name. Each layer has a separate arm and can fail independently. Local tiles may improve perception but not correspondence. Correct polylines may still enter the wrong diameter row. Correct page metres may still undercount a building bill. The experiment must expose those boundaries instead of averaging them into one opaque score.

  • Localized perception: recover small symbols, local linework, tags and diameter text at native resolution.
  • Instance and edge ledgers: force every count and metre to reconcile with positioned evidence.
  • Attribution: decide which service, material and diameter owns each accepted edge.
  • Scope: convert page evidence to floor, house or building quantity only when explicit drawing evidence supports that conversion.

B0–B8: the causal arm ladder

Every arm uses the same quantity-blind boundary, frozen recipe set and denominator. Ground-truth quantity joins only after the inference artefact has been sealed. Changing the denominator, recipe set or answer-join timing invalidates the run and requires the ladder to restart.

ArmTreatmentQuestion
B0Frozen Phase 15 A0Can the old baseline reproduce exactly?
B1Complete localized 4×3 tilesDoes native local evidence improve detection and count QA@20?
B2B1 + instance ledger and seam dedupeHow much error comes from duplicate or unlocated totals?
B3Localized short polylinesDoes measured geometry improve raw page metres?
B4B3 + junction/seam edge ledgerWhat is the effect of graph merge and duplicate prevention?
B5B4 + diameter/material attributionDo correct edges reach the correct SOR identity?
B6B5 + evidenced scopeDoes supported floor/building scope recover final metres?
B7Joint development systemDoes the combined route earn access to the blind set?
B8Frozen five-project blind runDid the platform actually reach the target?

Stage 1 — finish the tiles before scoring them

The complete campaign contains 132 tiles: 120 from ten non-P2 pages and 12 from the repeatedly failing P page 2. Eighteen usable responses already exist and remain immutable. Phase 16 resumes only cache misses. Every request records the prompt, model, image hash, tile bounds, attempts and terminal state. A provider failure remains visible and never removes the underlying page or SOR rows from the denominator.

Tile ledgerStarting stateScorable contract
Cached usable18Preserved unchanged
Missing non-P2102All retried under the frozen input
P page 20/12No permanent page exclusion
Total18/132 usable132/132 accounted; ≥95% usable; every page ≥10/12

Accounted means a success or an explicit terminal failure. Usable means evidence can actually be constructed. If usable coverage is below 95% or any page has fewer than ten successful tiles, B1 receives no comparative accuracy decision. Its status remains INCOMPLETE_PROVIDER_BLOCKED. Partial coverage is operational progress, not a scientific treatment.

Stage 2 — count instances, not totals

Each count recipe becomes a targeted local search keyed by family, diameter group, material and service context. The model returns positioned instances rather than a project total. Every accepted record carries a page, source tile, display-normalized point or box, evidence kind, identity candidates and fingerprint. Without a point or box, it cannot enter the count ledger.

  • Legend symbols, detail examples, notes and abbreviation tables are not installed instances.
  • Overlap evidence is deduplicated by geometry, family, diameter and fingerprint.
  • One evidence object may belong to only one SOR row.
  • Schedule rows and plan symbols are not added unless the drawing proves they cover mutually exclusive scopes.
  • The displayed count must equal the number of accepted unique pins.

Stage 3 — build a local polyline edge ledger

The length route no longer asks for total metres. Each tile returns short centerline segments with normalized points, endpoints, tangents, a scale region, provisional service and attribution evidence. The server transforms them to page-global coordinates and constructs a graph. Compatible endpoints may join; tee and cross junctions become graph nodes; reducers, valves and equipment connections stop unsafe diameter propagation.

Overlap is a measurement hazard. Two tiles may describe the same pipe with slightly different sampled curves. The ledger therefore combines curve distance, tangent compatibility, service and diameter evidence into a seam fingerprint. A final edge keeps every source edge ID, while accepted unique length is counted once. Legend lines, leaders, dimensions and detail-only samples are excluded from the installed network.

For every metre shown in the result, the arithmetic must be reproducible as the sum of accepted unique edge lengths under a verified scale.

Stage 4 — attribute diameter without guessing

Geometry alone cannot decide which SOR row owns an edge. A 100 mm drainage run and a 50 mm branch may be connected but must not be aggregated into one identity. Phase 16 ranks attribution evidence and preserves unknown where support is insufficient.

PriorityAttribution evidenceConstraint
1Inline Ø / DIA / mm textMust be spatially associated with the edge
2Branch label with leaderLeader endpoint must meet the edge tolerance
3Riser scheduleRiser or tag identity must link to the graph node
4Legend linkLine style and abbreviation must both match
5Bounded inheritanceOnly within a continuous component without reducer or conflict
6UnknownPreferred over a confident unsupported guess

The attribution test reports a confusion matrix for each diameter group, not only a micro average dominated by common sizes. The development gate is at least 90% diameter macro accuracy, at least 90% material macro accuracy and no more than 3% false confident attribution. Lower coverage is acceptable; false certainty is not.

Stage 5 — prove scope before multiplying it

A correct page measurement can still be far below a building bill when one typical plan represents several floors or houses. That does not give the system permission to infer the missing multiplier from the Schedule C difference. Every factor greater than one must be supported by positioned drawing evidence and must expose its arithmetic to the reviewer. Without such evidence, the multiplier is exactly one.

EvidenceInterpretation / decision
“Typical floor plan, 3/F–12/F”One plan explicitly represents ten floors
Floor schedule listing 2/F, 3/F and 5/FFactor three for the linked service only
“House 1–4 typical plan”Factor four when the plan is explicitly typical
AI 100 m vs SOR 1,000 mNo scope evidence
“Typical detail”A repeated detail, not a typical floor
Project name suggests a towerNo storey factor may be guessed

Peak is a particularly important guardrail because its low-rise, multi-house and roof-plan-heavy drawing set may not provide a simple typical-floor multiplier. If the evidence says factor one, factor one is the correct result even when it does not improve the score. The scope treatment is being tested, not guaranteed a favourable opportunity.

The result must remain visible on the drawing

Evidence is useful only if a user can inspect it. The right-hand inspector will expose identity, family, material, diameter, measurement basis, count pins or accepted edges, raw page quantity, seam and junction decisions, scope source, multiplier arithmetic, confidence and failure reason. Selecting a ledger entry must navigate to the correct page and geometry.

All pins, boxes, label anchors, polylines and scope-evidence boxes use display_normalized_v2. The browser matrix covers rotations 0/90/180/270, zoom 0.25× to 8×, device pixel ratios 1 and 2, three pan positions, overview and selected states, and the transition before and after PDF re-render. The label and line drift gate is at most one CSS pixel at P95. Evidence from another page is never allowed on the current page.

Development stop/go gates

The five-project blind set remains sealed until the joint B7 development system passes every gate below. A strong count score cannot compensate for incomplete tiles, unsupported scope or broken position evidence.

Development criterionGate
Tile usable coverage≥95%; every page ≥10/12
Count detection≥85%
Count QA@20≥75%
Length detection≥95%
Length-weighted QA@20≥70%
Evidence position correctness≥98%
Diameter macro accuracy≥90%
Unsupported scope multipliers0
Duplicate evidence allocation0

The only arm that may claim 80–90%

B8 is a frozen five-project blind evaluation containing at least 300 count rows and 120 length rows. Peak-only development performance is never the headline. Results will be reported overall, by project, by discipline and with Wilson confidence intervals. A denominator cannot shrink because a project is difficult or a provider request failed.

Final blind metricRequiredStretch
Count detection≥95%≥97%
Count QA@20≥90%≥93%
Count QA@0≥80%≥85%
Length detection≥97%≥98%
Length row QA@20≥80%≥85%
Length-weighted QA@20≥85%≥90%
Overall drawable QA@20≥80%≥85%
Position correctness≥98%≥99%

Passing point estimates is necessary but not sufficient. Overall QA@20 must have a Wilson lower bound of at least 75%; count and length row QA@20 lower bounds must each reach 75%; no single project may fall below 65% overall QA@20; and leakage, duplicate allocation and unsupported multipliers must all remain zero. Only then is “80–90%” an earned platform claim.

Pre-registered falsification

The central length hypothesis is falsified if B3–B6 fail to improve length-weighted QA@20 by at least five percentage points over the frozen baseline, or if detection falls below 90%. If raw edge metres become accurate while SOR assignment remains wrong, the failure is correspondence — especially diameter and heading context — and must not be relabelled as perception. If scope evidence does not exist on Peak, the scope arm may correctly produce no uplift.

  • No ground-truth quantity, expected metres or comparison verdict may enter a model prompt.
  • No scope factor may be inferred from the gap between AI and SOR quantity.
  • No failed page or difficult row may be removed from the denominator.
  • No partial tile campaign may be described as the complete treatment.
  • No unpositioned count or duplicated overlap edge may receive credit.
  • No engineering rate, price or total-value optimisation is part of this phase.

What the next journal must contain

The execution journal will not be published as a victory narrative written after the fact. It must report the final 132-tile coverage ledger, every failed page, B0–B8 results, count and edge reconciliation, diameter confusion matrices, every multiplier source, overlay drift, blind confidence intervals and the promote decision. If the required gates are missed, the title and conclusion must say that Phase 16 did not reach the target.

This post does not announce a breakthrough. It defines the evidence that a future breakthrough must survive.