Blog
EngineeringAI AccuracyResearchPeakPhase 10Promote ReadinessEvidence-first

Peak Phase 10 — Promote Readiness: Earning Default-On (and Why We Did Not)

Teraquant Team11 min read
Phase 9 asked whether the evidence path works. Phase 10 asked whether we may turn it on by default. Working is not the same as ready.

We previously froze a control vs tiled-evidence ablation on Peak P&D (Phase 9): oneshot detection 33.8% / QA@20 14.1% versus tiles 71.8% / 21.1%. Pooled evidence held Phase 6 at 78.2% / 38.0%. Both experiment flags stayed opt-in — cost, offline-only FS/MVAC, and no fresh product-path re-run blocked default-on. Phase 10 (Tasks CA–CH) attacked those promote blockers and the residuals that keep P&D quantity accuracy stuck. Numbers below trace to .tmp/phase10/ — not re-invented for the write-up.

Headline: promote not earned; flags stay off

DecisionResult
CA product-path on BA freezeRan — det 72.5% / QA@20 20.4% on 142 drawable
vs oneshotDet +38.7 pp; QA@20 +6.3 pp — product wins
vs Phase 9 tilesDet +0.7 pp; QA@20 −0.7 pp — hold, not a big new win
Cost envelopeFail — ~$4.53 / ~98 min API vs ~$0.79 / ~25 min
CB live FS/MVACREGRESS — FS QA@20 −49.2 pp; MVAC −27.3 pp
Promote either flag default-on?No
Regression / CI floorsPASS on continuity; floors remain Phase 6

Headline KPIs remain detection and QA@20 on every drawable ground-truth line — the honest Phase 5 denominator. We do not revive legacy soft-among-eligible as the public score.

Product-path P&D: oneshot vs Phase 9 tiles vs CA

Task CA ran a fresh product-path tiled extract on the Phase 9 BA frozen 142 drawable P&D rows — not a silent reuse of Phase 6 tile inventory. Code path: drawing-extract runTiledEvidencePass defaults (maxEdge 2048, maxTiles 12, overlap 0.15) under FORCE_XAI_VISION. Production defaults stayed off outside the measured arm.

ArmDetectionQA@20QA@0Paired / 142
Oneshot (BA freeze)33.8%14.1%8.5%48
Phase 9 tiles (P6 inventory)71.8%21.1%14.1%102
Product path (CA fresh)72.5%20.4%12.7%103
Δ product − oneshot+38.7 pp+6.3 pp+4.2+55
Δ product − P9 tiles+0.7 pp−0.7 pp−1.4+1

Fresh product path wins the primary comparison against oneshot and holds Phase 9 tiles within −2 pp. It is not a large new accuracy win over the inventory freeze. Stretch P&D QA@20 ≥25% was not earned (20.4%). Fittings and pumps remain hard; detection still outruns usable quantity.

Cost and latency: the hard promote gate

ArmTiles / APILatencyEst. cost (USD)
Phase 9 inventory (heavy cache resume)132 planned / ~27 live~25 min~$0.79
CA product path132 planned / 108 tile API~98 min API / ~142 min wall$4.53

Phase 9’s ~$0.79 reflected heavy Phase 6 cache resume, not a fair cold product budget. Even so, a multi-dollar multi-hour cold path is not acceptable as silent default-on without a documented product-accepted budget. Beating oneshot accuracy does not authorize that spend. CF marks costEnvelope.beatsPhase9 = false — a hard promote fail independent of the det/QA@20 hold.

Live FS / MVAC tiles: REGRESS

Phase 9 only offline-held FS and MVAC. Phase 10 Task CB ran live EVIDENCE_TILED_EXTRACT on Peak FS (5 pages) and AC/MVAC (2 pages). Native renders are 7021 px on the long edge — well above the 2400 min-edge no-op skip — so product default-on would tile these sheets.

SheetMetricOffline (P9)Live CBΔ (pp)Hold (−2pp)
FSDetection81.4%79.7%−1.7yes (det only)
FSQA@2064.4%15.3%−49.2no
MVACDetection100%72.7%−27.3no
MVACQA@2063.6%36.4%−27.3no

Verdict: REGRESS. Combined cost ~$2.01 / ~55.5 min does not rescue the hold failure. Default-on tiles would move FS/MVAC quantity accuracy the wrong way. This alone blocks EVIDENCE_TILED_EXTRACT promotion.

Peak FS page 1 with integrity-safe evidence pins (prior annotate-benchmark path)
Integrity-safe annotation from scripts/annotate-benchmark.ts (no demo fallback). Offline FS still scores well; live tiles in Phase 10 collapsed QA@20 — do not confuse the two arms.

Pooled compositions: label the arms

CompositionPooled detPooled QA@20Role
P9 evidence (P&D tiles + FS/MVAC offline)78.2%38.0%Phase 9 / Phase 6 baseline
CA product P&D + offline FS/MVAC78.6%37.6%Headline continuity (−2pp hold)
CA product P&D + live FS/MVAC74.4%21.4%Default-on risk — regresses

Report 78.6% / 37.6% only with the continuity label. Do not headline the full live composition as promote-ready. Stretch pooled QA@20 ≥42% remains unearned (37.6% continuity; 21.4% live).

Residuals: legend clustering and Cov@P90

ResidualControlAfter Phase 10Δ / verdict
Legend-seeded QA@20 (P&D)21.1%21.8%+0.7 pp — stretch ≥25% not earned
Legend detection71.8%71.8%0.0 pp
P&D Cov@P900%10.6%NONZERO at ~93% precision
FS Cov@P90 (sanity)~66.1%67.8%must not collapse

Legend work is aggregation-only template matching on Phase 6 tiled AI raw (111 seeds; plan pins preferred on legend–plan conflicts). Cov@P90 uses structural confidence floors without leaking ground-truth quantities. Both inform residual narrative; neither alone promotes tiles.

Peak MVAC page 1 with coordinate-consistent evidence and measured runs
Integrity-safe Peak AC page 1 (prior live cache). Offline MVAC still holds high detection; Phase 10 live tiles did not — report arms separately.

Cross-project INFO

Project / bankResultCI floor?
Kwai OnClassification 84.5% reconfirmed; det/QA@20 null (no scorable qty path)No
CPSStill INFO (det 61.5% / QA@20 38.5% freeze); graduateToCi=false; GT delta stub onlyNo
Corrections seedsv1→v2 rebalance (15/15); live after deferred; seedHelped false (n=1 before)No

Regression and promote path

Task CG reports overall PASS against Phase 6 continuity floors on product-path P&D plus offline FS/MVAC (pooled det ≥76.2%, QA@20 ≥36.0%, P&D QA@20 ≥19.1%). Floors are not raised: promote requires CF yes plus clean CA+CB. Cost failed; live FS/MVAC regressed. Both flags stay opt-in. Targeted recount had no redesigned arm this phase — Phase 9 already showed −0.7 pp QA@20 at extra cost.

Phase 11 backlog

  • Cost cuts for product tiles — adaptive skip, fewer tiles, or a documented product-accepted default-on budget (cold path ~$4.53 / multi-hour is the main P&D promote blocker after accuracy holds).
  • Recover FS/MVAC live tile hold within −2 pp of offline floors — default-on must not collapse quantity accuracy on those sheets.
  • P&D QA@20 residual — product 20.4% / legend 21.8% still short of stretch ≥25%; fittings, valves, pumps remain hard.
  • Full Kwai On quantity path — classification 84.5% reconfirmed; det/QA@20 still null until estate-vs-block GT is designed.
  • CPS CI graduation after a frozen GT delta table, parent-heading contract tests, and fresh vision under the graduated parser.
  • Redesigned targeted recount only if it beats tiled QA@20/det at justified cost (Phase 9 BD failed).
  • Network-aware length summation and remaining S33 unclassified MEP families (older backlog).

Phase 10 closes the promote-readiness gate Phase 9 opened. Fresh product tiles still beat oneshot on Peak P&D when you pay for them; cold cost and live FS/MVAC regression mean they are not yet the right silent default. Residuals moved a little (legend +0.7 pp, Cov@P90 10.6%); they did not buy promote. The next claims will need the same discipline: freeze rows, label every arm, and refuse default-on without measurement.