Blog
EngineeringAI AccuracyResearchPeakPhase 11Promote RecoveryEvidence-first

Peak Phase 11 — Promote Recovery: Cost Plateau, FS/MVAC NOOP, Flags Still Off

Teraquant Team12 min read
Phase 10 asked whether we may turn evidence tiles on by default. Phase 11 asked whether we can recover the two gates that said no. We recovered one. Cost still says no.

Phase 10 (Tasks CA–CH) tried to earn promote for the opt-in evidence path: fresh product-path P&D tiles beat oneshot (det 72.5% / QA@20 20.4%) but cost ~$4.53 / ~98 min API, and live FS/MVAC tiles regressed quantity accuracy (FS QA@20 −49.2 pp; MVAC −27.3 pp). Flags stayed off. Phase 11 (Tasks D0 + DA–DH; DE skipped) attacks those two hard gates and stretches residual P&D quantity. Numbers below trace to .tmp/phase11/ — not re-invented for the write-up.

Headline: promote not earned; one gate recovered

DecisionResult
D0 prompt reviewPASS — no task patches required
DA cost envelopeFail — $2.25 / 64.5 min API vs ~$0.79 / ~25 min; det 54.9% (−17.6 pp)
DB FS/MVAC livePASS — designed NOOP_SKIP (discipline gate; default-on respects)
DC P&D QA@20 stretchEarned offline 26.8% (≥25%); narrative only
Continuity pooled78.6% det / 37.6% QA@20 (0.0 pp vs Phase 10)
Promote either flag default-on?No
Regression / CI floorsPASS on continuity; floors remain Phase 6

Headline KPIs remain detection and QA@20 on every drawable ground-truth line — the honest Phase 5 denominator. We do not revive legacy soft-among-eligible as the public score. We do not claim default-on.

Cost before / after (DA)

Task DA re-measured a cold product-path arm on the Phase 9 BA freeze (142 drawable P&D rows) — not a silent re-label of Phase 10 caches. Control is Phase 10 CA (maxTiles 12). Optimized arm halves the planner budget (maxTiles 6), uses jpeg q50, and keeps maxEdge 2048 / overlap 0.15. Production planner defaults stay at maxTiles 12.

ArmDetectionQA@20Cost (USD)API minutesTiles / API
Phase 9 inventory (heavy cache resume)71.8%21.1%~$0.79~25132 / ~27 live
Phase 10 product (control)72.5%20.4%$4.53~98132 / 108
Phase 11 DA optimized54.9%19.7%$2.2564.566 / 66
Δ DA − Phase 10−17.6 pp−0.7 pp−~$2.28 (~50%)cuthalved

Envelope verdict: envelopeOk=false. ~2.8× Phase 9 inventory dollars and ~2.6× API minutes; acceptedBudget=null. Plateau: offline densest-2 projection (~$0.92 / ~20 min) only approaches the minute target by discarding tile content toward oneshot detection (~33.8%). Adaptive blank skip scored zero skips on content-heavy Peak P&D. Cost is still a hard promote gate — a real 50% cut that collapses detection is not promote-ready.

FS / MVAC: offline vs live vs Phase 11 NOOP (DB)

Phase 10 Task CB ran live tiles on Peak FS and AC/MVAC and REGRESSed. Native renders are 7021 px on the long edge — well above the 2400 min-edge skip — so product default-on would have tiled those sheets. Phase 11 does not pretend a merge-only fix recovered −2 pp hold. It ships a designed discipline gate: when the tile flag is on, FS / MVAC / EL skip the evidence tile pass and keep overview-only; P&D still tiles.

SheetMetricOffline (P9)P10 liveP11 liveHold
FSDetection81.4%79.7%null (NOOP)yes (skip)
FSQA@2064.4%15.3% (−49.2)null (NOOP)yes (skip)
MVACDetection100%72.7% (−27.3)null (NOOP)yes (skip)
MVACQA@2063.6%36.4% (−27.3)null (NOOP)yes (skip)

Verdict: NOOP_SKIP with defaultOnRespects=true. Do not read this as “live FS/MVAC tile accuracy recovered.” Live tile scores are null because tiles are skipped. Continuity pooled Peak metrics keep offline FS/MVAC floors (same composition as Phase 6–10 continuity). Gate code lives in packages/domain tiling helpers and drawing-extract runTiledEvidencePass. Flags remain opt-in; the gate only shapes flag-on behaviour.

Peak FS page 1 with integrity-safe evidence pins (prior annotate-benchmark path)
Integrity-safe annotation from scripts/annotate-benchmark.ts (no demo fallback). Offline FS still scores well; Phase 10 live tiles collapsed QA@20; Phase 11 skips those tiles when the flag is on.

P&D QA@20 stretch (DC)

Task DC is offline aggregation on cached Phase 10 product AI raw — no live re-extract, BA denominator fixed at 142. Baseline reproduces product-path 72.5% det / 20.4% QA@20. Best eligible arm (multipass spatial pin de-dupe + description aliases + pairing minSimilarity 0.25) reaches 83.8% det / 26.8% QA@20.

ArmDetectionQA@20Paired / 142Role
P10 product path (baseline)72.5%20.4%103Continuity product
DC best offline83.8%26.8%119Stretch ≥25% earned offline
Δ+11.3 pp+6.3 pp+16Not product default

Stretch ≥25% is earned offline. Pooled narrative composition with offline FS/MVAC reaches 41.5% QA@20 — still short of the aspirational ≥42% pooled stretch. DC does not change tile defaults or the cost envelope, and is not alone sufficient to promote.

Pooled compositions: label the arms

CompositionPooled detPooled QA@20Role
P9 evidence (P&D tiles + FS/MVAC offline)78.2%38.0%Phase 9 / Phase 6 baseline
P10/P11 continuity (product P&D + offline FS/MVAC)78.6%37.6%Headline floors / hold
DA optimized P&D + offline FS/MVAC67.9%37.2%Failed cost experiment — not continuity
DC stretch offline + offline FS/MVAC85.5%41.5%Narrative residual only
P10 product + live FS/MVAC (history)74.4%21.4%REGRESS — DB NOOP removes as default-on path

Report 78.6% / 37.6% only with the continuity label. Do not headline DA optimized or P10 full live as promote-ready. Stretch pooled QA@20 ≥42% remains unearned on continuity (37.6%) and short on the offline stretch composition (41.5%).

Peak MVAC page 1 with coordinate-consistent evidence and measured runs
Integrity-safe Peak AC page 1 (prior live cache). Offline MVAC still holds high detection; Phase 10 live tiles did not; Phase 11 skips MVAC tiles when the evidence flag is on.

Cross-project INFO (DD)

Project / bankResultCI floor?
Kwai OnClassification 84.5% reconfirmed; det/QA@20 null (estate vs block drawings)No
CPSStill INFO (det 61.5% / QA@20 38.5% freeze); graduateToCi=falseNo
Corrections seedsv2 (15); live n=1 on Peak P p2: seed ON mixed behavioural help, not Schedule C QA@20No

Regression and promote path

Task DG reports overall PASS against Phase 6 continuity floors on product-path P&D plus offline FS/MVAC under DB NOOP (pooled det ≥76.2%, QA@20 ≥36.0%, P&D QA@20 ≥19.1%). Floors are not raised: promote requires DF yes plus clean DA cost and DB hold/NOOP. Cost failed; DB passed. Both flags stay opt-in. Targeted recount had no redesigned arm this phase — Phase 9 already showed −0.7 pp QA@20 at extra cost.

Strict promote for EVIDENCE_TILED_EXTRACT needed DA envelopeOk and DB HOLD/NOOP and continuity hold. Gate results: cost no, NOOP yes, continuity yes → recommend no. Discipline gate shipped shapes flag-on behaviour only; production env defaults unchanged.

Phase 12 backlog

  • Redesigned targeted recount that beats tiles at justified cost (Phase 9 BD failed; no redesign in Phase 11).
  • Network-aware length summation — use junction data to merge connected polylines.
  • Remaining S33 unclassified MEP families (~281 items / ~59% of S33).
  • Default-on rollout playbook — only if a future promote decision is yes (cost OK and DB HOLD/NOOP preserved).
  • Further Kwai On / CPS CI graduation if still INFO (estate-vs-block qty path; CPS frozen GT + fresh vision under graduated parser).
  • Cost envelope for cold product tiles — product-accepted budget with sign-off, or planner redesign that holds ~72.5% det without multi-dollar multi-hour spend (DA plateau documented).
  • Ship DC-class offline aggregation into product matcher only after product-path re-measure earns ≥25% P&D QA@20.

Phase 11 closes the promote-recovery loop Phase 10 opened. We made default-on safer for FS/MVAC by refusing to tile those disciplines, and we measured how much cheaper tiles can get before detection falls apart. Cheaper still is not cheap enough; safer is not the same as ready. The next claims need the same discipline: freeze rows, label every arm, and refuse default-on without measurement.