Blog
EngineeringAI AccuracyResearchPeakPhase 12Promote CloseEvidence-first

Peak Phase 12 — Promote Close: Soft-Hold Det Recovered, Cost Still Blocks, Flags Stay Off

Teraquant Team13 min read
Phase 11 made default-on safer. Phase 12 asked whether safer plus a smarter planner is ready. Soft-hold detection came back. The dollar envelope did not.

Phase 11 (Tasks D0 + DA–DH) recovered FS/MVAC default-on safety via a designed discipline NOOP_SKIP and cut cold P&D tile cost ~50% ($4.53→$2.25) — still failed the ~$0.79/~25 min envelope, and detection collapsed −17.6 pp on the cheap maxTiles=6 arm. Offline DC QA@20 26.8% was narrative only. Phase 12 (Tasks E0 + EA–EH; EE skipped) tries to close the remaining promote path: a planner that holds ~72.5% det, product-path DC aggregation, and a redesigned recount that beats tiles. Numbers below trace to .tmp/phase12/ — not re-invented for the write-up.

Headline: promote not closed; soft-hold recovered; EB shipped

DecisionResult
E0 prompt reviewPASS — no task patches required
EA cost envelopeFail — $3.69 / 79.7 min API vs ~$0.79 / ~25 min; acceptedBudget=null
EA soft-hold det (≥70.5%)PASS — det 71.1% (−1.4 pp vs P10 72.5%); +16.2 pp vs DA 54.9%
DB FS/MVAC NOOPPASS — still NOOP_SKIP; defaultOnRespects=true (EF re-verified)
EB DC product-path shipYes — P&D QA@20 26.8% ≥25%; matcher shipped (not EVIDENCE_*)
EC redesigned recountBeats tiles +0.7 pp QA@20; cost not justified → promote recount no
Continuity pooled78.6% det / 37.6% QA@20 (0.0 pp vs Phase 11)
Promote either flag default-on?No
Regression / CI floors / rollout playbookPASS on continuity; floors Phase 6; playbook deferred

Headline KPIs remain detection and QA@20 on every drawable ground-truth line — the honest Phase 5 denominator. We do not revive legacy soft-among-eligible as the public score. We do not claim default-on.

Cost before / after (EA densest budget)

Task EA redesigns the cold planner toward densest-budget on a fine grid — keep maxTiles=12 plan, then select densest 8 tiles per page — instead of Phase 11’s maxTiles=6 coarsen (the accuracy bug). Optimized arm is an offline oracle densest-of-P10-fine-grid projection (not a live cold re-OCR). Production planner defaults stay maxTiles=12 with no densest budget default.

ArmDetectionQA@20Cost (USD)API minutesTiles / API
Phase 9 inventory (heavy cache resume)71.8%21.1%~$0.79~25132 / ~27 live
Phase 10 product (control)72.5%20.4%$4.53~98132 / 108
Phase 11 DA optimized (maxTiles=6)54.9%19.7%$2.2564.566 / 66
Phase 12 EA densest-8 offline71.1%19.0%$3.6979.788 / 88
Δ EA − Phase 10−1.4 pp−1.4 pp−$0.84cut−44 / −20
Δ EA − Phase 11 DA+16.2 pp−0.7 pp+$1.44+15.2+22

Envelope verdict: envelopeOk=false. Soft-hold det PASS (71.1% ≥70.5%). ~4.7× Phase 9 inventory dollars and ~3.2× API minutes; acceptedBudget=null. Plateau: densest fine grid holds soft-hold det only at ~$3.69/~80 min; densest-2 approaches P9 $ (~$0.92/~20 min) but det falls to 60.6% (fails soft hold). maxTiles=6 remains the accuracy bug — do not promote it. Soft hold alone does not open default-on.

DC product-path: offline vs product (EB)

Phase 11 DC earned offline P&D QA@20 26.8% but marked it narrative-only until product-path re-measure. Task EB wires multipass spatial pin de-dupe, description aliases, and pairing minSimilarity 0.25 into the domain product matcher and re-scores P10 product AI raw on the BA freeze (142 drawable).

ArmDetectionQA@20Paired / 142Role
P10 product path (baseline)72.5%20.4%103Continuity product
P11 DC offline (reference)83.8%26.8%119Narrative only in P11
P12 EB product-path83.8%26.8%119Ship gate earned
Δ vs product baseline+11.3 pp+6.3 pp+16Not EVIDENCE_* promote

Ship gate ≥25% is earned on the product path. Pooled EB composition with offline FS/MVAC reaches 41.5% QA@20 — still short of aspirational ≥42%. EB ship alone is not sufficient to promote EVIDENCE_TILED_EXTRACT or EVIDENCE_TARGETED_RECOUNT. Evidence flags stay unchanged by EB.

Recount redesign vs tiled-only (EC)

Phase 9 BD failed: FIFO cap 8 + blind override cost extra ~$0.21 and moved QA@20 −0.7 pp. Task EC redesigns targets (ranked, noise-filtered), multi-ROI symbol counts, seam merge, and a strict apply gate (only 14 overrides of 264 successful recount calls).

ArmDetectionQA@20Cost (USD)Verdict
Tiled-only (P6/P9 inventory)71.8%21.1%~$0.79Baseline
P9 BD recount (history)71.8%20.4% (−0.7)extra ~$0.21Failed
P12 EC redesigned recount71.8%21.8% (+0.7)~$1.82 (extra ~$1.02)beatsTiles yes; cost no

beatsTiles=true (+0.7 pp QA@20, 0 pp det). costJustified=false (~$1.46 and multi-hour scale per QA@20 point). promoteRecountCandidate=false. Never claim a recount win without both halves. EVIDENCE_TARGETED_RECOUNT stays opt-in.

Pooled compositions: label the arms

CompositionPooled detPooled QA@20Role
P9 evidence (P&D tiles + FS/MVAC offline)78.2%38.0%Phase 9 / Phase 6 baseline
P10–P12 continuity (product P&D + offline FS/MVAC)78.6%37.6%Headline floors / hold
EA densest-8 P&D + offline FS/MVAC77.8%36.8%Cost experiment — not continuity
P11 DA optimized P&D + offline FS/MVAC67.9%37.2%History — det collapse
EB DC product matcher + offline FS/MVAC85.5%41.5%Matcher ship narrative
EC recount + offline FS/MVAC78.2%38.5%Evaluation only
P10 product + live FS/MVAC (history)74.4%21.4%REGRESS — DB NOOP removes as default-on path

Report 78.6% / 37.6% only with the continuity label. Do not headline EA densest-8, EB matcher ship, or EC recount as default-on product accuracy. Stretch pooled QA@20 ≥42% remains unearned on continuity (37.6%) and short on the EB composition (41.5%).

Peak FS page 1 with integrity-safe evidence pins (prior annotate-benchmark path)
Integrity-safe annotation from scripts/annotate-benchmark.ts (no demo fallback). Offline FS still scores well; Phase 10 live tiles collapsed QA@20; Phase 11–12 skip those tiles when the flag is on (NOOP_SKIP).

Side tracks INFO (ED)

TrackResultCI floor?
Network lengthPer-page merge ready; cross-page still plateau; AC cache Δ0 mNo
S33 familiesClassification 52.2% → 86.9% (+34.7 pp); 21 unclassified leftNo
Kwai OnClassification 84.5% hold; det/QA@20 null (estate vs block)No
CPSStill INFO (det 61.5% / QA@20 38.5% freeze); graduateToCi=falseNo
Corrections seedsv2 (15); live n=2 on Peak P p2+p3: seedHelped behavioural, not Schedule C QA@20No

Regression and promote path

Task EG reports overall PASS against Phase 6 continuity floors on product-path P&D plus offline FS/MVAC under DB NOOP (pooled det ≥76.2%, QA@20 ≥36.0%, P&D QA@20 ≥19.1%). Floors are not raised: promote requires EF yes plus clean EA cost and DB hold/NOOP (tiles) or EC promoteRecountCandidate (recount). Cost failed; soft-hold det and DB passed; recount cost failed. Both flags stay opt-in. Default-on rollout playbook is deferred — we do not invent a production flip.

Strict promote for EVIDENCE_TILED_EXTRACT needed EA envelopeOk and soft-hold det and DB HOLD/NOOP and continuity hold. Gate results: cost no, soft-hold yes, NOOP yes, continuity yes → recommend no. Strict promote for EVIDENCE_TARGETED_RECOUNT needed beatsTiles and costJustified: beats yes, cost no → recommend no. Discipline gate still shapes flag-on behaviour only; production env defaults unchanged.

Peak MVAC page 1 with coordinate-consistent evidence and measured runs
Integrity-safe Peak AC page 1 (prior live cache). Offline MVAC still holds high detection; Phase 10 live tiles did not; Phase 11–12 skip MVAC tiles when the evidence flag is on.

Phase 13 backlog

  • Cost envelope / planner hold — soft-hold det recovered (71.1%); envelope still fails. Product-signed acceptedBudget + live cold densest re-measure, or cheaper planner holding ≥70.5% det. Never maxTiles=6 without soft hold.
  • Recount cost justification — EC +0.7 pp QA@20 vs tiles but not justified (~$1.46 per point). Cheaper ROI selection or larger accuracy delta before promote.
  • Default-on rollout playbook — still deferred until a future promote decision is yes.
  • Network-aware length — per-page merge ready; cross-page connected component still plateau.
  • S33 residual families — 21 unclassified remaining after 52.2%→86.9% classification (INFO only).
  • Kwai On / CPS / corrections — qty path null; CPS not graduated; corrections n=2 behavioural only (all INFO).
  • Pooled QA@20 stretch ≥42% — still short (continuity 37.6%; EB composition 41.5%).

Phase 12 closes the promote-close loop Phase 11 opened. We recovered soft-hold detection without replaying the maxTiles=6 collapse, shipped DC-class aggregation into the product matcher on a real product-path score, and proved a recount redesign can beat tiles — then refused to promote when dollars and minutes still said no. Safer and smarter still is not the same as ready. The next claims need the same discipline: freeze rows, label every arm, and refuse default-on without measurement.