Peak Phase 12 — Promote Close: Soft-Hold Det Recovered, Cost Still Blocks, Flags Stay Off
Phase 11 made default-on safer. Phase 12 asked whether safer plus a smarter planner is ready. Soft-hold detection came back. The dollar envelope did not.
Phase 11 (Tasks D0 + DA–DH) recovered FS/MVAC default-on safety via a designed discipline NOOP_SKIP and cut cold P&D tile cost ~50% ($4.53→$2.25) — still failed the ~$0.79/~25 min envelope, and detection collapsed −17.6 pp on the cheap maxTiles=6 arm. Offline DC QA@20 26.8% was narrative only. Phase 12 (Tasks E0 + EA–EH; EE skipped) tries to close the remaining promote path: a planner that holds ~72.5% det, product-path DC aggregation, and a redesigned recount that beats tiles. Numbers below trace to .tmp/phase12/ — not re-invented for the write-up.
Headline: promote not closed; soft-hold recovered; EB shipped
| Decision | Result |
|---|---|
| E0 prompt review | PASS — no task patches required |
| EA cost envelope | Fail — $3.69 / 79.7 min API vs ~$0.79 / ~25 min; acceptedBudget=null |
| EA soft-hold det (≥70.5%) | PASS — det 71.1% (−1.4 pp vs P10 72.5%); +16.2 pp vs DA 54.9% |
| DB FS/MVAC NOOP | PASS — still NOOP_SKIP; defaultOnRespects=true (EF re-verified) |
| EB DC product-path ship | Yes — P&D QA@20 26.8% ≥25%; matcher shipped (not EVIDENCE_*) |
| EC redesigned recount | Beats tiles +0.7 pp QA@20; cost not justified → promote recount no |
| Continuity pooled | 78.6% det / 37.6% QA@20 (0.0 pp vs Phase 11) |
| Promote either flag default-on? | No |
| Regression / CI floors / rollout playbook | PASS on continuity; floors Phase 6; playbook deferred |
Headline KPIs remain detection and QA@20 on every drawable ground-truth line — the honest Phase 5 denominator. We do not revive legacy soft-among-eligible as the public score. We do not claim default-on.
Cost before / after (EA densest budget)
Task EA redesigns the cold planner toward densest-budget on a fine grid — keep maxTiles=12 plan, then select densest 8 tiles per page — instead of Phase 11’s maxTiles=6 coarsen (the accuracy bug). Optimized arm is an offline oracle densest-of-P10-fine-grid projection (not a live cold re-OCR). Production planner defaults stay maxTiles=12 with no densest budget default.
| Arm | Detection | QA@20 | Cost (USD) | API minutes | Tiles / API |
|---|---|---|---|---|---|
| Phase 9 inventory (heavy cache resume) | 71.8% | 21.1% | ~$0.79 | ~25 | 132 / ~27 live |
| Phase 10 product (control) | 72.5% | 20.4% | $4.53 | ~98 | 132 / 108 |
| Phase 11 DA optimized (maxTiles=6) | 54.9% | 19.7% | $2.25 | 64.5 | 66 / 66 |
| Phase 12 EA densest-8 offline | 71.1% | 19.0% | $3.69 | 79.7 | 88 / 88 |
| Δ EA − Phase 10 | −1.4 pp | −1.4 pp | −$0.84 | cut | −44 / −20 |
| Δ EA − Phase 11 DA | +16.2 pp | −0.7 pp | +$1.44 | +15.2 | +22 |
Envelope verdict: envelopeOk=false. Soft-hold det PASS (71.1% ≥70.5%). ~4.7× Phase 9 inventory dollars and ~3.2× API minutes; acceptedBudget=null. Plateau: densest fine grid holds soft-hold det only at ~$3.69/~80 min; densest-2 approaches P9 $ (~$0.92/~20 min) but det falls to 60.6% (fails soft hold). maxTiles=6 remains the accuracy bug — do not promote it. Soft hold alone does not open default-on.
DC product-path: offline vs product (EB)
Phase 11 DC earned offline P&D QA@20 26.8% but marked it narrative-only until product-path re-measure. Task EB wires multipass spatial pin de-dupe, description aliases, and pairing minSimilarity 0.25 into the domain product matcher and re-scores P10 product AI raw on the BA freeze (142 drawable).
| Arm | Detection | QA@20 | Paired / 142 | Role |
|---|---|---|---|---|
| P10 product path (baseline) | 72.5% | 20.4% | 103 | Continuity product |
| P11 DC offline (reference) | 83.8% | 26.8% | 119 | Narrative only in P11 |
| P12 EB product-path | 83.8% | 26.8% | 119 | Ship gate earned |
| Δ vs product baseline | +11.3 pp | +6.3 pp | +16 | Not EVIDENCE_* promote |
Ship gate ≥25% is earned on the product path. Pooled EB composition with offline FS/MVAC reaches 41.5% QA@20 — still short of aspirational ≥42%. EB ship alone is not sufficient to promote EVIDENCE_TILED_EXTRACT or EVIDENCE_TARGETED_RECOUNT. Evidence flags stay unchanged by EB.
Recount redesign vs tiled-only (EC)
Phase 9 BD failed: FIFO cap 8 + blind override cost extra ~$0.21 and moved QA@20 −0.7 pp. Task EC redesigns targets (ranked, noise-filtered), multi-ROI symbol counts, seam merge, and a strict apply gate (only 14 overrides of 264 successful recount calls).
| Arm | Detection | QA@20 | Cost (USD) | Verdict |
|---|---|---|---|---|
| Tiled-only (P6/P9 inventory) | 71.8% | 21.1% | ~$0.79 | Baseline |
| P9 BD recount (history) | 71.8% | 20.4% (−0.7) | extra ~$0.21 | Failed |
| P12 EC redesigned recount | 71.8% | 21.8% (+0.7) | ~$1.82 (extra ~$1.02) | beatsTiles yes; cost no |
beatsTiles=true (+0.7 pp QA@20, 0 pp det). costJustified=false (~$1.46 and multi-hour scale per QA@20 point). promoteRecountCandidate=false. Never claim a recount win without both halves. EVIDENCE_TARGETED_RECOUNT stays opt-in.
Pooled compositions: label the arms
| Composition | Pooled det | Pooled QA@20 | Role |
|---|---|---|---|
| P9 evidence (P&D tiles + FS/MVAC offline) | 78.2% | 38.0% | Phase 9 / Phase 6 baseline |
| P10–P12 continuity (product P&D + offline FS/MVAC) | 78.6% | 37.6% | Headline floors / hold |
| EA densest-8 P&D + offline FS/MVAC | 77.8% | 36.8% | Cost experiment — not continuity |
| P11 DA optimized P&D + offline FS/MVAC | 67.9% | 37.2% | History — det collapse |
| EB DC product matcher + offline FS/MVAC | 85.5% | 41.5% | Matcher ship narrative |
| EC recount + offline FS/MVAC | 78.2% | 38.5% | Evaluation only |
| P10 product + live FS/MVAC (history) | 74.4% | 21.4% | REGRESS — DB NOOP removes as default-on path |
Report 78.6% / 37.6% only with the continuity label. Do not headline EA densest-8, EB matcher ship, or EC recount as default-on product accuracy. Stretch pooled QA@20 ≥42% remains unearned on continuity (37.6%) and short on the EB composition (41.5%).

Side tracks INFO (ED)
| Track | Result | CI floor? |
|---|---|---|
| Network length | Per-page merge ready; cross-page still plateau; AC cache Δ0 m | No |
| S33 families | Classification 52.2% → 86.9% (+34.7 pp); 21 unclassified left | No |
| Kwai On | Classification 84.5% hold; det/QA@20 null (estate vs block) | No |
| CPS | Still INFO (det 61.5% / QA@20 38.5% freeze); graduateToCi=false | No |
| Corrections seeds | v2 (15); live n=2 on Peak P p2+p3: seedHelped behavioural, not Schedule C QA@20 | No |
Regression and promote path
Task EG reports overall PASS against Phase 6 continuity floors on product-path P&D plus offline FS/MVAC under DB NOOP (pooled det ≥76.2%, QA@20 ≥36.0%, P&D QA@20 ≥19.1%). Floors are not raised: promote requires EF yes plus clean EA cost and DB hold/NOOP (tiles) or EC promoteRecountCandidate (recount). Cost failed; soft-hold det and DB passed; recount cost failed. Both flags stay opt-in. Default-on rollout playbook is deferred — we do not invent a production flip.
Strict promote for EVIDENCE_TILED_EXTRACT needed EA envelopeOk and soft-hold det and DB HOLD/NOOP and continuity hold. Gate results: cost no, soft-hold yes, NOOP yes, continuity yes → recommend no. Strict promote for EVIDENCE_TARGETED_RECOUNT needed beatsTiles and costJustified: beats yes, cost no → recommend no. Discipline gate still shapes flag-on behaviour only; production env defaults unchanged.

Phase 13 backlog
- Cost envelope / planner hold — soft-hold det recovered (71.1%); envelope still fails. Product-signed acceptedBudget + live cold densest re-measure, or cheaper planner holding ≥70.5% det. Never maxTiles=6 without soft hold.
- Recount cost justification — EC +0.7 pp QA@20 vs tiles but not justified (~$1.46 per point). Cheaper ROI selection or larger accuracy delta before promote.
- Default-on rollout playbook — still deferred until a future promote decision is yes.
- Network-aware length — per-page merge ready; cross-page connected component still plateau.
- S33 residual families — 21 unclassified remaining after 52.2%→86.9% classification (INFO only).
- Kwai On / CPS / corrections — qty path null; CPS not graduated; corrections n=2 behavioural only (all INFO).
- Pooled QA@20 stretch ≥42% — still short (continuity 37.6%; EB composition 41.5%).
Phase 12 closes the promote-close loop Phase 11 opened. We recovered soft-hold detection without replaying the maxTiles=6 collapse, shipped DC-class aggregation into the product matcher on a real product-path score, and proved a recount redesign can beat tiles — then refused to promote when dollars and minutes still said no. Safer and smarter still is not the same as ready. The next claims need the same discipline: freeze rows, label every arm, and refuse default-on without measurement.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read