Phase 17 — Cross-Project Standard: Coverage Before Accuracy Headlines
STATUS: Platform complete; cross-project accuracy claim not earned. Sufficiency code: INSUFFICIENT_CROSS_PROJECT_COVERAGE. JI verification: PASS_WITH_BLOCKED_LIVE_OVERLAY. Every percentage below carries n/N or project coverage. No 80–90% claim.
Phases 9 through 16 taught us a great deal about The Peak: tiled evidence, cost envelopes, vector geometry nulls, quantity-blind failure modes, and the difference between a soft score and QA@20 on all drawable rows. They did not produce a comparable score across the training corpus. Phase 17 changes the research unit. The object of study is every training project’s readiness record and standard metrics—not another Peak-only arm dressed up as a product result.
Investigation question
Can we report classification, detection, QA@20, QA@0, count, length (row and metre-weighted), diameter/scope, and position correctness for every project in the training set under one fixed-denominator contract—and only then ask whether cross-project accuracy is high enough to claim?
The pre-registered sufficiency gate required at least 10 fully quantity-evaluated projects, ≥1,000 drawable rows, ≥300 count rows, ≥120 length rows, ≥2 regions, ≥3 disciplines, and no single project above 40% of the micro denominator. Until that gate passes, the corpus headline must read INSUFFICIENT_CROSS_PROJECT_COVERAGE even if individual projects look strong.
Protocol in one page
- Roster: dynamic discovery under training_data/ — 21 projects, 259 assets (221 PDF). Exact-21 hard-fail removed.
- Statuses: evaluated | partial | blocked | not_applicable. Blocked is never encoded as 0%.
- Method version: phase5_fixed_denominator_v1. Quantity-blind: GT quantity never enters model prompts before seal.
- Aggregates: micro (sum n / sum N), macro (mean of project %), Wilson CI, project-bootstrap CI (seed 20260810, 2000 iters), worst/P10, largest-project share.
- Canonical artefact: apps/web/data/training-project-metrics.v2.json (contentHash c29dae8f…). Full internal report: /docs path PHASE17_CROSS_PROJECT_ACCURACY_REPORT.
Coverage first — the executive scoreboard
| Layer | n / N | Reading |
|---|---|---|
| Projects inventoried | 21 / 21 | Every folder has a readiness record |
| SOR ingestible | 13 / 21 | 8 need OCR / structured bills |
| Classification evaluated | 13 / 21 (61.9%) | Only multi-project accuracy family with mass |
| Count / overall quantity evaluated | 1 / 21 (4.8%) | Central Police Station only |
| Length fully evaluated | 0 / 21 | 6 partial · 6 blocked · 9 N/A |
| Position evaluated | 0 / 21 | All PHASE17_POSITION_NOT_RUN |
| Sufficiency gate | FAIL | INSUFFICIENT_CROSS_PROJECT_COVERAGE |
Statistical uncertainty and data-readiness blockers are different things. Wilson intervals describe binomial noise on measured rows. Blocked projects describe missing SOR text layers, unverified scale, or evaluation arms not run. Converting either into a silent 0% would invent accuracy the system did not earn.
Classification across 13 projects
| Statistic | Value | Notes |
|---|---|---|
| Micro | 3,502 / 6,067 = 57.7% | Wilson 95% [56.5, 59.0] |
| Macro | 68.3% | Bootstrap [59.2, 76.8]; equal project weight |
| Worst / P10 | 40.1% / 43.7% | Worst: Asia expo |
| Largest project share | 45.9% | Asia expo of classification row mass |
The Peak still classifies well under this contract: 463/512 = 90.4% (Wilson [87.6, 92.7]). That is not the corpus. Asia expo contributes 1,117/2,783 = 40.1% and nearly half the micro mass. Kwai On reaches 567/678 = 83.6%; S33 493/688 = 71.7%; N23 classification-only 9/14 = 64.3%. Blind-split classification micro sits at 43.1% versus 80.9% on development_burned—another reason Peak-burned headlines overstate generalization.
Discipline slices (classification only): FS 123/138 = 89.1%; EL&ELV 292/381 = 76.6%; P&D 330/461 = 71.6%; MVAC 115/183 = 62.8%. Eight inventory-blocked projects remain needs_ocr and contribute zero accuracy rows while staying in the 21-project coverage denominator.
Count and overall quantity — one project, seven rows
| KPI | n / N | Display % | Eval projects |
|---|---|---|---|
| Detection (= count) | 6 / 7 | 85.7% | 1 — CPS |
| QA@20 (= count) | 5 / 7 | 71.4% | 1 — CPS |
| QA@0 (= count) | 5 / 7 | 71.4% | 1 — CPS |
Overall detection and QA metrics in the Phase 17 manifest mirror the discrete count family. They are not a revival of Peak Phase 12–14 continuity (78.6% det / 37.6% QA@20 on 234 drawable rows). The Peak’s Phase 17 count arm is blocked with PHASE17_COUNT_NOT_RUN—explicitly not zero. Largest-project share on count mass is 100% because only Central Police Station entered the accuracy numerator.
Negative result: the programme’s most studied project (Peak) has no Phase 17 count evaluation yet. Treating Peak legacy continuity as the corpus answer would reverse the point of Phase 17.
Length — partial, Peak-dominated, QA@20 zero
| KPI | Micro n / N | % | Status |
|---|---|---|---|
| Length row detection | 13 / 100 | 13.0% | partial aggregate |
| Length row QA@20 | 0 / 100 | 0.0% | partial aggregate |
| Length-weighted detection | 1,310 / 139,345 m | 0.9% | partial; Peak ~99.99% share |
| Length-weighted QA@20 | 0 / 139,345 m | 0.0% | partial aggregate |
Peak length is partial at 11/98 row detection and 0/98 row QA@20 (1,295/139,330 m detected, 0 m within ±20%). N23 is partial at 2/2 detection and 0/2 QA@20. Several other projects are partial with empty geometry (covered 0 of eligible rows)—honest incompleteness, not a fabricated fail rate. Scale blockers (N.T.S., unverified scale) keep other drawings out of the metre denominator entirely.
Diameter attribution on the scored subset is 13/13 correct of 224/224 attributed edges—useful but narrow. Scope multipliers greater than 1.0 were never evidenced (100/100 rows defaultOne). Position correctness is blocked on all 21 projects.
Sufficiency gate — why we refuse a corpus success headline
| Gate | Need | Observed | Pass? |
|---|---|---|---|
| Fully evaluated projects (count) | ≥ 10 | 1 | No |
| Drawable rows | ≥ 1,000 | 107 | No |
| Count rows | ≥ 300 | 7 | No |
| Length rows | ≥ 120 | 100 | No |
| Regions / disciplines | ≥ 2 / ≥ 3 | 2 / 4 | Yes |
| Largest project share | ≤ 40% | 100% | No |
Five of seven gates fail. Regions and disciplines pass, which is necessary but nowhere near sufficient. Until fully evaluated count projects and row mass rise, any single headline percentage would be a presentation choice, not a scientific claim.
Legacy figures — labelled non-comparable
Readers who followed earlier journals will recognise Peak continuity 78.6%/37.6%, Peak P&D tiles ~72.5%/20.4%, N23 A4 88.9%/77.8%, CPS Phase 6 61.5%/38.5%, and Peak classification 90.2%. Those remain historical programme results. They used different pipelines, denominators, or models. Phase 17 deliberately does not import them into corpus micro/macro. Soft-among-eligible scores (for example the old ~70.7% soft framing) stay banned as headlines.
- Peak Phase 1 class 90.2% (462/512) vs Phase 17 Peak 90.4% (463/512): near, one-row delta.
- Peak continuity quantity: non-comparable — Phase 17 Peak count blocked.
- CPS Phase 6 (13 drawable) vs Phase 17 CPS count (7 rows): related stress test, not identical protocol.
What changed in the product `/projects` surface
The app no longer pretends training folders are a side gallery. A unified project list merges live Supabase projects and training-origin rows, de-duplicates Peak as live_and_training, and shows Classification, Detection, QA@20, and Coverage with statuses. Blocked cards expose reason codes and next actions instead of 0%. Each training project has a detail route with n/N, Wilson intervals, method version, model, split, and provenance. Drawing pages separate frozen 專案基準 (project benchmark) from 本圖紙即時 (this drawing live)—live QA@0 is not relabelled as corpus QA@20, and classification is not invented on the live track.
Overlay position — verified offline, blocked live
JI regression verdict is exactly PASS_WITH_BLOCKED_LIVE_OVERLAY. Domain unit tests for display_normalized_v2 coordinates and page-scoped overlays pass (22 tests). Signed-in browser zoom/pan P95 drift on a real Peak drawing was not measured: local HTML routes require Clerk keys and a session. That is a verification blocker, not a certified pass and not a proven geometric failure. We publish the caveat rather than invent a green matrix.
Falsifications worth keeping
- “We already have cross-project quantity accuracy.” — False; 1/21 count-evaluated, 107 drawable rows.
- “Peak’s published quantity score is the training corpus.” — False; Peak count blocked in Phase 17.
- “Length detection implies length QA@20.” — False; 13% row detection with 0% row QA@20 on the partial micro cohort.
- “Blocked means zero accuracy.” — False by contract; null n/N with reasonCode.
Path toward an honest 80–90% claim
An 80–90% claim is only discussable after the sufficiency gate passes on the same contract. Priority order by cross-project error budget: (1) finish quantity-blind count evaluation on the remaining classification-ready projects—especially Peak, N23, Kwai On, S33, and Asia expo—until ≥10 projects and ≥300 count rows; (2) OCR or re-export the eight needs_ocr bills; (3) verified scales and non-empty geometry so length rows exceed 120 without Peak owning >40% of the mass; (4) adjudicated position ledgers and live overlay P95; (5) only then resume Peak-local perception arms (Phase 16 tiles, attributed edges, evidenced scope) as treatments inside a multi-project denominator. Peak Phase 14/15 diagnostics (counting oracle on paired rows; full-page quantity-blind failure) remain diagnostic, not product proof.
Promotion decision
- Promote new default-on extract flags: no.
- Claim cross-project accuracy success: no — INSUFFICIENT_CROSS_PROJECT_COVERAGE.
- Claim 80–90% overall QA@20: no.
- Ship standardized metrics, roster, and this journal: yes.
Reproducibility
Canonical metrics: apps/web/data/training-project-metrics.v2.json. Aggregation audit: docs/phase17/AGGREGATION_AUDIT.md. Verification: docs/PHASE17_VERIFICATION.md (verdict PASS_WITH_BLOCKED_LIVE_OVERLAY). Full report: docs/PHASE17_CROSS_PROJECT_ACCURACY_REPORT.md. Related journals: quantity-blind Phase 15, Phase 16 pre-registration, Phase 14 estimand, Phase 6 multi-project sprint. No new inference was run for this article; no screenshots were fabricated.
The durable result of Phase 17 is not a higher percentage. It is the inability to hide incomplete coverage behind a Peak headline.
Related articles
Phase 28 — Coverage Ladder 5/21 Held; Count CP Still 3/63, Not 80/80
Phase 28 is the execute-close of the same qty_3d coverage + S33 count-CP families: inner COVERAGE_LADDER_PASS, five training_data graphs kept (S33, N23, Kwai On, S27, N22), fabricatedGeometry false. Extra S33 plan pages 12/14/16/19 raised world elements to 41 on the same slug — not a sixth coverage credit. Count CP stayed evaluated at 3/63 on S33 BQ sheet E (detection 5/63; N=63 held; reasonCode null). Frozen loop RSI_PLATEAU. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
14 min readPhase 27 — Coverage Ladder 5/21; Count CP 3/63, Still Not 80/80
Phase 27 is the execute-close of the coverage-ladder 3→5 + S33 count-CP pack: COVERAGE_LADDER_PASS, five training_data graphs (S33, N23, and Kwai On preserved; S27 and N22 new), fabricatedGeometry false. Count CP is evaluated at 3/63 on S33 BQ sheet E — unblocked, not IDENTITY_ONLY_NO_GT, and not an accuracy win. S27 geometric takeoff was empty (NO_PLAN_POLYLINES); fittings-only still qualified. Live 圖則.pdf is a residual, not a coverage credit. Not Peak P&D 3D 80/80. Not corpus 80–90%.
12 min read