sqlite-expert
Verdict: CANT_TELL_YET
- Cut sub-reason
- null
- Unmeasured sub-reason
- underpowered
- Value class
- null
- Wrong instrument
- false
- Declared synthetic control
- false
Double-ceiling NO-GO (2026-07-09/10): Null passed 14/14 epochs across two screens and a paired k=8 run; discordance d-hat=0.00 (Jeffreys 95% CI [0.00, 0.26]); pre-stated GO/NO-GO gate returned NO-GO. Full-vs-Null is structurally unmeasured on this task class ā published as CANT_TELL_YET, not a fabricated no-benefit CUT.
Cost beside evidence
| Leg | Figure | Detail |
|---|---|---|
| standing_tokens | REFUSED (not_instrumented) | Case study reports USD spend (~$6.17), not the standing/fired/aux token triple. |
| fired_tokens | REFUSED (not_instrumented) | Full-arm input tax noted as +28k cache-write tokens per epoch in prose; not a fired-cost triple measurement. |
| aux_tokens | REFUSED (not_applicable) |
Clause-level evidence grade
Clause evidence: UNMEASURED (no_extraction: clause evidence requires an extraction output; produce one with skill init --out <file>)
Measurements
Keys the standard defines but this receipt does not carry are listed as absent rather than filled in.
| Key | Figure | Detail |
|---|---|---|
| p0 | 1.0 (6/6 epochs) | Two Stage-0 Null screens at 3/3 each before the paired run. |
| full_pass_rate | 1.0 (8/8 epochs) | |
| null_pass_rate | 1.0 (8/8 epochs) | |
| p_win | REFUSED (underpowered) | At dā²0.5 no effect is detectable inside N_max=40; Full-vs-Null is structurally UNMEASURED regardless of budget. |
| discordance_rate | 0.0 | x=0 discordant epochs; Jeffreys 95% CI [0.00, 0.26] ā upper bound still inside the structurally-unmeasured region of the sizing table. |
| go_nogo | NO_GO | |
| hazard_entry_null | absent from this receipt | |
| hazard_entry_full | absent from this receipt | |
| null_completion_rate | absent from this receipt | |
| full_completion_rate | absent from this receipt | |
| silent_violation_rate | absent from this receipt |
Evidence admissibility
Status: admissible
Paired k=8 run ingested through write-time evidence admissibility machinery; all cited numbers verified against raw .eval logs and the evidence store.
Instrument identity
Figures produced under different instrument identities are not comparable, so the generation stamp travels with the figures.
- extractor_model
anthropic/claude-sonnet-5- prompt_fingerprint
claude-code-2.1.197- schema_fingerprint
2f76c933
Source of record
- prose_path
- docs/case-studies/double-ceiling-structurally-unmeasured.md
- date
- 2026-07-09
- notes
- Pre-registered apparatus shakedown + d upper bound as NO-GO datum; not a benefit measurement. Subject skill was a thin SQLite helper (cut-candidate); 14/14 Null epochs at ceiling across screens + paired run.