Skill Harness: published receipts

append-only-evidence-design

Verdict: CANT_TELL_YET

Cut sub-reason
null
Unmeasured sub-reason
null
Value class
calibration
Wrong instrument
true
Declared synthetic control
false

append-only-evidence-design reclassified to CANT_TELL_YET (wrong instrument): Null p0=1.00 is above the transformative bar, but value_class=calibration means the transformative-lift instrument cannot see this skill's value. CUT(subsumed) withheld.

Cost beside evidence

Cost triple
Leg Figure Detail
standing_tokensREFUSED (not_instrumented)
fired_tokensREFUSED (not_instrumented)
aux_tokensREFUSED (not_applicable)

Clause-level evidence grade

Clause evidence: UNMEASURED (no_extraction: clause evidence requires an extraction output; produce one with skill init --out <file>)

Measurements

Keys the standard defines but this receipt does not carry are listed as absent rather than filled in.

Measured values and typed refusals
Key Figure Detail
p01.0 (3/3 epochs)Bare-arm / Null screen ceiling under the live end-to-end run.
full_pass_rateabsent from this receipt
null_pass_rateabsent from this receipt
p_winabsent from this receipt
discordance_rateabsent from this receipt
go_nogoNOT_APPLICABLE
hazard_entry_nullabsent from this receipt
hazard_entry_fullabsent from this receipt
null_completion_rateabsent from this receipt
full_completion_rateabsent from this receipt
silent_violation_rateabsent from this receipt

Evidence admissibility

Status: admissible

Live end-to-end screen store-backed; evidence admissibility applied at write time. Historical pre-guard disposition was CUT(subsumed); this receipt is the post-guard reclassification.

Instrument identity

Figures produced under different instrument identities are not comparable, so the generation stamp travels with the figures.

extractor_model
anthropic/claude-sonnet-5
prompt_fingerprint
claude-code-2.1.197
schema_fingerprint
2f76c933

Source of record

prose_path
README.md
date
2026-07-20
notes
Value-class guard reclassification. Pre-guard live screen returned CUT(subsumed) at p0=1.00; skill is registered calibration, so above-bar p0 is wrong-instrument CANT_TELL_YET. Screen counts also in docs/observations/OBS-0005-append-only-evidence-design.md (batch-1) and the 2026-07-20 live re-screen.