Evaluation receipt / decision-ledger

Decision Ledger

The fully new skill passed one pair, then missed a weakly phrased proposal twice.

Inconclusive
3Attempted pairs
1Valid pair
1/3Treatment verifier passes
0/3Baseline verifier passes

What this proves

Proved the release pipeline can publish an honest miss from a brand-new skill.

Status

inconclusive. Only one of three attempted pairs remained valid, so the acceptance gate rejected an efficacy claim.

Method

Three attempted paired Harbor runs with a frozen verifier and invalid pairs excluded.

Model

openai/gpt-5.6-sol

Limitations
  • One representative task
  • One model
  • Only one valid pair after objective-gate failures
Exact packagesha256:2ba9d3b76b8ac66b57ff7bed01827c25d4298c1cdf21a65988590c967dbe301a