Evaluation receipt / agent-receipt
Agent Receipt
Both arms passed the corrected evidence-binding checks, and all three blind comparisons tied.
Inconclusive0/3Treatment wins
3/3Treatment verifier passes
3/3Baseline verifier passes
0Tool errors
What this proves
Produced public-safe receipts without demonstrating a measurable advantage.
Status
inconclusive. The corrected comparison demonstrated no lift over the empty baseline.
Method
Three blind paired Harbor runs against the same agent without the skill.
Model
openai/gpt-5.6-sol
Limitations
- One representative task
- One model
- Three valid pairs
- Earlier undisclosed-schema comparison was invalidated
Exact package
sha256:fa93aa2dc0b9b4bc32c421ef972762c3f62c968a37a1b5497b73dfa81d2e0139