Evaluation receipt / agent-receipt

Agent Receipt

Both arms passed the corrected evidence-binding checks, and all three blind comparisons tied.

Inconclusive
0/3Treatment wins
3/3Treatment verifier passes
3/3Baseline verifier passes
0Tool errors

What this proves

Produced public-safe receipts without demonstrating a measurable advantage.

Status

inconclusive. The corrected comparison demonstrated no lift over the empty baseline.

Method

Three blind paired Harbor runs against the same agent without the skill.

Model

openai/gpt-5.6-sol

Limitations
  • One representative task
  • One model
  • Three valid pairs
  • Earlier undisclosed-schema comparison was invalidated
Exact packagesha256:fa93aa2dc0b9b4bc32c421ef972762c3f62c968a37a1b5497b73dfa81d2e0139