Evaluation receipt / system-prompt-doctor

System Prompt Doctor

The tested skill reduced performance on a conflicting-prompt repair task.

Negative in this pilot
1/3Treatment wins
1/3Treatment verifier passes
2/3Baseline verifier passes
0Tool errors

What this proves

Exposed a concrete regression that must be fixed before efficacy can be claimed.

Status

negative. The baseline won two of three valid pairs. The package carries no efficacy claim.

Method

Three blind paired Harbor runs against the same agent without the skill.

Model

openai/gpt-5.6-sol

Limitations
  • One representative prompt-repair task
  • One model
  • Three paired samples
Exact packagesha256:a4764d030b749eb4043d4b26ed248b6ef630c9155afcbace8696e4a66ea4c25e