Evaluation receipt / system-prompt-doctor
System Prompt Doctor
The tested skill reduced performance on a conflicting-prompt repair task.
Negative in this pilot1/3Treatment wins
1/3Treatment verifier passes
2/3Baseline verifier passes
0Tool errors
What this proves
Exposed a concrete regression that must be fixed before efficacy can be claimed.
Status
negative. The baseline won two of three valid pairs. The package carries no efficacy claim.
Method
Three blind paired Harbor runs against the same agent without the skill.
Model
openai/gpt-5.6-sol
Limitations
- One representative prompt-repair task
- One model
- Three paired samples
Exact package
sha256:a4764d030b749eb4043d4b26ed248b6ef630c9155afcbace8696e4a66ea4c25e