Evaluation receipt / last30days
/last30days
A controlled three-pair evaluation of the pinned /last30days skill on a date-bounded browser-agent landscape report.
No measured lift ยท rejected0 / 3 / 0wins / ties / losses
3 of 3skill outputs passed
3 of 3baseline outputs passed
6recorded agent runs
What this proves
On this date-bounded research task, all three blind pairs tied and both arms passed every objective check; this run did not measure an advantage from installing the skill.
Status
REJECTED. All three valid pairs tied. Under Edge's preregistered gate, a tie without an approved probation case is rejected rather than called inconclusive.
Method
Three blind baseline-versus-skill pairs in isolated Harbor containers. Same model, prompt, date window, and objective checker; only the skill installation changed.
Model
OpenAI GPT-5.6 Sol
Limitations
- One task family with three paired samples; no confidence interval can support a catalog-wide efficacy claim.
- The harness confirmed installation but could not determine whether the agent loaded the skill instructions.
- The blind judge could verify artifact and run metadata but could not compare the full report contents, so substantive differences may be under-detected.
Exact package
sha256:bdde67d5df7fd7f73dbfff4ee7f9fea919babedd5a5b17690e6d216fc1769bc6