Evaluation receipt / last30days

/last30days

A controlled three-pair evaluation of the pinned /last30days skill on a date-bounded browser-agent landscape report.

No measured lift ยท rejected
0 / 3 / 0wins / ties / losses
3 of 3skill outputs passed
3 of 3baseline outputs passed
6recorded agent runs

What this proves

On this date-bounded research task, all three blind pairs tied and both arms passed every objective check; this run did not measure an advantage from installing the skill.

Status

REJECTED. All three valid pairs tied. Under Edge's preregistered gate, a tie without an approved probation case is rejected rather than called inconclusive.

Method

Three blind baseline-versus-skill pairs in isolated Harbor containers. Same model, prompt, date window, and objective checker; only the skill installation changed.

Model

OpenAI GPT-5.6 Sol

Limitations
  • One task family with three paired samples; no confidence interval can support a catalog-wide efficacy claim.
  • The harness confirmed installation but could not determine whether the agent loaded the skill instructions.
  • The blind judge could verify artifact and run metadata but could not compare the full report contents, so substantive differences may be under-detected.
Exact packagesha256:bdde67d5df7fd7f73dbfff4ee7f9fea919babedd5a5b17690e6d216fc1769bc6