Evaluation receipt / frontend-slides
Frontend Slides
A controlled three-pair evaluation of the pinned Frontend Slides skill on a complete, pre-approved seven-slide investor-deck brief.
Passed evaluation ยท adopted2 / 1 / 0wins / ties / losses
3 of 3skill outputs passed
1 of 3baseline outputs passed
6recorded agent runs
What this proves
On this fixed-stage deck task, the skill won two blind pairs and tied one; its outputs passed every objective check, versus one of three baselines.
Status
ADOPTED. The skill passed Edge's preregistered three-pair adoption gate: a strict majority of blind wins, treatment verification passed, and treatment tool errors did not exceed baseline.
Method
Three blind baseline-versus-skill pairs in isolated Harbor containers. Same model, prompt, files, and objective checker; only the skill installation changed.
Model
OpenAI GPT-5.6 Sol
Limitations
- One task family with three paired samples; no confidence interval can support a catalog-wide efficacy claim.
- The brief supplied an already-approved visual direction; the interactive style-selection phase was not evaluated.
- Visual quality was judged from recorded artifacts, not audience outcomes.
Exact package
sha256:832994fe1dfcce2aa7ceca9a1b7b708eca94becef242713999855f4e946cf4d5