Evaluation receipt / frontend-slides

Frontend Slides

A controlled three-pair evaluation of the pinned Frontend Slides skill on a complete, pre-approved seven-slide investor-deck brief.

Passed evaluation ยท adopted
2 / 1 / 0wins / ties / losses
3 of 3skill outputs passed
1 of 3baseline outputs passed
6recorded agent runs

What this proves

On this fixed-stage deck task, the skill won two blind pairs and tied one; its outputs passed every objective check, versus one of three baselines.

Status

ADOPTED. The skill passed Edge's preregistered three-pair adoption gate: a strict majority of blind wins, treatment verification passed, and treatment tool errors did not exceed baseline.

Method

Three blind baseline-versus-skill pairs in isolated Harbor containers. Same model, prompt, files, and objective checker; only the skill installation changed.

Model

OpenAI GPT-5.6 Sol

Limitations
  • One task family with three paired samples; no confidence interval can support a catalog-wide efficacy claim.
  • The brief supplied an already-approved visual direction; the interactive style-selection phase was not evaluated.
  • Visual quality was judged from recorded artifacts, not audience outcomes.
Exact packagesha256:832994fe1dfcce2aa7ceca9a1b7b708eca94becef242713999855f4e946cf4d5