Evaluation receipt / taste-skill
Taste Skill
A controlled three-pair evaluation of the pinned Taste Skill on a technical B2B landing-page redesign.
Passed evaluation ยท adopted3 / 0 / 0wins / ties / losses
3 of 3skill outputs passed
0 of 3baseline outputs passed
6recorded agent runs
What this proves
On this landing-page redesign, the skill won all three blind pairs and passed all three objective checks; no baseline did.
Status
ADOPTED. The skill passed Edge's preregistered three-pair adoption gate: a strict majority of blind wins, treatment verification passed, and treatment tool errors did not exceed baseline.
Method
Three blind baseline-versus-skill pairs in isolated Harbor containers. Same model, prompt, source page, and objective checker; only the skill installation changed.
Model
OpenAI GPT-5.6 Sol
Limitations
- One task family with three paired samples; no confidence interval can support a catalog-wide efficacy claim.
- The skill explicitly excludes dashboards; this evaluation stays within its landing-page scope.
- No production conversion or user study was measured.
Exact package
sha256:aa194351b246b8b4799099d4ed7b033d29eab6e6e3d58d8d2172978be7b3ec89