Evaluation receipt / taste-skill

Taste Skill

A controlled three-pair evaluation of the pinned Taste Skill on a technical B2B landing-page redesign.

Passed evaluation ยท adopted
3 / 0 / 0wins / ties / losses
3 of 3skill outputs passed
0 of 3baseline outputs passed
6recorded agent runs

What this proves

On this landing-page redesign, the skill won all three blind pairs and passed all three objective checks; no baseline did.

Status

ADOPTED. The skill passed Edge's preregistered three-pair adoption gate: a strict majority of blind wins, treatment verification passed, and treatment tool errors did not exceed baseline.

Method

Three blind baseline-versus-skill pairs in isolated Harbor containers. Same model, prompt, source page, and objective checker; only the skill installation changed.

Model

OpenAI GPT-5.6 Sol

Limitations
  • One task family with three paired samples; no confidence interval can support a catalog-wide efficacy claim.
  • The skill explicitly excludes dashboards; this evaluation stays within its landing-page scope.
  • No production conversion or user study was measured.
Exact packagesha256:aa194351b246b8b4799099d4ed7b033d29eab6e6e3d58d8d2172978be7b3ec89