Same model. Same prompt.
Different skill.
We give one agent one task, change only the skill, and keep every first attempt. No rerolls, no human edits. A blind judge compares each output with the no-skill run in both presentation orders.
Website: a launch page for an AI wearable




| Skill (told to load it) | With skill | No skill | Delta | Blind preference |
|---|---|---|---|---|
| frontend-design anthropics/skills @ 34040c9 | 74.5 | 63.5 | +11.0 | Won both orders |
| ui-ux-pro-max nextlevelbuilder/ui-ux-pro-max-skill @ 15de38f | 80.0 | 73.0 | +7.0 | Won both orders |
| taste-skill leonxlnx/taste-skill · design-taste-frontend @ e79ca9e | 72.5 | 80.0 | -7.5 | Lost both orders |
Installed is not used
| Skill (installed only) | Agent loaded it | With skill | No skill | Delta |
|---|---|---|---|---|
| frontend-design | No | 78.5 | 72.0 | +6.5 |
| ui-ux-pro-max | No | 81.5 | 77.5 | +4.0 |
| taste-skill | Yes | 77.5 | 78.0 | -0.5 |
Prompt, verbatim:
Build the homepage for ORBIT, a new AI wearable computer. ORBIT is a small screenless device you wear. You talk to it naturally and it handles messages, research, reminders, navigation and everyday tasks without pulling out your phone. Price: $399. Audience: design-conscious early adopters. Core idea: “Your AI, without the screen.” Build a launch-ready one-page website as a single HTML file with responsive CSS and minimal JavaScript. Do not ask me questions. You choose the entire visual direction. No stock imagery. If you need visuals, create them with HTML/CSS/SVG/code.
- One attempt per arm, so each delta is a single paired comparison, not a mean. The same no-skill page scored between 63.5 and 80 against different opponents: treat differences under about 8 points as noise.
- “Told to load it” adds one disclosed sentence to the system prompt of the skill arm: “The <name> skill is installed. Load it with the Skill tool before you start, and follow it.” The user prompt is identical.
- Criteria: visual hierarchy, originality, coherence, polish, launch-readiness.
Launch video: 10 public video skills
| Skill | Score | vs no skill |
|---|---|---|
| happy-skills-video-producer notedit/happy-skills | 91 | +8.0 |
| 30x-web-to-video norahe0304-art/30x-video | 89 | +6.0 |
| edge-launch-video-engine Edge (our own skill) | 87 | +4.0 |
| community-remotion-suite phamthanhnghia/remotion-agent-skills | 85 | +2.0 |
| No skill (baseline) same agent, no skill | 83 | |
| hyperframes heygen-com/hyperframes | 81 | -2.0 |
| amplitude-launch-video amplitude/builder-skills | 77 | -6.0 |
| memex-product-launch-video memex-lab/product-launch-video-skill | 75 | -8.0 |
| every-product-launch-video EveryInc/product-launch-video | 73 | -10.0 |
| remotion-official remotion-dev/skills | 69 | -14.0 |
| serena-launch-video serenakeyitan/launch-video-skill | 66 | -17.0 |











- One task, one attempt per skill, one AI judge. Frame review cannot see motion timing, transitions, audio quality or text dwell time; the planned human playback round has not run yet.
- edge-launch-video-engine is Edge's own skill: read its rank with that in mind.
- VideoZero Motion Canvas is excluded: its clean run hit an account usage limit and left an invalid MP4.
- Four of the eleven films rendered without audio and two came in under the length specification; those handicaps are inside the visual scores.