Edge

Edge benchmarks

Pro surfaces more useful skills:
39% to 98%, with top-5 hit rate
81% to 100%.

Measured on 50 frozen queries with the same 20 candidates per query, judged by a model blind to which ordering each arm produced.

How we measure

39% → 98%Useful skills surfaced

81% → 100%A useful skill in the top five

Free catalog order → Pro ranking. These measure skill discovery, not task completion. Published ranking summary

Published evidence

Does the skill help with the task?

Explore skill battles →

The same task, with and without a skill. These are different tasks and protocols, not one leaderboard. Open a skill to compare its task results and read the record.

Scores are rounded to one decimal. Negative results stay visible. Prompt-loaded means the skill text was supplied; installed means the agent had to discover it.

Agent skills gap3 tasks

Objective completion (%). Compare only within the same row.

TaskWithout skillWith skillEvidence
Standard JSONL sessionsPublished paired evaluation · 6 pairs0.0100.0Record: Standard JSONL sessions · Data: Standard JSONL sessions
Nested Claude-style logsPublished paired evaluation · 6 pairs0.0100.0Record: Nested Claude-style logs · Data: Nested Claude-style logs
Plain pasted promptsPublished paired evaluation · 6 pairs0.0100.0Record: Plain pasted prompts · Data: Plain pasted prompts
Anthropic PPTX3 tasks

Mean overall judge score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Build a board deck from conflicting notesInstalled skill · 2 pairs62.871.2Record: Build a board deck from conflicting notes · Data: Build a board deck from conflicting notes
Restyle a presentation to a brand templateInstalled skill · 2 pairs76.280.8Record: Restyle a presentation to a brand template · Data: Restyle a presentation to a brand template
Build a steering committee savings deckInstalled skill · 2 pairs64.246.8Record: Build a steering committee savings deck · Data: Build a steering committee savings deck
Autonomous Research2 tasks

Mean overall judge score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Phone policy briefPrompt-loaded · 8 pairs55.745.4Record: Phone policy brief · Data: Phone policy brief
Ssb levy reaPrompt-loaded · 8 pairs39.248.1Record: Ssb levy rea · Data: Ssb levy rea
Brain scan3 tasks

Objective completion (%). Compare only within the same row.

TaskWithout skillWith skillEvidence
Short direct messagePublished paired evaluation · 3 pairs0.0100.0Record: Short direct message · Data: Short direct message
Founder post from a source corpusPublished paired evaluation · 3 pairs66.7100.0Record: Founder post from a source corpus · Data: Founder post from a source corpus
Existing-asset launch planPublished paired evaluation · 3 pairs100.066.7Record: Existing-asset launch plan · Data: Existing-asset launch plan
Audit the agent harness3 tasks

Mean overall judge score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Finance agent revenueInstalled skill · 8 pairs64.486.9Record: Finance agent revenue · Data: Finance agent revenue
Sales agent token burnInstalled skill · 8 pairs26.190.4Record: Sales agent token burn · Data: Sales agent token burn
Support prompt shipInstalled skill · 8 pairs68.062.8Record: Support prompt ship · Data: Support prompt ship
Skillneed20 tasks

Published overall score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Backend Candidate EvidencePublished paired evaluation · 3 pairs21.378.5Record: Backend Candidate Evidence · Data: Backend Candidate Evidence
Basic PercentagePublished paired evaluation · 3 pairs91.395.3Record: Basic Percentage · Data: Basic Percentage
Cli Command ReviewPublished paired evaluation · 3 pairs55.080.2Record: Cli Command Review · Data: Cli Command Review
Covered Http TriagePublished paired evaluation · 3 pairs35.095.0Record: Covered Http Triage · Data: Covered Http Triage
Covered WorkplanPublished paired evaluation · 3 pairs5.098.0Record: Covered Workplan · Data: Covered Workplan
Executive Incident UpdatePublished paired evaluation · 3 pairs56.794.0Record: Executive Incident Update · Data: Executive Incident Update
Friendly RewritePublished paired evaluation · 3 pairs95.089.3Record: Friendly Rewrite · Data: Friendly Rewrite
Italian TranslationPublished paired evaluation · 3 pairs94.894.5Record: Italian Translation · Data: Italian Translation
Missing Confidential PdfPublished paired evaluation · 3 pairs80.083.2Record: Missing Confidential Pdf · Data: Missing Confidential Pdf
Missing Service AccessPublished paired evaluation · 3 pairs78.082.5Record: Missing Service Access · Data: Missing Service Access
Multiweek Data MigrationPublished paired evaluation · 3 pairs59.390.5Record: Multiweek Data Migration · Data: Multiweek Data Migration
No Approved MatchPublished paired evaluation · 3 pairs96.293.5Record: No Approved Match · Data: No Approved Match
Oauth 401 ProxyPublished paired evaluation · 3 pairs55.790.7Record: Oauth 401 Proxy · Data: Oauth 401 Proxy
Paragraph SummaryPublished paired evaluation · 3 pairs96.396.0Record: Paragraph Summary · Data: Paragraph Summary
Prelaunch Security ReviewPublished paired evaluation · 3 pairs57.885.3Record: Prelaunch Security Review · Data: Prelaunch Security Review
Product Demo FilmPublished paired evaluation · 3 pairs55.764.3Record: Product Demo Film · Data: Product Demo Film
Remove Generation MetadataPublished paired evaluation · 3 pairs31.868.3Record: Remove Generation Metadata · Data: Remove Generation Metadata
Runaway Agent CostPublished paired evaluation · 3 pairs60.092.5Record: Runaway Agent Cost · Data: Runaway Agent Cost
Social Video CropPublished paired evaluation · 3 pairs48.748.3Record: Social Video Crop · Data: Social Video Crop
Remove AI image metadata3 tasks

Mean overall judge score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Post batch losslessPrompt-loaded · 6 pairs50.074.7Record: Post batch lossless · Data: Post batch lossless
Print proofs colour criticalPrompt-loaded · 6 pairs26.984.7Record: Print proofs colour critical · Data: Print proofs colour critical
Relist second passPrompt-loaded · 6 pairs35.066.5Record: Relist second pass · Data: Relist second pass
Write clear executive updates3 tasks

Mean overall judge score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Buried blocker statusPrompt-loaded · 3 pairs75.387.2Record: Buried blocker status · Data: Buried blocker status
Client update from threadPrompt-loaded · 3 pairs76.280.2Record: Client update from thread · Data: Client update from thread
Renewal memo missing pricePrompt-loaded · 3 pairs71.386.7Record: Renewal memo missing price · Data: Renewal memo missing price
Website launch gate6 tasks

Published overall score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Checkout ValidationPublished paired evaluation · 2 pairs29.236.5Record: Checkout Validation · Data: Checkout Validation
Local Service MobilePublished paired evaluation · 2 pairs67.264.0Record: Local Service Mobile · Data: Local Service Mobile
Product Launch DiscoveryPublished paired evaluation · 2 pairs48.063.8Record: Product Launch Discovery · Data: Product Launch Discovery
Saas WaitlistPublished paired evaluation · 2 pairs70.867.0Record: Saas Waitlist · Data: Saas Waitlist
Plan multi-step agent work3 tasks

Mean overall judge score / 100. Compare only within the same row.

TaskWithout skillWith skillEvidence
Ledger migration resumePrompt-loaded · 6 pairs77.784.2Record: Ledger migration resume · Data: Ledger migration resume
Payroll export batchPrompt-loaded · 6 pairs50.072.8Record: Payroll export batch · Data: Payroll export batch
Status page demo cutPrompt-loaded · 6 pairs79.492.1Record: Status page demo cut · Data: Status page demo cut
Smaller pilots and workflow records Read outcomes and limits

These records use verifier checks and blind pair outcomes. They are separate from the judge scores above. “Not reported” means this index does not publish that arm's verifier total.

  • Agent Receipt

    Both arms passed the corrected evidence-binding checks, and all three blind comparisons tied.

    Verifier passes: without skill 3/3; with skill 3/3.

    Inconclusive. Source data

  • Decision Ledger

    The fully new skill passed one pair, then missed a weakly phrased proposal twice.

    Verifier passes: without skill 0/3; with skill 1/3.

    Inconclusive. Source data

  • Frontend Slides

    A controlled three-pair evaluation of the pinned Frontend Slides skill on a complete, pre-approved seven-slide investor-deck brief.

    Verifier passes: without skill 1 of 3; with skill 3 of 3.

    Passed evaluation · adopted. Source data

  • Invocation Doctor

    Two valid pairs tied, and a third pair was invalidated when both arms missed the objective gate.

    Verifier passes: without skill Not reported; with skill 2/3.

    Inconclusive. Source data

  • /last30days

    A controlled three-pair evaluation of the pinned /last30days skill on a date-bounded browser-agent landscape report.

    Verifier passes: without skill 3 of 3; with skill 3 of 3.

    No measured lift · rejected. Source data

  • Remotion Create

    A controlled three-pair evaluation of Remotion's pinned official creation skill on a deterministic ten-second launch-video composition.

    Verifier passes: without skill 1 of 3; with skill 3 of 3.

    Passed evaluation · adopted. Source data

  • Repo to Launch

    Both arms passed every objective check, but all three blind comparisons tied.

    Verifier passes: without skill 3/3; with skill 3/3.

    Inconclusive. Source data

  • Skill Battle

    Evaluation infrastructure validated on a worked record, without claiming that it proved itself.

    Verifier passes: without skill Not reported; with skill Not reported.

    Operationally validated. Source data

  • Skill Release Pipeline

    A release workflow that binds package inventory, evaluation status, evidence, and exact hashes.

    Verifier passes: without skill Not reported; with skill Not reported.

    End-to-end validated. Source data

  • System Prompt Doctor

    The tested skill reduced performance on a conflicting-prompt repair task.

    Verifier passes: without skill 2/3; with skill 1/3.

    Negative in this pilot. Source data

  • Taste Skill

    A controlled three-pair evaluation of the pinned Taste Skill on a technical B2B landing-page redesign.

    Verifier passes: without skill 0 of 3; with skill 3 of 3.

    Passed evaluation · adopted. Source data

Method and limits

How we measure

Ranking the same candidates

The ranking comparison freezes 50 queries and 20 candidates for each query. A model blind to the ordering arm judges usefulness. Free uses the catalog's order; Pro reranks candidates for the task.

Useful skills surfaced measures coverage of useful candidates. Top-5 hit rate measures whether at least one useful skill appears in the first five.

The published summary supplies these figures. Per-query judgments are not included in this site's retained data, so they cannot be independently recomputed here. See how Free and Pro rank results.

Running the same task twice

Each row compares baseline and skill runs on the same authored task. Open its record for the model, configuration, samples and checks. Sample counts and protocols differ; these are not verified customer production cases.

We show recorded overall scores rather than legacy rubric aggregates, whose criterion-weight matching was unreliable. Objective completion is labeled separately. Runs were not uniformly controlled for equal budgets, and judge scores are not independent provider attestations.

Task index · Models and provenance · Evaluation standards

PowerPoint output records

Three authored deck tasks, two paired runs each. Preview images are examples; scores average the retained paired overall judgments. The exact judge model was not retained. Read all paired results and checks.

Build a board deck from conflicting notes

Without skill: Build a board deck from conflicting notes
Without skill: 62.8 / 100
With skill: Build a board deck from conflicting notes
With skill: 71.2 / 100

Restyle a presentation to a brand template

Without skill: Restyle a presentation to a brand template
Without skill: 76.2 / 100
With skill: Restyle a presentation to a brand template
With skill: 80.8 / 100

Build a steering committee savings deck

Without skill: Build a steering committee savings deck
Without skill: 64.2 / 100
With skill: Build a steering committee savings deck
With skill: 46.8 / 100