How we choose.
Edge evaluates a versioned skill package as an installable capability: what it enables, how reliably it works, and what evidence supports the claim.
A verdict needs receipts.
Opportunity, request match and evidence quality remain separate signals. A relevant skill is not automatically a good one, and a polished example is not a measured result.
Find work where judgment matters.
We begin with a user outcome, not a repository count.
Opportunity measures whether a task is common, consequential, underserved, or difficult for agents. It decides where the library looks. It never becomes evidence that a particular skill is good.
Candidates come from original publisher repositories, public registries, observed requests, and research task sets. Forks and copied packages are deduplicated before review.
Inspect the complete package.
A skill is every instruction, script, reference, asset, dependency, and permission it brings with it.
We identify the canonical source, pin a commit, inventory every file, and record licensing and setup requirements. We inspect network destinations, filesystem access, credential handling, remote instructions, dependencies, and executable paths.
Security and distribution rights are gates. A publisher identity, download count, or text scan is not proof of safety.
Keep request match separate.
A relevant skill is not automatically a high-quality skill.
Request match measures how closely a package fits the work described now. It can use task wording, intended output, constraints, and runtime compatibility. It does not include popularity or evaluation results.
When no package fits, Edge returns no recommendation. Adding more instructions is not always useful.
Compare one change at a time.
Same task, same capable starting agent, one intended difference: the skill package.
We pin model, runtime, tools, inputs, package hash, rubric, thresholds, and stopping rule. Baseline and treatment receive the same source material and permissions. Additional tools, context, time, and compute are disclosed.
Held-out tasks cover normal cases, difficult cases, and failures. Deterministic checks come first. Blind pairwise judging and human adjudication handle qualitative disagreements.
Publish the evidence and uncertainty.
Evidence quality describes what a result can support.
We retain inputs, outputs, failures, ties, losses, judge agreement, task count, run count, and uncertainty. A polished example can explain a workflow, but it cannot establish performance.
SkillBench informs which domains deserve investigation. Its bundle-level results are not proof for an individual Edge skill without isolated testing.
Score quality without hiding gaps.
The Edge score is package quality supported by evidence. It is distinct from request match and opportunity.
The proposed weighting is 50% measured evaluation, 20% domain-expert review, 15% package reliability, 10% source and attribution, and 5% usability. Missing dimensions remain visibly missing.
A public score requires passed gates, a matching package hash, documented evidence, and a current named approval. The launch catalog displays no fabricated scores.
Trust belongs to a version.
Changed code does not inherit an earlier review.
The reviewed, evaluated, and installed package must match. Upstream changes trigger a new scan and, when relevant, a new evaluation. Stale, superseded, and withdrawn records remain explicit.
Every recommendation keeps its provenance, reviewer, review date, evidence references, and removal path.