Edge Copy setup link

Benchmarks / Methodology

Results pending. No results are shown here.

How we test. Edge Pro and no skills.

The same agent attempts each eligible specialist task twice, once without skills and once with Edge Pro. Task inputs, tools, model and budget are matched. An objective verifier decides whether the task was completed.

How we test

Public specialist skills and task candidates are selected before agent results. A separate pilot pool checks whether the tasks and runtime are executable. The evaluation set remains fixed once eligibility, inputs, verifiers and runtime are locked.

The two arms

ArmTreatment
A · No skillsCodex GPT-6 Sol, medium reasoning, without skills or Edge MCP.
E · Edge ProThe same agent with the frozen Edge routing instruction and Edge Pro connector. Specialist content reaches it through the recorded Edge transport.

Model, package, catalog, backend, limits and token-accounting details must be pinned in the final runtime lock before the pilot begins.

Tasks

The preparation pool contains 42 candidates in six specialist domains, with 18 pilot candidates and 24 evaluation candidates. These are candidates, not a final eligible denominator. Each task needs a public source, frozen inputs, an independently calibrated objective verifier, and a relevant pre-existing public skill in the catalog. Retrieval success does not determine eligibility.

SkillsBench is a separate later study. No SkillsBench task or outcome enters this pool.

Fairness rules

  • No skill is written or edited after the tasks are known.
  • Both arms receive the same task prompt, base checkout, ordinary tools, network policy and resource limits. Edge access is the intended treatment difference.
  • Task eligibility and evaluation membership are fixed before pilot results. Retrieval failures, easy baseline tasks and Edge losses remain visible.
  • One primary solve attempt per task and arm. An infrastructure pause can be repaired and resumed without silently replacing tasks or selecting the best retry.

Verification

Frozen, task-specific tests or format checks establish pass or fail. An LLM judge does not decide task success. Missing or invalid verifier results stay unresolved until verification is recovered; they are never silently counted as passes, failures or exclusions.

What we report

We publish a row for every selected evaluation task and each arm: objective result, verifier evidence, runtime, attempts and failure reason. Paired Edge wins, ties and losses include both-pass and both-fail ties. Pilots appear separately.

For Edge runs we also show whether the agent searched, received a nonempty verified skill payload, and followed applicable guidance. Search or load success never removes a task from the denominator. We report model tokens for the agent and Edge provider where available, with coverage clearly marked if provider usage cannot be verified.

Current status

The task pool and preparation protocol exist. Final eligibility, verifier calibration, Pro runtime and token accounting still need their launch lock. No pilot or evaluation result is reported here yet.

Limitations

Results will describe the tested tasks, model and catalog snapshot. The evaluation set is small within each domain. Every loss and unresolved infrastructure case will be shown alongside any aggregate.