← Back to Arena

Open reproduction · Final public run

Six benchmarks.
No hidden losses.

We gave an AI agent 48 hours to reproduce TypeSafe’s Jev using public information and an open 2B model. This is the complete public scorecard—including the regressions.

Jev48 Open model benchmark
Independent · not affiliated with TypeSafe
Benchmark suites6
Evaluated cases6,300
Suite wins1
Additional spend<$1

Jev48 vs. published Jev results.

Accuracy is shown where it is the common metric. Protocol notes disclose paired, unpaired, and repeated-reference comparisons.

2,000 decisions

Typed decisions

-15.0 pts
Jev4857.7%
Jev72.7%
Aggregate / unpaired
2,000 emails

PhishNChips

-12.6 pts
Jev4850.0%
Jev62.6%
Aggregate / unpaired
231 tasks

JevBench public

-16.9 pts
Jev4869.7%
Jev86.6%
Paired outcomes
300 texts

BTZSC pilot

+8.0 pts
Jev4883.3%
Jev75.3%
Aggregate / unpaired
480 rule decisions

Code review

-17.2 pts
Jev4881.9%
Jev99.0%
Repeated Jev reference
1,289 cases

CLASH conflicts

-98.6 pts
Jev480.0%
Jev98.6%
Aggregate / unpaired

One win. Five losses. All published.

Jev48 beats the published Jev reference on BTZSC and reaches the same ceiling on JevBench Easy. It falls short elsewhere. The point is an auditable result, not a selected one.

01 · FreezeCode before outcomes
02 · IncludeFull suites only
03 · VerifyImmutable receipts