2,000 decisions
-15.0 ptsTyped decisions
Jev4857.7%
Jev72.7%
Open reproduction · Final public run
We gave an AI agent 48 hours to reproduce TypeSafe’s Jev using public information and an open 2B model. This is the complete public scorecard—including the regressions.
Accuracy is shown where it is the common metric. Protocol notes disclose paired, unpaired, and repeated-reference comparisons.
Jev48 beats the published Jev reference on BTZSC and reaches the same ceiling on JevBench Easy. It falls short elsewhere. The point is an auditable result, not a selected one.