Jev in production › Evaluation and testing

jevbench (Gxutxm)

An independent preregistered study on 1,500 ChaosNLI pairs finds Jev calibrated on consensus items but overconfident on contested ones, at 81% agreement (author).

Open on GitHub ↗

81%measured, as published by the source
Use
Evaluation and testing
Industry
Education and research
Form
Write-up
Stage
Announced
Listed
2026-09-29
Found via
discord
Repository
GautamTalksDev/jevbench
Stars
1
Forks
0
Last push
2026-10-02
Language
Python
License
MIT

The README opens with

Paper (latest version): doi.org/10.5281/zenodo.22971491 · v1.1: zenodo.org/records/23032384 · v1.0: zenodo.org/records/22971492 Preregistration: zenodo.org/records/22971413 (DOI 10.5281/zenodo.22971413) Plain-language write-up: docs/WRITEUP.md · 11-minute video: youtu.be/C6chAhTWvP4

v1.1 (29 Sep 2026) adds missing related work and credits consistency resampling (Bröcker and Smith 2007) as the basis of the bias correction. No data or results changed.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing