Jev in production › Evaluation and testing

LangChain

LangChain tested Jev as an agent-eval judge against GPT-5.6 and Claude Sonnet 4.6 judges on identical traces, then added Jev-as-a-judge evaluators to LangSmith Evals. Jev cost $0.00035/call (company).

Open on x.com ↗

$0.00035/callmeasured against a baseline, as published by the source
Use
Evaluation and testing
Industry
AI infrastructure
Form
Write-up
Stage
In production
Plugs into
LangSmith
Listed
2026-09-21
Found via
x

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing