Jev in production › Evaluation and testing

LangWatch Instant Evals

An eval() function in LangWatchQL, the CLI and the REST API asks Jev a question over every stored conversation, trace or LLM call and returns a calibrated verdict per row. 97% agreement with human labels on a support-chat benchmark (company).

Open on langwatch.ai ↗

97%measured against a baseline, as published by the source
Use
Evaluation and testing
Industry
AI infrastructure
Form
Commercial product
Stage
In production
Plugs into
LangWatch
Listed
2026-09-19
Found via
discord

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing