Jev in production › Evaluation and testing
An eval() function in LangWatchQL, the CLI and the REST API asks Jev a question over every stored conversation, trace or LLM call and returns a calibrated verdict per row. 97% agreement with human labels on a support-chat benchmark (company).
For the project's own README, linking back here: