Jev in production › Evaluation and testing
Skill-testing CLI with an optional Jev grader (--backend typesafe) that judges each agent run against plain-English task criteria and returns a calibrated pass probability and failure kind. ~30x cheaper than Claude grading (company).
A CLI tool to help you write tests for skills, test them across agents, and iterate until they work, from the Pinecone DevRel team.
Test how well skills work against tasks, across agents, locally and remotely. Customize sandboxes for how agents should start, and graders for how agents should work.
Use traces to iteratively refine skills and optimize them against tasks.
Benchmark against skills, docs, and baselines. And, even run in parallel simulatenously for faster execution.
For the project's own README, linking back here: