Jev in production › Agent tool and action selection
Jev replaces LLM-based tool selection in MCP agents across 612 tools; author reports 98.8% accuracy on 400 tickets (author).

Code Mode for MCP, where the sub-model is a decision model, not an LLM.
Your agent doesn't need an LLM for every decision. 612 tools: 77% fewer tokens. 400 tickets: 5.3x cheaper, 98.8% accurate.
I benchmarked toolJev with hosted Jev on MCPToolBench++, LiveMCPBench, When2Call and live Claude Haiku agents. Several results went against my first design, and the design changed to match.
Question Result (hosted Jev) Takeaway --------- Does it pick the right tool? Right tool ranked first: 83% on MCPToolBench++, 54% on LiveMCPBench, 99% on When2Call. Retrieval alone: 73%, 40%, 92%. Jev alone, with no retrieval: 18% Retrieval shortlists, Jev picks Does it know when no tool fits? AUROC 0.94 on When2Call near-misses (embeddings 0.74), 0.72 on...
For the project's own README, linking back here: