Jev in production › Model and agent routing
A local dual-system LLM server pairs a 3B model with Jev; author's benchmark shows tokens per second up but latency up 36% (author).

A dual-system, OpenAI-compatible local LLM server that runs small models (Phi-4, Qwen, Gemma, Llama) privately on your own machine, almost free.
DuoMind combines System 2 (a small local LLM via llama.cpp) with System 1 (a fast classification API) so a 3B model reasons like a much larger one — generation stays 100% local.
DuoMind has been benchmarked on real-world prompts with measurable results. The statistics below are from actual runs on a 3B model (Qwen2.5-Coder-3B-Instruct-Q4KM), comparing performance with Jev enabled versus disabled.
For the project's own README, linking back here: