Jev in production › Evaluation and testing
Inference-time loop in which Jev scores each reasoning path a larger model writes and returns the highest-scoring answer, branching more paths when confidence is low. Lifted accuracy from 76.7 % to 86.7 % on GPQA Diamond and AIME rows (author).
MetaCog wraps any language model — local or API — in an inference-time metacognition loop. The model writes one answer; a separate judge scores it. If the judge is confident, that is the answer. If not, the model writes several more thought paths in parallel, the judge scores each path and each distinct final answer, and the best one is returned. No training, no weights touched.
Measured: on 180 paired GPQA Diamond + AIME rows across three hosted thinkers, the thinker alone scores 76.7 %, MetaCog 86.7 % (+11 fixed / −2 lost vs the previous MetaCog default, p = 0.02, 0.89× wall time, 1.39× thinker tokens). Pooled over 856 rows on four benchmarks, the earlier configuration lifted 75.0 % → 79.4 %. Full tables, ledger of every idea tried and charts:...
For the project's own README, linking back here: