Jev in production › Evaluation and testing

MetaCog

Inference-time loop in which Jev scores each reasoning path a larger model writes and returns the highest-scoring answer, branching more paths when confidence is low. Lifted accuracy from 76.7 % to 86.7 % on GPQA Diamond and AIME rows (author).

Open on GitHub ↗

86.7 %measured against a baseline, as published by the source
Use
Evaluation and testing
Industry
AI infrastructure
Form
Open-source tool
Stage
Beta
Listed
2026-09-22
Found via
github
Repository
ItIsCuthNotCup/MetaCog
Stars
25
Forks
2
Last push
2026-09-30
Language
Python
License
MIT

The README opens with

MetaCog wraps any language model — local or API — in an inference-time metacognition loop. The model writes one answer; a separate judge scores it. If the judge is confident, that is the answer. If not, the model writes several more thought paths in parallel, the judge scores each path and each distinct final answer, and the best one is returned. No training, no weights touched.

Measured: on 180 paired GPQA Diamond + AIME rows across three hosted thinkers, the thinker alone scores 76.7 %, MetaCog 86.7 % (+11 fixed / −2 lost vs the previous MetaCog default, p = 0.02, 0.89× wall time, 1.39× thinker tokens). Pooled over 856 rows on four benchmarks, the earlier configuration lifted 75.0 % → 79.4 %. Full tables, ledger of every idea tried and charts:...

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing