VibeThinker-3B, a 3B-parameter small model, achieves frontier-level verifiable reasoning performance (AIME26 94.3, LiveCodeBench v6 80.2%) rivaling much larger models. It proposes the Parametric Compression-Coverage Hypothesis, showing compact models can excel in reasoning.
VibeThinker-3B, a 3B-parameter small dense model, achieves frontier-level performance on verifiable reasoning tasks: AIME26 94.3, LiveCodeBench v6 80.2% Pass@1, and 96.1% acceptance rate on recent LeetCode contests. It also maintains instruction controllability with IFEval 93.4.
Previously, reasoning ability was considered the domain of large models. VibeThinker-3B shows that small models can reach frontier performance through an optimized pipeline (curriculum SFT, multi-domain RL, offline self-distillation), extending prior 1.5B work.
The Parametric Compression-Coverage Hypothesis suggests verifiable reasoning can be compressed into compact reasoning cores, while open-domain knowledge requires broad parameter coverage. This positions small models not just as efficient substitutes but as a complementary path to frontier performance in reasoning.