back to top
HomeTechMiniCPM5-1B Shows Why the Small-Model Race Isn't Over

MiniCPM5-1B Shows Why the Small-Model Race Isn’t Over

- Advertisement -

A 1B model scoring 40.42 on AIME 2025 should not be possible. AIME is the American Invitational Mathematics Examination, the kind of test that filters out most humans who attempt it. Qwen3-0.6B scores 16.25 on the same benchmark. LFM2.5-1.2B, a larger model, scores 31.88. MiniCPM5-1B, at roughly one billion parameters, beats both.

OpenBMB dropped MiniCPM5-1B, the first model in their MiniCPM5 series, and it’s built specifically for the scenarios like on-device deployment, resource-constrained environments, local inference on consumer hardware.

The AIME score is surprising. The telecom agent benchmark is even more surprising. And then there’s the desktop pet. We’ll get to that.

One checkpoint, two personalities

With most models you get the fast version or the reasoning version. MiniCPM5-1B gives you both from the same weights.

Switch enable_thinking to True and the model enters deliberate reasoning mode, working through problems step by step before answering. Switch it off and you get a fast conversational assistant.

Running two separate models on constrained hardware is often not practical. Having one model that adapts to the task without requiring you to manage multiple downloads and contexts is a real convenience.

The recommended settings reflect the difference in what each mode is doing. Think mode runs at temperature 0.9 with top_p 0.95 to give the reasoning process room to explore. No Think mode drops to 0.7 for more predictable, focused responses. Both are sensible defaults and both can be adjusted through standard inference parameters.

The training recipe that explains the numbers

OpenBMB published the full training stack and it’s worth understanding because it explains why MiniCPM5-1B outperforms models of similar or larger size on specific tasks.

The post-training process runs in three stages: supervised fine-tuning, reinforcement learning, and On-Policy Distillation. The SFT stage used 400B tokens total, split between deep-thinking and hybrid-thinking data, establishing the baseline reasoning and chat capabilities. Then RL trained specialized teachers for math, code, closed-book QA, writing, and instruction following separately. OPD then distilled those teachers back into the single release model.

The result of RL plus OPD combined is a 16 point average score improvement over the SFT-only checkpoint, alongside a 29 percentage point drop in responses that hit the maximum token budget. That second number matters as much as the first. A model that reasons efficiently and stops when it has an answer is more useful in production than one that pads responses until it runs out of context. Overlong responses waste compute, slow inference, and often indicate the model is uncertain rather than thorough.

The training data is also fully released alongside the model as Ultra-FineWeb, Ultra-FineWeb-L3, UltraData-Math, and UltraData-SFT-2605 for anyone who wants to study or reproduce the recipe.

Benchmarks

minicpm5 1b evaluation results
via: MiniCPM5 1B HF

The benchmark that stands out most isn’t AIME. It’s τ2-Bench Telecom at 79.53.

τ2-Bench tests multi-step agentic task completion in realistic enterprise environments, specifically telecom workflows which are notoriously complex and procedural. Qwen3-0.6B scores 21.10 on the same benchmark. LFM2.5-1.2B scores 19.60. MiniCPM5-1B nearly quadruples both of them. That gap is large enough that it’s worth treating as a genuine capability difference rather than benchmark noise.

Here’s a tight comparison across the benchmarks that tell the story:

BenchmarkMiniCPM5-1B (Thinking)Qwen3-0.6BQwen3.5-0.8BLFM2.5-1.2B
AIME 202540.4216.251.0431.88
AIME 202640.4212.290.2131.67
MATH-50091.6072.6030.4089.00
τ2-Bench Telecom79.5321.1047.7019.60
IFEval80.4159.8959.8984.84
BBH71.8947.8654.5857.32
LCB-v633.5216.005.3321.33

All figures from OpenBMB’s evaluation. Self-reported benchmarks should be read with that caveat.

The coding numbers deserves to be noted as well. LCB-Pro Easy at 22.68 against Qwen3-0.6B’s 4.12 and Qwen3.5-0.8B’s flat zero. OJBench at 7.33 against both Qwen models below 1. These are competitive programming benchmarks and the gap at 1B scale is unusually wide.

GPQA-Diamond at 26.26 trails LFM2.5-1.2B’s 34.85, which is the clearest benchmark where a larger model pulls ahead on difficult scientific reasoning.

Related: MiniCPM-V 4.6: The 1.3B Model Running on Your Phone That Challenges Much Larger Rivals

How to run it, and about that desktop pet

MiniCPM5-1B uses standard LlamaForCausalLM architecture which means mainstream inference engines load it directly without custom kernels or model-code forks. Ollama, vLLM, SGLang, llama.cpp, LM Studio, and MLX for Apple Silicon all work out of the box. For tool calling specifically, SGLang is the recommended backend since it handles MiniCPM5’s XML-style tool calls natively through a built-in parser.

The model is on Hugging Face under Apache 2.0, which is as permissive as licenses get.

Now about the desktop pet. OpenBMB ships MiniCPM-Desk-Pet alongside the model, a local desktop companion driven entirely by MiniCPM5-1B running on your own hardware. It supports Apple Silicon, NVIDIA GPU, and CPU paths, integrates with coding agents like Cursor and Claude Code, and supports LoRA persona switching. It’s genuinely unusual to see a serious reasoning model release bundled with something this playful and it says something about who OpenBMB is building for. Not just researchers and enterprise teams. People who want capable AI running locally in whatever form makes sense for them.

If you’re evaluating small models for local deployment, agentic tool use, or constrained inference scenarios, MiniCPM5-1B belongs on your shortlist.

Want more stories worth your time?

Add us to your Google favorites. We cover the tech stories, AI developments, and open-source projects that are easy to miss in the noise.

Add as a preferred source on Google

Don’t miss any Tech Story

Subscribe To Firethering NewsLetter

You Can Unsubscribe Anytime! Read more in our privacy policy

LEAVE A REPLY

Please enter your comment!
Please enter your name here

YOU MAY ALSO LIKE
OpenAI Cuts Off Cursor SpaceX Deal Triggers Nov 12 Cutoff

OpenAI Cuts Off Cursor After SpaceX Acquisition With Nov. 12 Deadline

0
OpenAI is cutting Cursor off. The company has notified SpaceX that it intends to end Cursor’s direct access to OpenAI models on November 12, 2026, following SpaceX’s acquisition of Anysphere, the company behind Cursor. For developers who rely on GPT models inside Cursor, that puts a clock on something that has become part of their daily workflow. But Cursor itself isn't going away. The question is what actually changes when one of the models behind your AI coding workflow suddenly disappears and whether Cursor can move on without OpenAI as easily as it might seem.
Can Twitter.now Survive The New Twitter Faces a Bigger Problem Than X

Can Twitter.now Survive? The New Twitter Faces a Bigger Problem Than X

0
The bird is back. After Elon Musk turned Twitter into X, a startup called Operation Bluebird has now launched Twitter.now using the Twitter name, the familiar bird, and an ambitious promise to build a new kind of public square. But this isn't really about bringing back an old brand. Twitter.now says it wants to build something different: a social network centered on trust, transparency, and user choice, where people can decide how much reach different posts deserve instead of leaving everything to an opaque algorithm. Early access costs $20, and the platform is already offering founding members the chance to claim handles and even become Founder #00001. The question, though, isn't whether Twitter.now can bring back the name. It's whether it can bring back the people. Because in social media, having the right name is one thing. Convincing millions of people to leave an established network, rebuild their communities somewhere else, and give a tiny new platform a reason to exist is an entirely different game. And that's where its actual challenge begins.
A Bluetooth Glitch Just Exposed How AliExpress Fingerprints Vistors' Browsers Through Audio

A Bluetooth Glitch Just Exposed How AliExpress Fingerprints Vistors’ Browsers Through Audio

0
Your headphones can tell you when a website is doing something strange. At least, that's what happened when a developer noticed his Bluetooth headphones stopped switching back to his phone whenever an AliExpress tab was open on his PC. There was no music playing from the computer. No video. No audible sound. Closing the AliExpress tab immediately fixed the problem. That strange Bluetooth glitch led him to something much more interesting: AliExpress was quietly running audio through the browser, not to listen to the user, but to measure how the browser and device processed a known sound. The technique is called Web Audio fingerprinting. And while it sounds like something from an older era of browser tracking, the discovery shows just how much information a website can extract from seemingly ordinary browser APIs. Brave later highlighted the case, explaining that AliExpress wasn't recording users' audio. Instead, it was generating a silent signal and measuring how a device processed it to help fingerprint the browser. So what exactly was happening inside that supposedly silent AliExpress tab and why did a fingerprinting technique end up interfering with a pair of Bluetooth headphones?