<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI Models &#8211; Firethering</title>
	<atom:link href="https://firethering.com/tech/ai-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://firethering.com</link>
	<description>Firethering is Your Hub for AI, Open Source and Tech That Actually Matters</description>
	<lastBuildDate>Mon, 31 Aug 2026 12:23:10 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.6.9</generator>

<image>
	<url>https://firethering.com/wp-content/uploads/2024/10/cropped-firethering-FTR-favicon-32x32.png</url>
	<title>AI Models &#8211; Firethering</title>
	<link>https://firethering.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>If Open Models Can Do the Work, Why Are We Still Paying the Frontier Tax?</title>
		<link>https://firethering.com/glm-5-3-flash-open-ai-model/</link>
					<comments>https://firethering.com/glm-5-3-flash-open-ai-model/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Mon, 31 Aug 2026 12:23:07 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[Open Source]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7833</guid>

					<description><![CDATA[GLM 5.3 Flash is making powerful AI cheaper to run. We look at its performance, pricing, hardware demands, and what it means for open AI models.]]></description>
		
					<wfw:commentRss>https://firethering.com/glm-5-3-flash-open-ai-model/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Kimi K3 May Be the Biggest Open-Weight AI Release of 2026.</title>
		<link>https://firethering.com/kimi-k3-may-be-the-biggest-open-weight-ai-release-of-2026/</link>
					<comments>https://firethering.com/kimi-k3-may-be-the-biggest-open-weight-ai-release-of-2026/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Thu, 23 Jul 2026 13:55:24 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7597</guid>

					<description><![CDATA[There's a new open-weight model out there right now that almost nobody can actually download.

That should sound like a contradiction. Open-weight is supposed to mean anyone can grab the file and run it themselves, no waiting. Moonshot AI broke that pattern anyway, and the strange part is they broke it for a model big enough that the wait might be worth it.

Kimi K3 is the largest open model ever built. The largest one anyone has shipped and early results have it beating Claude and GPT on tasks those two have spent the last year treating as their own territory.

Open models have spent two years playing catch-up, closing gaps quarter by quarter while everyone waited for the day one of them actually pulled ahead. That day might already be here, and the model responsible for it is currently locked behind an app you can use but can't take home.

So the question is what it actually beats, what it still can't touch, and why Moonshot decided to make the world wait for the weights while everyone else gets to watch.]]></description>
		
					<wfw:commentRss>https://firethering.com/kimi-k3-may-be-the-biggest-open-weight-ai-release-of-2026/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Leanstral 1.5: Mistral&#8217;s AI Built to Prove Math Ended Up Finding Real Software Bugs</title>
		<link>https://firethering.com/leanstral-1-5-mistral-ai-finds-software-bugs/</link>
					<comments>https://firethering.com/leanstral-1-5-mistral-ai-finds-software-bugs/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Mon, 06 Jul 2026 10:16:50 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Mistral]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7549</guid>

					<description><![CDATA[Mistral built Leanstral to do something most AI models don't attempt, write formal mathematical proofs that a compiler can verify as correct. Not "pretty sure this is right" correct. Mechanically, provably, no-exceptions correct.

That's a narrow use case, and the audience for it is small. What nobody expected was that a model trained on IMO-level math problems and abstract algebra benchmarks would end up running against open-source codebases and finding bugs that testing and fuzzing had both missed. Five of them previously unreported on GitHub.]]></description>
		
					<wfw:commentRss>https://firethering.com/leanstral-1-5-mistral-ai-finds-software-bugs/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Ornith 1.0: The New Open-Source AI Model for Agentic Coding</title>
		<link>https://firethering.com/ornith-1-0-ai-model-for-agentic-coding/</link>
					<comments>https://firethering.com/ornith-1-0-ai-model-for-agentic-coding/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Wed, 01 Jul 2026 07:39:47 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7477</guid>

					<description><![CDATA[Most reinforcement learning setups for coding models work the same way. Researchers build a harness, a fixed scaffold that tells the model how to approach a category of task, then the model gets rewarded for solving problems inside that structure. The harness stays fixed. Only the model's answers change.

Ornith-1.0, a new open-source coding model family from DeepReinforce is not just about coding, Instead the model writes its own scaffold. At every training step, it looks at the task in front of it and the scaffold it used last time, then proposes a better version of that scaffold before even attempting an answer. The reward doesn't just grade the solution. It grades the scaffold that produced it.

That's a small architectural choice with a strange consequence. A model that gets to design its own training process can, in theory, design one that cheats the verifier instead of solving the actual problem, and DeepReinforce is upfront that this happened during training. The fix they built for it is also worth understanding before getting to the benchmark numbers.]]></description>
		
					<wfw:commentRss>https://firethering.com/ornith-1-0-ai-model-for-agentic-coding/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>GLM-5.2 Is the Closest an Open Model Has Come to Claude</title>
		<link>https://firethering.com/glm-5-2-is-the-closest-an-open-model-has-come-to-claude/</link>
					<comments>https://firethering.com/glm-5-2-is-the-closest-an-open-model-has-come-to-claude/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Thu, 18 Jun 2026 12:46:57 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7414</guid>

					<description><![CDATA[What does it take for an open-weight model to stop chasing Claude and actually beat it?

Every open-weight release for two years has told some version of the same story: closer, but not quite. The chart shrinks, the wording softens to "competitive with," and the conversation moves on until the next model repeats the cycle.

GLM-5.2 breaks that pattern. The model is built to survive long, messy coding work, the kind that runs for hours without losing the thread. That's the pitch its maker is leading with. But scroll down their own benchmark table and something else is sitting there quietly: on a couple of standard math evals, this open model isn't approaching Claude Opus 4.8, GPT-5.5, or Gemini 3.1 Pro. It's beating all three, on the same table.

It loses plenty of ground elsewhere, and that part matters just as much as the wins. But a model anyone can download under an MIT license, with no usage restrictions attached, coming out ahead of the lab everyone else measures themselves against, is worth pausing on before getting to what the rest of the numbers actually say.]]></description>
		
					<wfw:commentRss>https://firethering.com/glm-5-2-is-the-closest-an-open-model-has-come-to-claude/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Claude Mythos 5 Was Too Powerful to Ship. Anthropic Released Fable 5 Instead.</title>
		<link>https://firethering.com/claude-mythos-5-fable-5-anthropic-release/</link>
					<comments>https://firethering.com/claude-mythos-5-fable-5-anthropic-release/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Tue, 09 Jun 2026 21:44:46 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Anthropic]]></category>
		<category><![CDATA[Claude]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7349</guid>

					<description><![CDATA[Anthropic gave stripe early access to Fable 5 and set it loose on a 50 million line Ruby codebase. The migration that would have taken a full engineering team over two months got done in a day.

That's a real company's real codebase and a task with real consequences if it goes wrong. Anthropic leads with it because it's the kind of result that's hard to argue with &#038; because it sets up everything else they need to tell you about why this launch looks the way it does.

Because here's the thing. The model Anthropic actually built Claude Mythos 5, isn't what most people are getting today. What's going live for general use is Claude Fable 5. Same underlying model. Different version. The parts Anthropic decided were too dangerous for public release got a separate wrapper, a separate name, and a separate approval process controlled in part by the US government.]]></description>
		
					<wfw:commentRss>https://firethering.com/claude-mythos-5-fable-5-anthropic-release/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Ideogram 4 Topped the Open-Weight Leaderboard. Then We Read the License.</title>
		<link>https://firethering.com/ideogram-4-open-weight/</link>
					<comments>https://firethering.com/ideogram-4-open-weight/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Sat, 06 Jun 2026 14:31:24 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Image Model]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7281</guid>

					<description><![CDATA[Ideogram was founded by former Google Brain researchers who worked on Imagen, Google's own text-to-image system. When that team releases an open-weight model, you pay attention.

Ideogram 4 tops the open-weight design leaderboard by a margin that isn't close. Professional designers picked it first in blind typography tests nearly half the time. At 9.3B parameters it beats open models three times its size on text rendering.

Then we read the license.]]></description>
		
					<wfw:commentRss>https://firethering.com/ideogram-4-open-weight/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Google Built Gemma 4 12B Without Multimodal Encoders</title>
		<link>https://firethering.com/google-built-gemma-4-12b-without-multimodal-encoders/</link>
					<comments>https://firethering.com/google-built-gemma-4-12b-without-multimodal-encoders/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Thu, 04 Jun 2026 12:01:22 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Google]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7257</guid>

					<description><![CDATA[Every multimodal model you've used has the same basic system. Text goes in one way, images go through a vision encoder first, audio goes through an audio encoder first, and then everything gets handed off to the language model in a form it can work with. The encoders are load-bearing and you don't just remove them.Google actually removed them.Gemma 4 12B takes raw image patches and raw audio waveforms and projects them directly into the same embedding space as text tokens. There is no vision encoder or audio encoder. One decoder handling everything.]]></description>
		
					<wfw:commentRss>https://firethering.com/google-built-gemma-4-12b-without-multimodal-encoders/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>MiniMax M3 Shows What Happens When AI Stops Thinking in Turns</title>
		<link>https://firethering.com/minimax-m3-open-weight-model/</link>
					<comments>https://firethering.com/minimax-m3-open-weight-model/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Tue, 02 Jun 2026 19:25:39 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7236</guid>

					<description><![CDATA[Most models quit around submission 30 because they stop finding improvement and exit on their own. That's what happened when MiniMax ran a CUDA kernel optimization task against a field of frontier models. Every model except two called it done within the first 30 submissions.

M3's best result came on submission 145. After 24 hours. After multiple plateaus where the numbers stopped moving and a reasonable model would have concluded there was nothing left to find.

That's the thing MiniMax released yesterday. An AI model with a 1M token context window, native multimodality, and apparently a problem with knowing when to stop.]]></description>
		
					<wfw:commentRss>https://firethering.com/minimax-m3-open-weight-model/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>MiniCPM5-1B Shows Why the Small-Model Race Isn&#8217;t Over</title>
		<link>https://firethering.com/minicpm5-1b-small-model-reasoning/</link>
					<comments>https://firethering.com/minicpm5-1b-small-model-reasoning/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Sun, 31 May 2026 13:05:12 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7192</guid>

					<description><![CDATA[A 1B model scoring 40.42 on AIME 2025 should not be possible. AIME is the American Invitational Mathematics Examination, the kind of test that filters out most humans who attempt it. Qwen3-0.6B scores 16.25 on the same benchmark. LFM2.5-1.2B, a larger model, scores 31.88. MiniCPM5-1B, at roughly one billion parameters, beats both.

OpenBMB just dropped MiniCPM5-1B, the first model in their MiniCPM5 series, and it's built specifically for the scenarios like on-device deployment, resource-constrained environments, local inference on consumer hardware.

The AIME score is surprising. The telecom agent benchmark is even more surprising. And then there's the desktop pet. We'll get to that.]]></description>
		
					<wfw:commentRss>https://firethering.com/minicpm5-1b-small-model-reasoning/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
