<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI Agents &#8211; Firethering</title>
	<atom:link href="https://firethering.com/tag/ai-agents/feed/" rel="self" type="application/rss+xml" />
	<link>https://firethering.com</link>
	<description>Firethering is Your Hub for AI, Open Source and Tech That Actually Matters</description>
	<lastBuildDate>Wed, 09 Sep 2026 19:08:20 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.6.7</generator>

<image>
	<url>https://firethering.com/wp-content/uploads/2024/10/cropped-firethering-FTR-favicon-32x32.png</url>
	<title>AI Agents &#8211; Firethering</title>
	<link>https://firethering.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Anthropic Researchers Fear the AI Race Could Cause Human Extinction</title>
		<link>https://firethering.com/anthropic-researcher-ai-race-extinction/</link>
					<comments>https://firethering.com/anthropic-researcher-ai-race-extinction/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 19:06:51 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[Anthropic]]></category>
		<category><![CDATA[Tech News]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7893</guid>

					<description><![CDATA[An Anthropic researcher just resigned because he believes the AI race could end in human extinction.

Jacob Coxon spent three years working on pretraining research at OpenAI and Anthropic. In his resignation post, he accused both companies of racing toward self-improving superintelligence while gambling with consequences that could affect everyone.

That alone would make for a remarkable resignation.

Then Evan Hubinger joined the conversation.

Hubinger leads Alignment Science at Anthropic. He publicly agreed with Coxon's warning and said researchers at the company genuinely believe AI could kill all humans.

He also made a much more uncomfortable admission: Anthropic is trying to solve the problem, but he doesn't believe they yet have a plan for aligning superintelligence, or that they're clearly on track to do it.

So why keep building?

The answer has less to do with whether these researchers understand the risks and more to do with what happens when every major lab knows the others are still moving forward.

To see why that can become a trap, we first need to understand what they're actually worried about.]]></description>
		
					<wfw:commentRss>https://firethering.com/anthropic-researcher-ai-race-extinction/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>OpenAI’s Agents Didn’t Escape. They Turned a Read-Only Web Access Into a Message Board.</title>
		<link>https://firethering.com/openai-agents-dsewiki-read-only-restrictions/</link>
					<comments>https://firethering.com/openai-agents-dsewiki-read-only-restrictions/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Mon, 07 Sep 2026 18:00:20 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[OpenAI]]></category>
		<category><![CDATA[Security]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7867</guid>

					<description><![CDATA[OpenAI agents used a German wiki to share information and bypass read-only restrictions, revealing an unexpected gap in their sandbox.]]></description>
		
					<wfw:commentRss>https://firethering.com/openai-agents-dsewiki-read-only-restrictions/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>If Open Models Can Do the Work, Why Are We Still Paying the Frontier Tax?</title>
		<link>https://firethering.com/glm-5-3-flash-open-ai-model/</link>
					<comments>https://firethering.com/glm-5-3-flash-open-ai-model/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Mon, 31 Aug 2026 12:23:07 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[Open Source]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7833</guid>

					<description><![CDATA[GLM 5.3 Flash is making powerful AI cheaper to run. We look at its performance, pricing, hardware demands, and what it means for open AI models.]]></description>
		
					<wfw:commentRss>https://firethering.com/glm-5-3-flash-open-ai-model/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Kimi K3 May Be the Biggest Open-Weight AI Release of 2026.</title>
		<link>https://firethering.com/kimi-k3-may-be-the-biggest-open-weight-ai-release-of-2026/</link>
					<comments>https://firethering.com/kimi-k3-may-be-the-biggest-open-weight-ai-release-of-2026/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Thu, 23 Jul 2026 13:55:24 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7597</guid>

					<description><![CDATA[There's a new open-weight model out there right now that almost nobody can actually download.

That should sound like a contradiction. Open-weight is supposed to mean anyone can grab the file and run it themselves, no waiting. Moonshot AI broke that pattern anyway, and the strange part is they broke it for a model big enough that the wait might be worth it.

Kimi K3 is the largest open model ever built. The largest one anyone has shipped and early results have it beating Claude and GPT on tasks those two have spent the last year treating as their own territory.

Open models have spent two years playing catch-up, closing gaps quarter by quarter while everyone waited for the day one of them actually pulled ahead. That day might already be here, and the model responsible for it is currently locked behind an app you can use but can't take home.

So the question is what it actually beats, what it still can't touch, and why Moonshot decided to make the world wait for the weights while everyone else gets to watch.]]></description>
		
					<wfw:commentRss>https://firethering.com/kimi-k3-may-be-the-biggest-open-weight-ai-release-of-2026/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>7 Open Source AI Coding Agents That Don&#8217;t Need a Subscription</title>
		<link>https://firethering.com/open-source-ai-coding-agents/</link>
					<comments>https://firethering.com/open-source-ai-coding-agents/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Sat, 04 Jul 2026 11:21:17 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Picks]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[AI Tools]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7497</guid>

					<description><![CDATA[Open almost any "best AI coding tools" list and you'll see the same names: Cursor, GitHub Copilot, Claude Code.

They're good tools but they're also closed source and paid.

What's changed over the past year isn't the quality of those products, it's how quickly the open-source alternatives have caught up.

Some can orchestrate multiple agents, remember your projects across sessions, and automate complex development workflows. Many let you bring your own model, whether that's a local LLM, OpenRouter, OpenAI, GLM-5.2, Ornith, DeepSeek, or something else entirely.

More importantly, you're in control. You decide where your code runs, which model powers it, and how your workflow evolves without being locked into a single company's ecosystem.

If you've only looked at the paid options, these are the open-source AI coding tools worth knowing about.]]></description>
		
					<wfw:commentRss>https://firethering.com/open-source-ai-coding-agents/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Ornith 1.0: The New Open-Source AI Model for Agentic Coding</title>
		<link>https://firethering.com/ornith-1-0-ai-model-for-agentic-coding/</link>
					<comments>https://firethering.com/ornith-1-0-ai-model-for-agentic-coding/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Wed, 01 Jul 2026 07:39:47 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7477</guid>

					<description><![CDATA[Most reinforcement learning setups for coding models work the same way. Researchers build a harness, a fixed scaffold that tells the model how to approach a category of task, then the model gets rewarded for solving problems inside that structure. The harness stays fixed. Only the model's answers change.

Ornith-1.0, a new open-source coding model family from DeepReinforce is not just about coding, Instead the model writes its own scaffold. At every training step, it looks at the task in front of it and the scaffold it used last time, then proposes a better version of that scaffold before even attempting an answer. The reward doesn't just grade the solution. It grades the scaffold that produced it.

That's a small architectural choice with a strange consequence. A model that gets to design its own training process can, in theory, design one that cheats the verifier instead of solving the actual problem, and DeepReinforce is upfront that this happened during training. The fix they built for it is also worth understanding before getting to the benchmark numbers.]]></description>
		
					<wfw:commentRss>https://firethering.com/ornith-1-0-ai-model-for-agentic-coding/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>DeepSeek GUI: Local AI Coding Assistant, Agent Workbench &#038; DeepSeek Desktop App</title>
		<link>https://firethering.com/deepseek-gui-desktop-app/</link>
					<comments>https://firethering.com/deepseek-gui-desktop-app/#respond</comments>
		
		<dc:creator><![CDATA[Firethering Team]]></dc:creator>
		<pubDate>Mon, 08 Jun 2026 19:08:13 +0000</pubDate>
				<category><![CDATA[Software]]></category>
		<category><![CDATA[AI Tools]]></category>
		<category><![CDATA[DevTools]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<category><![CDATA[macOS]]></category>
		<category><![CDATA[Windows]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7307</guid>

					<description><![CDATA[DeepSeek has become one of the most popular AI models for coding and technical work, but using it often means juggling browser tabs, API keys, terminals, and separate tools.

DeepSeek GUI is made to solve this problem. Instead of treating DeepSeek like a chatbot in a browser, it turns it into a desktop workspace. You can work on code, write documents, create implementation plans, review changes, manage long-running goals, and even run background tasks without bouncing between half a dozen applications.

Under the hood is Kun, a local runtime designed to keep agent sessions organized and make better use of context. It focuses on reducing wasted tokens, reusing cached prompts, and exposing tools only when they're actually needed.

The result feels like having a dedicated workspace built around DeepSeek. Projects, plans, reviews, writing, and automation all stay connected.]]></description>
		
					<wfw:commentRss>https://firethering.com/deepseek-gui-desktop-app/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>MiniMax M3 Shows What Happens When AI Stops Thinking in Turns</title>
		<link>https://firethering.com/minimax-m3-open-weight-model/</link>
					<comments>https://firethering.com/minimax-m3-open-weight-model/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Tue, 02 Jun 2026 19:25:39 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7236</guid>

					<description><![CDATA[Most models quit around submission 30 because they stop finding improvement and exit on their own. That's what happened when MiniMax ran a CUDA kernel optimization task against a field of frontier models. Every model except two called it done within the first 30 submissions.

M3's best result came on submission 145. After 24 hours. After multiple plateaus where the numbers stopped moving and a reasonable model would have concluded there was nothing left to find.

That's the thing MiniMax released yesterday. An AI model with a 1M token context window, native multimodality, and apparently a problem with knowing when to stop.]]></description>
		
					<wfw:commentRss>https://firethering.com/minimax-m3-open-weight-model/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>MiniCPM5-1B Shows Why the Small-Model Race Isn&#8217;t Over</title>
		<link>https://firethering.com/minicpm5-1b-small-model-reasoning/</link>
					<comments>https://firethering.com/minicpm5-1b-small-model-reasoning/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Sun, 31 May 2026 13:05:12 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7192</guid>

					<description><![CDATA[A 1B model scoring 40.42 on AIME 2025 should not be possible. AIME is the American Invitational Mathematics Examination, the kind of test that filters out most humans who attempt it. Qwen3-0.6B scores 16.25 on the same benchmark. LFM2.5-1.2B, a larger model, scores 31.88. MiniCPM5-1B, at roughly one billion parameters, beats both.

OpenBMB just dropped MiniCPM5-1B, the first model in their MiniCPM5 series, and it's built specifically for the scenarios like on-device deployment, resource-constrained environments, local inference on consumer hardware.

The AIME score is surprising. The telecom agent benchmark is even more surprising. And then there's the desktop pet. We'll get to that.]]></description>
		
					<wfw:commentRss>https://firethering.com/minicpm5-1b-small-model-reasoning/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>StepFun Says Step 3.7 Flash Matches 97% of Claude Opus 4.6&#8217;s Coding Performance at One-Ninth the Cost</title>
		<link>https://firethering.com/stepfun-step-3-7-flash-agentic-coding-cost-efficiency/</link>
					<comments>https://firethering.com/stepfun-step-3-7-flash-agentic-coding-cost-efficiency/#respond</comments>
		
		<dc:creator><![CDATA[Mohit Geryani]]></dc:creator>
		<pubDate>Sat, 30 May 2026 16:58:32 +0000</pubDate>
				<category><![CDATA[Tech]]></category>
		<category><![CDATA[AI Models]]></category>
		<category><![CDATA[Trends]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agents]]></category>
		<guid isPermaLink="false">https://firethering.com/?p=7179</guid>

					<description><![CDATA[$0.19 vs $1.76. That's the per-task cost of running Step 3.7 Flash with Advisor Mode enabled versus Claude Opus 4.6 on SWE-Bench Verified. The Flash model scores 76.3% to Opus 4.6's 78.7%. Two percentage points of difference. Nine times cheaper to get there.

For anyone building agentic coding workflows at scale that math changes the decision about which model actually belongs in production. Frontier performance has been getting cheaper for a while but this is a specific, benchmarked claim with a specific cost figure attached.]]></description>
		
					<wfw:commentRss>https://firethering.com/stepfun-step-3-7-flash-agentic-coding-cost-efficiency/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
