back to top
HomeTechAI ModelsGPT-5.4 Is Outperforming Humans at Work. But the Real Story Is What...

GPT-5.4 Is Outperforming Humans at Work. But the Real Story Is What OpenAI Isn’t Telling You

- Advertisement -

OpenAI dropped their latest model yesterday and buried inside the benchmarks is a number that deserves more attention than it’s getting. On GDPval, a test that puts AI agents through real professional tasks across 44 actual occupations, GPT-5.4 matched or outperformed human professionals 83% of the time. The previous version sat at 71%. That’s not a small jump.

And this isn’t GPT writing emails or summarizing documents anymore. This version can move a mouse, click buttons, fill out forms, and work across applications the way a person sitting at a desk would. It scored 75% on OSWorld, a benchmark that tests exactly that. The average office worker scores 72.4%.

The model is already better at operating a computer than most people who use one for a living & 83% is just the beginning of what this release actually means.

The GDPval Number Nobody Is Talking About

The tasks GPT-5.4 was tested on are things real people get hired to do like sales presentations, accounting spreadsheets, urgent care schedules, manufacturing diagrams. The kind of output a junior hire would spend their first few months learning to produce.

The finance number is the one that stopped me. On investment banking modeling tasks, the Excel heavy work that junior analysts spend most of their first two years doing, GPT-5.4 scored 87.3%. GPT-5.2 was at 68.4%. Nearly 19 points in a single release.

To be fair, GDPval tests specific tasks, not entire careers. A job is more than its deliverables. But when the deliverables are exactly what junior roles are hired for, that distinction starts to feel thinner than it used to.

GPT-5.4 Is Not Just Answering Questions Anymore

Think about what a junior analyst actually does on a given day. They open a PDF, pull numbers from it, drop them into a spreadsheet, build a model, then paste results into a presentation. That’s not one task. That’s four applications, a lot of switching, and hours of work.

GPT-5.4 can now do that sequence without stopping. Not by generating text about it. By actually doing it across the applications, the same way a person would.

On an internal benchmark of spreadsheet modeling tasks specifically the kind a junior investment banking analyst would handle, it scored 87.3%. On presentations, human raters preferred GPT-5.4’s output 68% of the time over GPT-5.2’s. The quality gap between versions is noticeable enough that people can see it without being told which is which.

For developers building on top of this, GPT-5.4 also supports up to 1 million tokens of context. That means an agent can hold an entire project in memory, plan across it, execute steps, check its own work, and keep going without losing track of where it started.

That’s a different kind of tool than what most people picture when they think of ChatGPT.

The Parts OpenAI Won’t Tell You About GPT-5.4

The 83% number is real. But there are three things buried in this release that quietly put a ceiling on how far that number actually reaches in the real world.

The 1M context trap

GPT-5.4 technically supports a 1 million token context window. What OpenAI didn’t put in the headline is that anything beyond 272K tokens gets charged at 2x the normal rate. That’s not a feature, that’s a tax. If your workflow genuinely needs that full window, you’re paying double for the privilege. Treat 272K as the real limit and build around it.

The Cost Problem (Nobody is doing the math on)

To get that 83% human level performance you need GPT-5.4 Pro. That runs $30 per million input tokens and $180 per million output tokens. At that price point, for high volume repetitive work like data entry or customer support, the math doesn’t always favor the AI. A junior hire handling straightforward volume tasks can still be cheaper than running Pro at scale. The ROI just isn’t there yet for every use case.

The Security Ceiling

OpenAI’s own safety documentation flags GPT-5.4 as high cyber capability and wraps it in significant restrictions around anything that looks like offensive security work. The model won’t think creatively outside those guardrails. For white hat hackers and security researchers, the kind of outside the box thinking that makes someone genuinely good at that work is exactly what the model is prevented from doing.

The 83% Trade-Off: Power vs. Privacy

GPT-5.4 might be the most capable model available right now. But it arrives at a complicated moment for OpenAI.

On February 28th, OpenAI signed a deal with the Pentagon to deploy AI on classified military networks. The same day, ChatGPT uninstalls in the US jumped 295% according to Sensor Tower. One star reviews surged 775%. Claude hit number one on the US App Store for the first time, with downloads up 51% day over day.

People voted with their phones.

The contrast is hard to ignore. Anthropic got blacklisted as a national security risk for refusing to allow mass domestic surveillance and autonomous weapons without human oversight. OpenAI signed a deal. And now GPT-5.4, a model with native computer use capabilities and access to classified networks, is the most powerful version yet.

For professionals in 2026 the question isn’t just “does GPT-5.4 perform better.” It’s “where does my data go and what is it being used for.”

If that question matters to your work, alternatives exist. Claude is one. Local models like GLM-5 that run entirely on your own machine are another. The performance gap is closing faster than most people expected.

The 83% efficiency gain is real. So is the trade-off that comes with it.

The Bottom Line on GPT-5.4

GPT-5.4 is genuinely impressive. The benchmarks are real, the computer use capabilities are real, and the jump from GPT-5.2 is significant enough that it’s hard to dismiss.

But impressive and right for everyone are two different things. The pricing ceiling and the data privacy question deserve a place in your decision making alongside the 83% headline.

Use it if it fits your workflow. If it doesn’t, the alternatives are better than they’ve ever been.

Don’t miss any Tech Story

Subscribe To Firethering NewsLetter

You Can Unsubscribe Anytime! Read more in our privacy policy

LEAVE A REPLY

Please enter your comment!
Please enter your name here

YOU MAY ALSO LIKE
Open Source AI Assistants You Can Run Locally

5 Best Open Source AI Assistants You Can Run Locally

0
Somewhere between "just use ChatGPT" and "compile this from source," there's a category of AI tools that don't get talked about enough. Apps you download once, run on your own hardware, and never pay a monthly fee to use. No data leaving your machine. No API key. No usage limits that reset on the first of the month. The tools in this list aren't compromises. Some of them have millions of downloads. One was built by Mozilla. Another turns a 1B model into a desktop companion that reacts to your coding sessions. What they share is that after the initial setup, they answer only to you. If your current AI workflow depends on a subscription staying affordable and a company deciding your use case still matters next quarter, these are worth knowing about.
apple-sues-openai-stolen-secrets

OpenAI Paid $6.5 Billion to Build an iPhone Rival. Apple Says It Was Built...

0
Last year, OpenAI acquired io, Jony Ive's hardware startup, for $6.5 billion. The deal was widely read as OpenAI's clearest signal yet that it was serious about building a physical device, something that could sit in your pocket the way an iPhone does, powered by AI agents instead of apps. A direct challenge to Apple's most important product. On Friday, Apple filed a lawsuit suggesting that challenge was built on a foundation of stolen confidential information through what Apple describes as a coordinated operation directed from the top of OpenAI's hardware division, the same division now tasked with building the device meant to compete with Apple. Apple isn't just alleging that some employees walked out with files they shouldn't have taken. It's alleging that the people now running OpenAI's hardware ambitions actively ran a system to extract Apple's most guarded technical knowledge, and that the $6.5 billion acquisition sits on top of that foundation.
Anthropic Secretly Tracked Claude Code Users. Then Called It an Experiment

Anthropic Secretly Tracked Claude Code Users. Then Called It an “Experiment.”

0
There's a version of this story where Anthropic was trying to protect itself from large-scale model theft. There's another where one of the AI industry's biggest privacy advocates quietly crossed a line its own users never expected. What makes this headline important isn't just that hidden tracking code existed. It's that the company behind it was Anthropic. Just months ago, Anthropic publicly refused to let the Trump administration use Claude to surveil American users. The company defended that position in court, arguing that AI companies shouldn't become tools for government surveillance. That stance became part of Anthropic's identity. Then came a very different decision. In March, Anthropic quietly added hidden tracking markers to Claude Code that flagged users' timezones, proxy connections, and potential ties to Chinese AI labs. The code remained unnoticed until security researcher Thereallo discovered it last week. After the discovery went public, an Anthropic engineer confirmed it on X, described it as an "experiment" intended to combat account abuse and model distillation, and said the company had already planned to remove it. The tracker was taken down shortly afterward. The bigger question isn't whether Anthropic had a reason. It's whether a company that built its reputation on privacy can afford to hide surveillance from the very developers it asks to trust its tools.