OpenAI Is Racing to Catch Up in the AI Coding Revolution It Started
Sam Altman folds his legs into a pretzel in an office chair, his gaze drifting upward to the high ceiling of OpenAI’s sleek new headquarters. To be fair, the glass-and-bleached-wood temple in San Francisco’s Mission Bay is built for exactly this kind of big-picture contemplation. Behind the front desk, a display rack holds pamphlets framing the “Eras of AI” as incremental steps toward something like technological enlightenment. Posters lining the staircase celebrate AI’s landmark wins: one marks the day thousands of viewers tuned into a livestream to watch an AI outplay a top professional esports team at Dota 2. Down the hallways, researchers wander past in the company’s unofficial unofficial swag, one of the most popular tees reading “Good research takes time.” Ideally, not too much of it.
Altman and I are seated in a sprawling glass conference room when I ask him about the AI coding boom— and why OpenAI, the company that kicked off the generative AI era, isn’t leading it. Today, millions of software developers offload day-to-day programming work to AI, forcing Silicon Valley’s core workforce to grapple with job automation for the first time. Coding agents have emerged as one of the rare enterprise AI categories that customers are willing to spend big on, a milestone that by all rights should deserve its own spot on OpenAI’s staircase of triumphs. But right now, the biggest name in the space belongs to someone else.
That name is Anthropic, a smaller rival founded by former OpenAI employees, whose standalone coding agent Claude Code has exploded in popularity. In February, Anthropic disclosed that Claude Code makes up nearly 20% of its total business, pushing the company’s annualized revenue past $2.5 billion. By comparison, OpenAI’s own coding offering, Codex, brought in just over $1 billion in annualized revenue by the end of January, according to a person with direct knowledge of the company’s internal numbers. So why is OpenAI playing catch-up here?
“First-mover advantage counts for a lot,” Altman says finally. “We had that with ChatGPT.” But he argues the moment is now ripe for OpenAI to go all in on coding, noting that the company’s flagship models are now powerful enough to support best-in-class coding agents— after the company poured billions into training them to reach this point. “This is going to be a massive business, just in terms of raw economic value, not to mention all the general-purpose work that widespread accessible coding unlocks,” Altman says. “I don’t throw this label around lightly, but this is one of those rare multitrillion-dollar markets.” What’s more, he calls Codex “probably the most likely path” to building artificial general intelligence, or AGI— which OpenAI defines as a system that can outperform humans at nearly all economically valuable work.
But for all Altman’s confident declarations from his relaxed cross-legged pose, the story inside OpenAI over the past few years is far messier. To get an inside look, I spoke with more than 30 people: current OpenAI leaders and employees who spoke with the company’s approval, plus other current and former staff who spoke on condition of anonymity to discuss internal operations of the private company. Their accounts paint a picture of OpenAI in a position it has rarely occupied: it’s racing to catch up to a competitor.
It all goes back to 2021, when Altman and other OpenAI leaders invited WIRED journalist Steven Levy to the company’s original office in San Francisco’s Mission District to preview a new project. The tool was an offshoot of OpenAI’s GPT-3 model, trained on billions of lines of open source code pulled from GitHub. In the demo, executives showed how Codex could take plain English commands and turn them into working, simple code snippets.
“It can actually act in the computer world on your behalf,” Greg Brockman, OpenAI’s president and co-founder, said at the time. “You actually have a system that can carry out your commands.” Even back then, OpenAI’s researchers saw clearly that Codex would be central to building a powerful “super assistant” for general use.
At that point, Altman and Brockman were consumed by negotiations and meetings with Microsoft, OpenAI’s largest investor. The tech giant was already planning to use Codex to power one of its first commercial AI products: GitHub Copilot, a code completion tool that integrated directly into a developer’s existing workflow. Back at this early stage, Codex “couldn’t do much more than autocomplete,” one early OpenAI employee told me, but Microsoft executives framed it as a preview of the coming AI future. When GitHub Copilot launched publicly in June 2022, it attracted hundreds of thousands of users within just a few months.
OpenAI’s original Codex team was soon split up and reassigned to other projects. The company planned to integrate coding capabilities directly into its future general-purpose models, so there was no need for a standalone dedicated effort, the early employee explained. Some engineers moved to work on DALL-E 2, OpenAI’s viral image generator. Others shifted to training GPT-4, which was widely viewed inside the company as the fastest path to AGI.
Then ChatGPT launched in November 2022, and racked up more than 100 million users in just two months. Every other non-ChatGPT project at OpenAI was put on hold. For years after that, OpenAI had no dedicated team working on a standalone AI coding product. Coding fell outside the company’s new, laser focus on consumer products, one former Codex team member says. It also “felt like the sector was ‘covered’ by GitHub Copilot,” they added. OpenAI would just supply updated models to power Microsoft’s tool— this was Microsoft’s lane, not OpenAI’s.
OpenAI spent most of 2023 and 2024 investing instead in multimodal AI models and agents, built to process text, images, video, and audio, and control a computer’s cursor and keyboard just like a human would. This aligned with the prevailing industry consensus at the time: after Midjourney went viral for its AI image models, most leaders believed large language models needed to see and hear the world to develop true intelligence.
Anthropic took a different tack. The company also experimented with chatbots and multimodal models, but it recognized the promise of standalone coding agents far earlier than OpenAI. On a recent podcast, Brockman commended Anthropic for staying “focused very hard on coding” from its early days. He noted that Anthropic trained its models not just on difficult academic coding competition problems, but also on real-world messy code from public repositories. “That was a lesson that we were delayed on,” Brockman admitted.
In early 2024, Anthropic trained Claude 3.5 Sonnet on troves of that real-world messy code. When the model launched that June, users were blown away by its coding capabilities. That momentum especially lifted Cursor, a startup founded by a team of twentysomethings that lets developers build software by requesting changes in plain English. After Cursor integrated Anthropic’s new model, its usage skyrocketed, according to a person close to the company. Within months, Anthropic began internal testing of its own standalone offering: Claude Code.
As Cursor grew in popularity, OpenAI approached the startup about an acquisition. The founders turned down the offer before talks ever reached an advanced stage, people close to Cursor told me, because they wanted to stay independent to capitalize on coding market’s growing potential.
At the time, OpenAI was training its first dedicated reasoning model, o1, which works through complex problems step-by-step before delivering an answer. At launch, OpenAI noted the model “excels at accurately generating and debugging complex code.” Andrey Mishchenko, OpenAI’s research lead for Codex, says a key reason AI models have improved so much at coding is that it is a uniquely verifiable task: code either runs or it doesn’t, giving models a clear signal when they make a mistake. OpenAI used that feedback loop to train o1 on increasingly difficult coding challenges. “Without the ability to crawl around a code base, implement changes, and test their own work— these are all under the umbrella of reasoning— coding agents would not be anywhere near as capable as they are today,” he says.
By December 2024, several small teams inside OpenAI began focusing explicitly on AI coding agents. One was led by Mishchenko and Thibault Sottiaux, a former Google DeepMind researcher who now serves as OpenAI’s head of Codex. Initially, the team viewed coding agents as a tool to speed up OpenAI’s own internal research, automating grunt work like managing training runs and monitoring GPU clusters. A separate effort was led by Alexander Embiricos, who previously worked on OpenAI’s multimodal agents and is now Codex’s product lead. Embiricos built an internal demo called Jam that spread rapidly across the company.
Unlike tools that control a computer via cursor and keyboard, Jam had direct programmatic access to the computer’s command line. Where the 2021 Codex demo only output code for a human to run, Embiricos’ version could execute the code itself. He recalls being awestruck watching a dashboard tracking Jam’s actions update over and over on his laptop.
“For a while, I had been thinking that multimodal interaction might be how we achieve our mission— like we would just be screen-sharing with AI all day,” Embiricos says. “Then it became super clear: Maybe giving models programmatic access to a computer is how we're going to get there.”
It took months for the separate internal projects to merge into a unified effort. By early 2025, OpenAI finished training o3, a model even more optimized for coding than o1, and finally had the foundation to build a full-fledged standalone AI coding product. But Claude Code was already poised to launch publicly.
Before Claude Code’s rollout— first a limited research preview in February 2025, then a full general release that May— the state of the art was so-called “vibe coding”: human developers steered projects, while AI filled in small details along the way. But Anthropic’s new product, just like OpenAI’s internal Jam demo, works directly from a developer’s command line, meaning it has access to all of a coder’s files and applications. This was no longer incremental help: developers could offload entire projects to the AI agent.
OpenAI scrambled to launch a competing product. Sottiaux says he formed a dedicated “sprint team” in March 2025, with a mandate to merge OpenAI’s internal groups and ship a product in just a few weeks. At the same time, Altman pursued a second acquisition to jumpstart OpenAI’s effort: buying AI coding startup Windsurf for $3 billion. OpenAI leadership expected the deal to deliver an established product, an experienced team, and an immediate base of enterprise customers.
But the Windsurf deal stalled for months. According to The Wall Street Journal, the holdup stemmed from Microsoft, OpenAI’s core business partner, which demanded access to Windsurf’s intellectual property. Microsoft has used OpenAI’s models to power GitHub Copilot since 2021, and the product has become a highlight of Microsoft’s quarterly earnings calls. But as new agentic coding tools like Cursor, Windsurf, and Claude Code gained traction, GitHub Copilot began to feel outdated. OpenAI launching its own standalone coding product would directly compete with Microsoft’s offering.
The Windsurf negotiations unfolded during a particularly fraught period in OpenAI and Microsoft’s partnership: the two companies were renegotiating their overall deal, with OpenAI pushing to loosen Microsoft’s control over its products and cloud computing resources. The Windsurf deal became a casualty of the process, and the acquisition fell apart by July. Google ultimately hired Windsurf’s founders, while the rest of the team was acquired by Cognition, another AI coding startup.
“I would have loved to get that done,” Altman says. “You can’t control every deal.” While he had hoped the acquisition “would have accelerated us somewhat,” Altman says he’s impressed with the Codex team’s trajectory. Sottiaux and Embiricos kept building and shipping updates throughout the acquisition negotiations, and by August, Altman says OpenAI hit full speed.
Greg Brockman’s favorite way to test AI performance is a custom computer game he built called the Reverse Turing Test. He hand-coded it years ago, and now challenges AI agents to build a working version from scratch. The rules are simple: Two humans on separate computers each see two chat windows. One connects to the other human, one connects to an AI. The goal is to guess which window is connected to the AI, while tricking your opponent into thinking you are the AI.
For most of 2025, Brockman says OpenAI’s best model took hours to build a working version of the game, and required constant explicit guidance from humans along the way. But by December, Codex was able to create a fully functional version from a single well-crafted prompt, powered by OpenAI’s new GPT-5.2 model.
Brockman wasn’t the only one noticing the leap forward. Developers around the world began noting that AI coding agents had suddenly become dramatically more capable. The public conversation, which initially centered on Claude Code, spread far beyond Silicon Valley and became a mainstream news story. Even casual users with no formal coding experience began building custom software projects for personal use.
This spike in usage was no accident. Both Anthropic and OpenAI spent heavily to acquire new customers for their coding agents in this period. Multiple developers told WIRED that their $200-per-month subscription plans for Codex and Claude Code often deliver well over $1,000 worth of usage. These generous rate limits are a deliberate strategy to get developers adopting the tools in their workplaces, where both companies can eventually charge premium enterprise pricing for heavy usage.
Back in September 2025, Codex accounted for just 5 percent of the total usage of Claude Code, according to people with direct knowledge of internal metrics. By January 2026, Codex’s user base had grown to nearly 40 percent of Claude Code’s, the sources say.
George Pickett, a developer who has worked at tech startups for 10 years, recently started organizing in-person meetups for Codex users. “I think it's clear we're going to replace white-collar work with agents,” Pickett says. “Societally, who fucking knows what this means. It’s going to be disruptive, but I’m pretty optimistic about what’s happening.”
Simon Last, cofounder of the $11 billion AI productivity startup Notion, says he and his top engineering team switched to Codex around the launch of GPT-5.2, largely due to its reliability. “I found that Claude Code just lies to me,” Last says. “It says it's working, but it actually isn't.”
Katy Shi, a research lead on the Codex team, says that while some users describe the agent’s default tone as “dry bread,” many have come to appreciate its less sycophantic style. “A lot of engineering work is about being able to take critical feedback without interpreting it as mean,” Shi says.
Dozens of major enterprises have also signed enterprise deals for Codex. “The fact that ChatGPT is synonymous with AI gives us a massive advantage in the B2B market,” says Fidji Simo, OpenAI’s CEO of applications. “Companies want to use technologies their workers are already familiar with.” OpenAI’s go-to-market strategy for Codex centers on packaging it alongside ChatGPT and other existing OpenAI products, Simo says.
Jeetu Patel, Cisco’s president and chief product officer, says he has told employees not to worry about the cost of using Codex, because they need to get comfortable with the tool to stay relevant. When employees ask if “they’re going to lose their job because they’re using these tools,” Patel says, “what we have to tell our people is no, but I guarantee you'll lose your job if you don't use them, because you won't be relevant. So you're going to be out.”
Today, panic over AI coding agents has spread far beyond Silicon Valley. The Wall Street Journal credited Claude Code with triggering a $1 trillion tech stock sell-off last month, as investors feared software development jobs would soon become largely obsolete. Weeks later, IBM posted its worst single-day stock drop in 25 years after Anthropic announced Claude Code can be used to modernize legacy COBOL systems that run on most IBM mainframes. OpenAI has worked aggressively to position Codex at the center of the national conversation, spending millions of dollars on a Super Bowl ad focused on Codex, rather than ChatGPT.
Inside OpenAI’s Mission Bay headquarters, no one needs to be sold on Codex’s value. Many OpenAI engineers I spoke with say they rarely type out full lines of code anymore. They spend most of their days prompting Codex to do the work for them, and sometimes even gather to build projects together with the tool.
During my visit, I sat in on a company-wide Codex hackathon: roughly 100 engineers crammed into a large conference room, with four hours to build the best demo using Codex. A senior OpenAI leader stood at the front of the room, announcing team names into a microphone. Team representatives walked nervously to the podium to give short presentations on their AI-built projects, and winners walked away with Patagonia backpacks.
Most of the projects were both built with Codex and designed to help other engineers use Codex better. One team built a tool that summarizes Slack messages into polished weekly reports. Another built an AI-generated Wikipedia-style guide to internal OpenAI services. What would have taken days or weeks to build just a few years ago can now be completed in a single afternoon.
On my way out of the building, I ran into Kevin Weil, the former Instagram executive who now leads OpenAI for Science, the company’s new unit building AI tools for researchers. He told me Codex was running several projects for him overnight, and he would check on their progress the next morning. That routine has become standard for Weil, and hundreds of other OpenAI employees. One of the company’s core 2026 goals is to build an automated AI intern that conducts research on… AI.
Simo tells me OpenAI wants Codex to eventually power task-completion features across ChatGPT and all of the company’s products, not just for programming. Altman says he’d love to release a general-purpose version of Codex for all types of work, but he worries about unaddressed safety risks. In late January, he says, a non-technical friend asked him to set up OpenClaw, a viral open-source AI coding agent. Altman told me he declined, because it was “clearly not a good idea yet” — OpenClaw can accidentally delete a user’s important files. A few weeks after Altman shared that anecdote with me, OpenAI announced it had hired OpenClaw’s creator.
Many developers I spoke with say the race between Codex and Claude Code is tighter than it
OpenAI Is Racing to Catch Up in the AI Coding Revolution It Started