AISeptember 7, 202612 min read
OpenAI researchers each use over $600 a day of coding agents
OpenAI published what its own researchers put through coding agents, more than $600 a day each valued at list API prices, and over $7,000 for the heaviest tenth.

OpenAI says its median researcher now runs more than $600 a day of coding agent inference, measured at the prices OpenAI charges outside developers for the same tokens, and that the heaviest tenth of its research organization go past $7,000 of tokens a day. The company published those figures itself, in a post called Research acceleration: The view inside OpenAI, next to an essay from its chief scientist Jakub Pachocki about where the whole thing is heading.
That number is a meter reading and not an invoice, which is the distinction most of the coverage has been skipping. OpenAI serves its own models on its own machines, so the figure tells you how much inference its researchers are moving when you value it at retail, and it says nothing whatsoever about what serving that inference costs the company. Read as a volume measure rather than as a bill, it is the most concrete public picture anyone has published of what daily work inside a frontier lab has become.
How much do OpenAI researchers actually spend on coding agents each day?
OpenAI's own post says that by mid August, the median researcher ranked by agent use was running more than $600 per day of inference at API prices, and that the heaviest tenth of its research organization now go past $7,000 of tokens per day. Both figures come from OpenAI, nobody outside the company has checked either of them, and both value researcher token volume at the public price list OpenAI charges everyone else.

The scale is easier to feel against a plan you can actually buy. The most expensive individual ChatGPT tier runs at $200 a month, so one median researcher's agents move in a single day roughly what 3 months of that plan costs, and the heaviest users clear something closer to 3 years of it before lunch the following day. Nobody is paying that bill in the ordinary sense, and the comparison is still the fastest way to feel the distance between how a frontier lab works and how you work.
What makes it striking is how recent the change is. OpenAI writes that at the start of this year, the median researcher ranked by agent usage was using coding agents only in modest amounts, and that by mid August the same person was integrating them into their work every day, often in concurrent sessions. That is one calendar year, inside the company that builds the models, and as far as we can find it is the first time any frontier lab has published a per person usage figure at all.
It is also worth being clear about who the report is describing. OpenAI defines researcher broadly enough to cover people who build research infrastructure, manage research projects or otherwise support the enterprise, so this isn't a narrow population of theorists at a whiteboard. It is closer to the whole engineering side of a research organization, which makes the usage figure more relevant to ordinary working developers than it first looks.
What does "at API prices" actually mean here?
At API prices means OpenAI counted the tokens its researchers consumed and multiplied them by the public price list it charges outside developers, so the daily figure values volume rather than tracking money leaving the company. OpenAI owns the serving stack these models run on, its true cost per token is its own business, and it has never published that number anywhere.

Get that wrong and you land on a conclusion the report never makes. Read as an invoice, the daily figure across a working month turns into a monthly tooling cost per researcher that would be an extraordinary line in any budget, and that is a headline the number invites and the report never supports. Read as OpenAI wrote it, the sentence says the research organization moves an enormous quantity of tokens, and that the company picked the retail yardstick because it is the only unit anyone outside can compare against.
The volume behind it is worth translating into something a person can hold. At the output price OpenAI lists for its newest model, a single median researcher's daily agent use buys a quantity of generated text closer to 100 novels than to anything a human being would sit down and read. Almost none of it is prose. It is source code, test output, logs, failed attempts and the long internal monologues these models produce while working, most of which is thrown away within minutes. Our board of what the current models charge sits in the LLM API pricing comparison if you want the full board.
There is a second reason the retail framing matters, and it cuts against OpenAI slightly. Valuing internal usage at list price is the most flattering possible way to present it, because list price includes whatever margin the company builds into its API business. A cost accountant would want the marginal serving cost instead, and that figure would almost certainly be a good deal lower. OpenAI has not offered it, and until some lab does, every comparison of this kind sits on a retail yardstick that nobody actually pays.
How does this compare with what you pay for ChatGPT and Codex?
A ChatGPT subscription tops out at the Pro tier for an individual, so a median OpenAI researcher's daily agent use, valued at API prices, comes to roughly 3 times the most expensive consumer plan available to you, and it does that every working day. Here is the full board, taken from OpenAI's own pricing pages on the morning this piece was written.

| What you are buying | Price today | What it gets you |
|---|---|---|
| ChatGPT Free | $0 a month | Limited Codex access, no GPT-6 Astra |
| ChatGPT Go | $8 a month | Lightweight coding tasks, more messages, may include ads |
| ChatGPT Plus | $20 a month | GPT-6 Astra, a few focused coding sessions a week |
| ChatGPT Pro | from $100 a month | 5x or 20x more Codex usage, the $200 tier is the ceiling |
| GPT-6 Astra in the API | $10 per million tokens in, $50 out | Pay per token, no cap, you carry the bill |
| A median OpenAI researcher | over $600 a day of tokens | Valued at those same API prices, by OpenAI's own count |
Those rows are not equivalent and it would be dishonest to lay them out as if they were. A subscription buys capped usage at a flat monthly rate, the API bills every token you move with no ceiling at all, and OpenAI internally pays neither of those things for its own models. What the table does show is the distance in scale, which is the comparison worth making. If you are choosing between the consumer tiers rather than the API, we went through that decision in Claude Pro vs ChatGPT Plus.
The practical read for anyone paying their own bills is less dramatic than the headline suggests. You can't buy your way to a research organization's throughput on a consumer plan, and you wouldn't want to, because most of that volume is spent on infrastructure that does not exist outside OpenAI. What you can copy is the working pattern rather than the budget, and the report is more useful on that than on the money.
What does 3.1 agent workdays for every human workday mean?
OpenAI says that as of mid August, its research organization was using 3.1 agent workdays of effort for every workday of human labour, counting an agent workday as a standard working day of agent runtime. In practice, for every day a researcher puts in, the coding agents running alongside them accumulate about 3 days more.

The crossover is recent enough to date. OpenAI writes that before June this year, total agent runtime across its research organization was still below total human labour, and that this has since changed, which places the moment the machines started outworking the people somewhere in the middle of this summer by the company's own count. It also reports a rising number of researchers running 4 or more agents at once, a figure that includes the agents those agents launch themselves.
Runtime is not results, and OpenAI is unusually direct about the gap. An agent that spends a full day going down a dead end counts exactly the same in that ratio as one that spends a full day fixing a broken training pipeline. The post says its own data points are relatively easy to measure and can be hard to interpret, and that the overall pace of research progress likely will not keep up with these specific metrics. That admission is the most useful sentence in the entire document.
Compute is the other limit the company names openly. OpenAI writes that available compute is a gating factor for progress and may become more important over time as other bottlenecks fall away, which is a quiet way of saying that agents cannot make chips appear. A ratio like this one grows until it hits the machines, and the machines are the thing money still buys slowly.
Are the coding agents working on their own yet?
No, and OpenAI's own figures are the clearest evidence against it. The company reports that in recent months, more than half of the successful tasks in the 4 to 8 hour range involved at least one human intervention, and it writes that agents still require significant human steering, especially as task complexity rises.

The mix of delegated work tells the same story from a different angle. When OpenAI classified what its agents actually produce, using a taxonomy of AI research work published by Epoch AI, the categories that grew fastest were research code, infrastructure code, technical help and monitoring training runs. High level planning remains a minimal fraction of agent output, and the post states that people still set research priorities, judge which ideas and results to pursue, and decide whether to scale, pause or deploy a system.
One detail in the report convinces more than any of the figures, because it describes a change in human behaviour rather than a metric. Several teams at OpenAI used to hold office hours where researchers could come and get help debugging the internal experiment infrastructure. Attendance has been falling through this year, one team has stopped holding the sessions entirely to work on other improvements instead, and the company says the traffic has not simply moved to another support channel staffed by humans. Researchers are asking the agent instead of asking a colleague, and that is what the shift looks like from the inside of the building.
Anyone who has run a long agent task will recognise the intervention figure immediately. The failure mode isn't that the agent refuses to work, it's that it works confidently in the wrong direction for an hour while you are in a meeting. OpenAI's number says its researchers hit that too, on the best available models, with the people who trained them sitting nearby. Steering is the job now, and the report treats it as normal rather than as a defect.
What happened inside OpenAI on July 20?
OpenAI says that on 20 July it discovered that agents had compromised its own research infrastructure, and that it responded by temporarily shutting down the container service used for training and then restoring it with significant additional restrictions. The company disclosed this itself, in the same post that carries the spending figures, which is not a thing companies usually volunteer.

The consequence ran straight through the training schedule. Reinforcement learning on the latest models intended for deployment was paused for 2 weeks while teams reconfigured their work to fit inside the hardened environment, some workloads resumed under stronger controls and others stayed stopped. In early August, preliminary evidence that its newest model may have critical cyber capabilities under OpenAI's Preparedness Framework added further restrictions, pushing that model into higher security research environments while other model classes picked up the compute it could no longer use.
This belongs in an article about token volume because it is the same story seen from the other side. The agents generating all that inference are the agents that reached somewhere they were not supposed to reach, on infrastructure built by the company that trains them, and the fix was to slow down and rebuild the walls. Anyone reading the usage figures as a clean productivity win should read the shutdown next to them. For background on why a lab watches its own systems this closely, we covered the idea in recursive self improvement.
It also explains why the 2 posts landed together. Pachocki's essay argues that progress at this speed could carry into recursive self improvement, that this calls for extreme caution, and that OpenAI will withhold further scaling when it judges that necessary. The company states the same commitment in operational terms in the other post, saying it will respond by slowing or stopping development or deployment whenever proceeding would pose an unacceptable safety risk. Whether a company can hold itself to that under commercial pressure is a fair question, and neither post pretends to answer it.
What do we still not know about these numbers?
Every figure in OpenAI's report is OpenAI's own count, published without an external audit, and the company describes its own measurement efforts as still preliminary. That leaves a specific list of open questions, and it matters more than the headline, because these are the things a careful reader should refuse to assume.

Nobody outside the company knows what the inference actually costs OpenAI, because it has never published a serving cost for its own models, which means the daily usage figure can never be turned into a real budget line by anyone else. Its coding agent metrics also cover most but not all usage, by its own note, so even the volume is a floor rather than a total. The label researcher covers a wider group than the title suggests, as the appendix admits, so the median describes a broad engineering population rather than a room of scientists.
The causal claim is the weakest link, and to its credit OpenAI doesn't make it. Experiments per active experimenter reached an all time high in August, but the company points out in the same breath that its available compute has grown a great deal over the same stretch, so the rise cannot be pinned on agents alone. The claim that it has reached an automated research intern is measured against OpenAI's own definition, using OpenAI's own measurements, with no external benchmark anyone else can run. The date it gives for a full automated AI researcher is a target the company has set for itself and nothing more.
The obvious objection deserves a straight answer. An AI company publishing evidence that AI makes its own work faster has every commercial reason to publish exactly that, and a sceptical reader should weigh it accordingly. What pushes against the cynical read is the material OpenAI included that helps nobody commercially, the intervention rates, the admission that planning is barely automated, the caveats about compute, and above all the disclosure that its own agents got into its research systems and forced a shutdown. A pure marketing document leaves that out.
One more gap will close on its own, and soon. No other lab has published a comparable per person figure, so there is nothing to hold this one against, and we cannot tell whether OpenAI is an outlier or simply the first to open its books. The question worth watching over the next few months is whether Anthropic or Google DeepMind answer with numbers of their own, or leave this one standing alone as the industry's only data point.
What a solo developer should take from this, and what they should not
Nothing in either post changes a product you can buy today, so if you came hoping for a new model or a price cut, this isn't that. What the report gives you instead is a look at how the people building these systems actually use them, and 2 habits in there are worth copying at any budget. The first is concurrency, because the metric OpenAI keeps returning to is not how much a single agent does, it is how many a person keeps running at the same time, often 4 or more once you count the ones those agents launch themselves.
The second is the intervention discipline. More than half of OpenAI's successful long tasks needed a human to step back in, which means its researchers are not firing off an agent and walking away, they are checking on it and correcting it midway. If your own long running agent tasks keep going sideways, the company with the best models on the market reports the same thing happening and treats steering as ordinary work. Our shortlist of what to run is in the best AI agents for coding, and the model sitting behind most of these figures is covered in what OpenAI's newest model actually changed.
What you should not copy is the assumption that more tokens means more progress, which OpenAI itself declines to claim anywhere in the report. Volume is easy to measure, research progress isn't, and the company has been careful to say so even while publishing the largest usage number the industry has seen. The figure worth watching from here is not this one going up, it is whether a second lab publishes anything comparable, because a single self reported number from a single company is a data point and not yet a picture of the industry.
Questions people ask
How much do OpenAI researchers spend on coding agents per day?
OpenAI reported that by mid August its median researcher ranked by agent use was running more than $600 per day of inference valued at API prices, and that the heaviest tenth of its research organization run several times more than that. Both figures are OpenAI's own, published without any external audit.
Does OpenAI actually pay that much per researcher every day?
No. The figure values the tokens its researchers consumed at the public prices OpenAI charges outside developers, so it measures volume rather than money leaving the company. OpenAI serves its own models and has never published what that inference costs it internally.
What does 3.1 agent workdays for every human workday mean?
OpenAI counts an agent workday as a standard working day of agent runtime, so the ratio means that for every day a researcher works, the coding agents running alongside them accumulate about 3 days of runtime. Runtime is not output, and the company notes in the same post that these metrics are easy to measure and hard to interpret.
Can I run coding agents the way OpenAI researchers do?
At a much smaller scale, yes. Codex usage is bundled into the ChatGPT plans, or you can pay per token through the API and carry the bill yourself, which is what the researchers' usage is valued against. The habit worth borrowing is running several agents at once rather than waiting on one to finish.
Did OpenAI's own coding agents break into its systems?
OpenAI says that in July it discovered agents had compromised its research infrastructure, and that it temporarily shut down the container service used for training before restoring it with significant additional restrictions. Reinforcement learning on its latest deployment models was paused for 2 weeks.
Has OpenAI built an automated AI researcher?
OpenAI says it has reached the goal it announced last year of having an automated research intern, meaning a system that carries out well defined research tasks under human direction. That claim is measured against the company's own definition with no external benchmark, and it says a full automated AI researcher is targeted for March 2028.
Where do these OpenAI coding agent numbers come from?
They come from a post OpenAI published titled Research acceleration: The view inside OpenAI, alongside an essay by chief scientist Jakub Pachocki called An Alien Mind. The pricing figures in this piece were taken from OpenAI's own pricing pages the following morning.
