AISeptember 18, 202611 min read
What Astra for Law costs, and what its 54% really means
OpenAI shipped a legal configuration of GPT-6 Astra on 17 September. No page in the launch carries a price, and the 54% everyone quoted was scored on questions the public leaderboard will not use.

Astra for Law is OpenAI's GPT-6 Astra model wrapped in a search index of American law and a set of instructions for legal analysis, announced on 17 September and given first to a small group of law firms through an application only program OpenAI calls Trusted Access. There is no published price for it, no date for the API version, and the 54.0% score quoted in every write up this morning was measured by OpenAI itself, on a set of questions the benchmark's own public leaderboard refuses to use.
That score is real and OpenAI published it openly, so this is not a story about a company hiding a number. It is a story about a benchmark whose public page explains rules almost nobody opened, and those rules change what the number means twice over. Everything below was checked at the source this morning, including the 4 pages where a price should have been and was not.
What is Astra for Law, and what did OpenAI actually ship?
Astra for Law is not a new model, it is the Astra model given a library and a house style. OpenAI's announcement describes 3 things bolted together, the flagship model it launched earlier this month, a search index built over American law, and custom instructions for legal analysis and writing. If you read our explainer on what Astra actually is, the brain here is the same one, and everything OpenAI added sits around it.

The index is the new material. OpenAI says it reaches United States case law, statutes, regulations, court rules and administrative decisions, and that it draws on the collection held by the Free Law Project, the nonprofit behind CourtListener. By the Free Law Project's own account, that collection covers effectively every published American precedential decision a lawyer would cite.
The instructions matter almost as much, even though a reader never sees them. OpenAI says they push the model toward the moves a lawyer makes without being asked, separating what a court actually held from the remarks it made along the way, dealing with the cases that damage an argument instead of walking past them, and explaining how an exception buried in a contract shifts risk from one side to the other. That is a description of behaviour rather than a measurement of it, and OpenAI presents it that way.
In ChatGPT it appears in the model picker as GPT-6 Astra Law, and OpenAI says the API version will carry its own name when it arrives. The same day, OpenAI launched more than 2 dozen connectors built by legal software companies, released its ChatGPT add in for Microsoft's document editor to business users, and named a set of community plugins written by lawyers and legal engineers. OpenAI does not present any of that as a research breakthrough, only as a better starting point for the firms building their own tools.
What does Astra for Law cost, and who can use it today?
Nobody outside OpenAI can tell you what Astra for Law costs, because no page in this launch carries a number, and today you can only reach it if OpenAI selects your firm. I checked the announcement itself, OpenAI's business pricing page, its ChatGPT pricing page and the website of Harvey, the legal AI company OpenAI named as a customer. The announcement quotes no figure at all, OpenAI's business pricing page lists Enterprise with a Contact sales button where a rate would go, and Harvey's own pricing address leads to a missing page.

Access runs through what OpenAI calls the Trusted Access Program, described as being for eligible law firms, covering lawyers and the people working under their supervision. For those firms, OpenAI says the offering includes Zero Data Retention on its API, and that ChatGPT Enterprise usage is excluded from human review by default. The API version is described as coming soon, with no date attached to the phrase.
So the practical answer for a solo practitioner, a small firm, or a person with a landlord problem is that this launch does not include you, and OpenAI has not said whether it ever will. A tool nobody can buy changes nobody's work this week, whatever it scores on a test. Most of yesterday's coverage treated this as a change already in effect, and for almost everybody reading, it is not.
What does the 54% score actually mean?
The 54.0% means Astra for Law produced an answer containing every single element a panel of practicing lawyers had marked as required, on that share of the questions it was given. It does not mean the system was wrong about half the time. The benchmark is Vals AI's Legal Research Bench, its primary metric is one Vals calls all pass, and Vals states the rule on its own page, a question counts as correct only when the response satisfies every rubric item, and scores nothing otherwise.

Those rubrics are long, because Vals says each one lists many required items, sometimes dozens, and that every item is something a correct answer must contain rather than a bonus. Miss a single required element and the answer scores zero, however good the rest of it is. Read that way, the figure describes a system that delivered everything a lawyer demanded on a bit over half of a set of hard questions, which is a very different claim from getting half of them right.
The questions themselves explain why the scores sit where they do. Vals publishes one sample openly, a Virginia haulage business that leased 2 refrigerated trailers through a finance company, watched both compressors fail, and then watched the manufacturer file for bankruptcy before honouring the warranty. Answering it properly means knowing that Virginia's version of the commercial code makes those lease payments irrevocable once the trailers were accepted, that the finance company owes no warranty of its own, and that the manufacturer's collapse lands on the customer rather than the lessor. Every one of those elements is required, and an answer carrying 2 of the 3 scores nothing at all.
It helps to know where that sits. Averaged across every model Vals has run on this benchmark, the all pass rate is under a third, so most answers come back incomplete by a lawyer's standard. The questions and their rubrics are written and reviewed by practicing lawyers, and Vals lists Fisher Phillips, McDermott and Reed Smith among its evaluation partners. Grading is done by a language model acting as judge, which Vals states openly, and which is worth remembering before treating any of these figures as a measurement of legal quality.
How does Astra for Law compare with models anyone can already use?
You cannot put Astra for Law's score on the same table as the public leaderboard, and the reason is printed on the benchmark's own page. Vals says the Legal Research Bench splits its questions 3 ways, a handful of open samples, a validation set of 200 that it licenses to anybody who asks, and a separate test set that it keeps entirely to itself.
The Test set will remain private. All results reported on this page are based solely on the Test set to prevent overfitting.
OpenAI's announcement names the set it used, and it is the validation set, the one available under licence. The leaderboard everybody else appears on is built on the other set. Both numbers are honest and both were published by the people who produced them, but they come from different exams, so a reader treating the announcement figure as a rank against the board is comparing 2 things that were never measured together.

| System | Score | Measured by | Which questions |
|---|---|---|---|
| Meta Muse Spark 1.3 Max | 55.29% | Vals AI | Test set, 208 questions held back |
| Claude Opus 5 | 55.29% | Vals AI | Test set, 208 questions held back |
| Claude Fable 5.1 | 55.29% | Vals AI | Test set, 208 questions held back |
| Astra for Law | 54.0% | OpenAI | Validation set, 200 licensed questions |
| GPT-6 Astra with web search | 38.7% | OpenAI | Validation set, 200 licensed questions |
Read the first column and the story stops being about one launch. 3 models from 3 different companies sit tied at the top of this test, not one of them is anywhere near reliable, and the configuration OpenAI built specially for law lands in the same band on a different set of questions. Legal research done properly is hard, and no system measured so far is close to doing it unsupervised.
The board also measures what each run costs and how long it takes, and the systems tied at the top are nowhere near each other on either. Meta's entry finishes a question in about 5 minutes for well under a dollar, while Anthropic's top entry needs nearly an hour and costs dozens of times more for exactly the same score. Identical accuracy at wildly different prices is the sort of spread that decides what a firm actually deploys, and it never appears in a launch post.
What does this do to Harvey, Legora and Thomson Reuters?
Nothing this week, because OpenAI named Harvey and Legora in its own announcement as API customers who will build on Astra for Law rather than compete against it. The same document says the new index complements the licensed content and specialist products firms rely on from providers such as Thomson Reuters, which is about as direct a reassurance as a supplier can put in writing.

Harvey announced a funding round at a $15.5 billion valuation earlier this month, and just over a week later the company whose models it builds on shipped a legal configuration of its own, with its own search index over American case law. Nothing in the announcement suggests a break between them, and Harvey stays on the connector list, so this reads as a supply relationship rather than a fight.
The named customers tell you who this was built for. OpenAI says Sullivan & Cromwell built an agreement analyzer that brings the firm's negotiating playbooks into the review of a new deal, that Ropes & Gray built a system for working through a data room, and that Cooley built a tool for preparing a company to go public. Each one involved forward deployed engineers from OpenAI sitting alongside the firm's own lawyers. That is a description of bespoke enterprise work rather than a product anybody downloads, and it is the clearest available signal about where Astra for Law sits today.
Thomson Reuters is bringing its HighQ matter context into ChatGPT and has previewed a CoCounsel Legal connector, so the incumbents are inside the building rather than locked out. What nobody can answer is whether this configuration is better than what those companies already sell, because no such test exists. Legal IT Insider, briefed before the launch, said so directly.
There is no benchmark available comparing it with leading legal research providers.
Westlaw, Lexis and CoCounsel have not been measured against Astra for Law by anybody, in either direction, and until somebody does it the ranking everybody wants simply does not exist.
What does Astra for Law change for you tonight?
If you are not a lawyer at a firm OpenAI selected, Astra for Law changes nothing you can touch tonight, and the sensible move is to treat a chatbot's legal answers exactly as you did last week. The reason is the scoring rule from earlier. The best systems anybody has measured return a complete answer on a bit over half of the questions lawyers wrote, and you have no way of knowing which half you just received.

One habit costs nothing and it works on every chatbot. When you ask any chatbot a legal question, ask it for the authority by name, the statute section or the case citation, then go and read that authority yourself. CourtListener, the database OpenAI licensed for this launch, is free and open to the public, so anybody can look up an American decision and check whether it says what the chatbot claimed. That check takes a few minutes and it catches the failure that gets people into real trouble, a confident summary of a case that was later overturned.
For a lawyer inside one of the selected firms, the change is real but narrower than the headlines suggested. They get research and drafting in one place, with the authorities linked, rather than searching in one tool and writing in another. John Savva, a partner at Sullivan & Cromwell, told Legal IT Insider that lawyers have generally had to finish their research first and then bring AI in for the analysis, and that this moves the 2 closer together. Whether it saves real hours is something only those firms can report, and none of them has published a figure yet.
What did ship for everyone that same day is smaller and more ordinary. OpenAI made its ChatGPT add in for Microsoft's document editor generally available, so business users can proofread and get suggested edits inside the file they already draft in. The legal connectors need a business or enterprise plan and a firm that installs them, which again leaves out most readers of this site.
What we do not know yet about Astra for Law
The list of open questions is longer than the list of answers, and every item below comes from something OpenAI did not say rather than something it got wrong. These are lines to watch rather than accusations, and any one of them could be answered next week.
- No price, for the ChatGPT route or for the API version.
- No date for the API, only the phrase coming soon.
- No independent measurement, because nobody outside OpenAI has run Astra for Law on anything.
- No test against Westlaw, Lexis or CoCounsel, in either direction.
- Nothing said about whether smaller firms or individuals will ever get access.

One more detail needs stating carefully, because it is a question rather than an accusation. OpenAI says it evaluated on the validation set, which Vals sells under licence, and it does not say whether that set was also used while tuning the configuration. Vals keeps a separate test set out of everybody's hands precisely because that distinction changes how a score should be read. Neither company has claimed anything improper, and licensing the validation set is exactly what it is sold for.
The comparison OpenAI published against a rival model sits in the same category. OpenAI's write up describes its own example, run by OpenAI, in which a Claude Fable model returned a holding that had been reversed on appeal. Nobody outside OpenAI has reproduced that, and a single example chosen by the company doing the comparing is an illustration rather than evidence. We covered a similar situation when a lab's own evaluation was the only source for a Claude Fable result.
What to watch over the next few weeks
4 things will tell you whether this launch is as large as yesterday's coverage suggested. The first is OpenAI's own pricing page, because the moment the legal model appears there with a figure beside it, the question every firm is asking gets answered. The second is the Vals leaderboard, since an entry for Astra for Law on the held back test set would be the first measurement nobody at OpenAI controlled.
The third is whether Harvey or Legora actually ship something built on it, which would say more about the configuration's value than any score. The fourth is Thomson Reuters, whose CoCounsel Legal connector is still in preview, because a company that competes with the index and connects to it at the same time will eventually have to choose. Our earlier piece on what OpenAI's Agents API costs followed exactly this pattern, a launch announced loudly with the meter published quietly afterwards.
Until then the defensible summary is short. OpenAI built a legal research configuration, published an honest number for it, and handed it to firms it picked. That number sits in the same band as models you can already rent by the token, which you can see in our comparison of what each model charges, and the scoring rule behind it says the whole field still fails the majority of questions that working lawyers write. That is a real step forward and it is nowhere near the end of the work.
Questions people ask
What is Astra for Law?
Astra for Law is OpenAI's flagship Astra model combined with a search index over United States case law, statutes and regulations, plus custom instructions for legal analysis and writing. OpenAI announced it on 17 September 2026. It appears in the ChatGPT model picker as GPT-6 Astra Law and will carry its own name in the API.
How much does Astra for Law cost?
OpenAI has published no price for Astra for Law. The announcement quotes no figure, and OpenAI's business pricing page lists Enterprise with a Contact sales button rather than a rate. The API version has no price and no launch date, only the phrase coming soon.
Can I use Astra for Law today?
Only if OpenAI selected your firm. Access runs through a Trusted Access Program that OpenAI describes as being for eligible law firms and the people working under their lawyers' supervision, inside ChatGPT and Codex. There is no public sign up and no consumer version of it.
What does the 54% score on the Legal Research Bench mean?
It means Astra for Law produced an answer containing every element a panel of practicing lawyers marked as required, on 54.0% of the questions it was given. Vals AI counts a question as correct only when every rubric item is satisfied and scores it nothing otherwise, and its rubrics list many required items, so a single omission scores the whole answer zero.
Is Astra for Law better than Harvey or Westlaw?
Nobody has measured that comparison yet. No benchmark comparing Astra for Law with established legal research providers exists, which Legal IT Insider noted on the day of the launch. Harvey and Legora are named by OpenAI as API customers who will build on Astra for Law rather than compete with it.
Does Astra for Law replace Westlaw or Lexis?
OpenAI does not claim that it does. Its announcement says the new legal search index complements the licensed content and specialist products firms rely on from providers such as Thomson Reuters. No test comparing Astra for Law with Westlaw, Lexis or CoCounsel has been published by anybody, so there is no evidence in either direction.
Can ChatGPT give me legal advice now?
No, and this launch does not change that for a member of the public. The best systems measured on the Legal Research Bench return a complete answer on a bit over half of the questions that lawyers wrote, and you cannot tell which half you received. Ask any chatbot for the case citation or statute section, then read that source yourself on a free database such as CourtListener.
