AIOctober 6, 202612 min read
Best long context LLM in 2026, priced per full window
8 models with a window near a million tokens, ranked by what one full window costs to send. Reflection announced Beam on 5 October, and this morning there were still no weights to download.

The best long context LLM for a single big job is DeepSeek's flash model, which reads a million token prompt for about 15 cents of input, and the best one for work you plan to repeat is Claude Sonnet 5.5, which holds a million tokens at its normal rate with no surcharge and no beta flag to switch on. Google's own long context documentation says a million tokens is about 8 average length English novels, so those are the prices for pushing a small library through a model in a single request.
Reflection AI announced Beam on 5 October with a million token window of its own and weights promised under the Apache licence, and it is the name the field is loudest about this morning. We searched Hugging Face and the OpenRouter model list before writing this, and there is nothing to download and nothing to call, so Beam sits at the bottom of this ranking with an empty price cell instead of at the top.
Every ranking of these models we could find scores them on tests, so this one prices them instead, because a window you can't afford to fill is a number on a marketing page and nothing more. Each rate below was read off the vendor's own page this morning, and the arithmetic in the last column is ours.
What is the best long context LLM for each job?
The best long context LLM for each job comes down to 3 questions, how much text you will send, how often you will send it, and whether you need the weights on your own disk. The table settles the money question by taking the window each vendor publishes and multiplying it by the input rate each vendor publishes, which gives what one full window of reading costs, once, before the model has written anything back.

| Model | Window the vendor publishes | One full window of input | Weights you can download |
|---|---|---|---|
| DeepSeek flash | 1M tokens | $0.15 off peak, $0.30 peak | Yes, DeepSeek V4.1 Flash, MIT |
| GLM-5.3 (Z.ai) | 1,048,576 (OpenRouter listing) | about $1.47 | Yes, licence listed as other |
| GPT-6.1 Sol (OpenAI) | 922,000 input tokens | about $1.84 | No |
| Claude Sonnet 5.5 (Anthropic) | 1M tokens | $2.00 | No |
| Claude Opus 5.5 (Anthropic) | 1M tokens | $4.00 | No |
| Gemini 3.1 Pro Preview (Google) | 1,048,576 (OpenRouter listing) | about $4.19 at the long prompt rate | No |
| GPT-6 Astra (OpenAI) | 922,000 input tokens | about $9.22 | No |
| Beam (Reflection AI) | 1M tokens, announced | no price published | Announced for later this month, Apache 2.0 |
DeepSeek's flash model wins on price by a margin that makes you read the page twice, and the same page carries the conditions. Its rate doubles outside the discount hours the price list calls off peak, and the much lower figure sitting in the same row is the cache hit rate, which applies only to text the service has already processed for you. Even at the higher of the 2 rates, reading a full window through it costs less than a coffee, which is the number that makes a document pipeline thinkable for somebody with no budget.
Claude Sonnet wins the category that starts to matter the moment you run something twice, a window whose price doesn't move as you fill it. Anthropic's pricing page says long context requests are billed at standard pricing, and its context window documentation adds that a million is the default with no beta header to set. The example Anthropic prints itself is that a nearly full request is billed at the same rate per token as a short one, and it is the only vendor in this table that commits to that in writing.
Gemini 3.1 Pro is where a headline rate and a real rate come apart, because Google publishes one input price for short prompts and a higher one for long ones, so the window costs more at the end than at the start. GPT-6 Astra is the model to keep away from a long document, not because it reads one badly but because OpenAI prices its strongest model at 5 times the Sonnet rate, and the cheaper Sol model on the same page has exactly the same window for a fifth of the money.
There is a cheaper window than all of these, and the price is your text. Google's pricing page lists a free tier with no charge for input or output, limited to certain models, and the same page says content on that tier is used to improve its products. For a public document that trade is fine and for a client contract it isn't, so the free tier belongs in the ranking as the thing you try first and never the thing you automate.
What does Reflection's Beam change for long context?
Reflection AI's Beam, announced on 5 October, is a sparse model of about 501 billion parameters that switches on only a fraction of itself on every request, the design we took apart in what a mixture of experts model is, and it holds a million tokens like the leaders above. It changes nothing for anybody reading this tonight, because Reflection's own post says the weights, the technical report and the model card all arrive later this month, with a sign up form for early access until then.

Reflection says Beam matches a Chinese rival on reasoning while using 3 to 4 times less compute to answer, which is the company's own measurement and nobody else's. Its post also says the model is still going through final red teaming and evaluations, which is the honest reason the files aren't out. There is no price because there is no endpoint to price, and no file list, which is the document that would tell a reader what machine the model needs.
We checked the 2 places a reader would actually look. The Hugging Face API returns no model named Beam from Reflection and no organisation page under the company's name, and the OpenRouter model list, which carries a hosted version of every other model in the table above, has no Beam row among the hundreds it serves right now. Both checks will be wrong the day the weights land, and that is the day this ranking gets a new first row rather than a new footnote.
The story around the launch is that an American lab is answering the Chinese open models, and the table above shows what that answer is walking into. A model with a million token window is already sitting on Hugging Face under the MIT licence, the most permissive licence anything here carries, and it has been there for weeks with hundreds of thousands of downloads recorded on its page. Beam arrives against something already free to take, already free to sell with, and already hosted in a dozen places, so the licence alone will not be the news when the files appear.
How much does it cost to fill a million token window?
Filling a million token window costs 2 things most comparisons leave out, a rate that applies only to long prompts and the output you pay for on top. Google's pricing page lists a higher input rate and a higher output rate for prompts over 200,000 tokens, so the back of a Gemini window costs double the front, while Anthropic charges one rate the whole way and OpenAI's window stops short of the round number everybody repeats.

OpenAI's own model pages give 922,000 input tokens for both of the models it has in this table, not a million, and that gap is wider than it looks when you are deciding whether a book fits in one request. Aggregator listings add the output allowance to the input limit, which is how the same model shows up elsewhere as a million token model, and it is a fair way to describe a conversation and a misleading way to describe a document you need to send whole.
The input rate is also only half of a bill. Every model here charges more for what it writes than for what it reads, often around 5 times more, so a full window plus a long answer costs more than the last column of the table says. That column prices the reading on purpose, because reading is the number that compares cleanly across vendors, while the writing depends entirely on how much you ask the model to produce.
The cheapest way to read a long document twice is to avoid paying for it twice, and every vendor here has already built that discount. Anthropic, Google and DeepSeek all publish a much lower rate for text they have already processed for you, which they call cached input, and on the Claude models a cached read costs a tenth of the standard rate or less. If you plan to ask a dozen questions about the same book, the honest price is one full window plus a dozen cheap reads, not a dozen full windows, and that reordering decides which model is actually cheapest for your job.
Waiting is the other discount, and it is the easiest one in this table to take. Anthropic prices its batch service at half the standard rate on every model it lists, and the batch listings for the OpenAI and Google models run at half too, so a job you are willing to collect tomorrow morning costs half of what it costs right now. Nothing about a long document usually needs an answer in 4 seconds, so the waiting costs you nothing you would miss.
Which long context models can you download and run yourself?
Only 2 of the models in this ranking have weights a reader can download today, and both come from Chinese labs. Z.ai's GLM-5.3 has been on Hugging Face since late August with over a million downloads recorded, under a licence its page labels only as other, which means you open the file and read it before you build a business on it. DeepSeek V4.1 Flash sits beside it under the MIT licence, which is short enough to read in a minute and lets you sell what you make with it.

Downloading the weights is not the same as running them. A model of this size needs hundreds of gigabytes of memory to hold at once, which is a rack of server chips or a rented machine rather than a laptop, and that applies to both downloadable models here and to Beam when it lands. The versions that fit a desktop are the compressed ones, and which model file you should download covers what you give up when you take a smaller file.
The licence is the document to read before the specification, because it decides what you may do with the model and with what it produces. We ranked the open models that way in the best open source AI model, ranked by its licence, and the order changes a lot once you stop scoring on tests. An MIT model you can host yourself is worth more to a small business than a higher scoring model with a clause about who is allowed to use it.
Owning the weights also fixes the price, because nobody can reprice or retire a file that sits on your own disk. Every hosted rate in the table above is a rate until the vendor changes it, and the preview label on the Gemini row is a reminder that a model in preview is a model that can still move. A file on your own disk keeps working at the same cost per hour whatever the vendor announces next week.
Does a long context LLM remember the middle of a long document?
A long context LLM doesn't read the middle of a very long document as reliably as the start and the end, and the 2 vendors with the biggest windows say so in their own documentation rather than leaving it to somebody else to discover. Anthropic's context window page gives the effect a name and tells you to curate what goes in instead of filling the space because it is there.

As token count grows, accuracy and recall degrade, a phenomenon known as context rot.
Google is just as direct on its long context page, where the warning is attached to the case that matters most for a real document, which is looking for several things at once instead of one. Google's own sentence is the condition that decides whether a million token window can do your job or only look like it can.
In cases where you might have multiple 'needles' or specific pieces of information you are looking for, the model does not perform with the same accuracy.
Google gives the encouraging half of that finding as well, saying a model pulls a single piece of information out of a large block of text with very high accuracy in many cases. So a million token window is dependable for a search and shakier for an audit, and the practical reading is to send the whole book when you want one answer out of it and to split it when you want every answer. Splitting is also the cheaper habit, because the pages you really need cost less than the book they came from.
That habit is exactly what the price column measures. A model you send one huge file to, for one question, is a model you'll call one request at a time, and one request at a time is the shape of bill the table above prices. The expensive mistake is stuffing the window on every single call and meeting the total at the end of the month.
Which long context LLM should you pick tonight?
For one long document and one question, send it to DeepSeek's flash model and spend cents, and if the document is confidential read the retention terms before anything leaves your machine. For anything you will run every day, take Claude Sonnet, because a rate that doesn't change with prompt length is worth more than a few cents a call once a script is making the calls for you. How to call a model through OpenRouter is the quickest way to try both side by side without opening an account with either vendor.

For Google, the window is excellent and the bill has a step in it, so measure a real prompt and find out which side of the threshold your work sits on before you commit a pipeline to it. For OpenAI, the strongest model is priced like the strongest model, and the cheaper one on the same page has an identical window, so long document work belongs on the cheaper one until you have watched it fail at something you care about.
For anybody who wants the weights, take the MIT licensed model and plan for server hardware or a rented machine, because the memory is the real cost and it doesn't shrink because the download was free. We priced the rest of the market in our comparison of model API prices across 3 real jobs, and the ordering there holds for short prompts, which is most work that is not a document.
Whatever you pick, put a ceiling on the account before the first long prompt, which is what AI spending limits, and which providers actually stop the bill walks through provider by provider. A million token prompt costs a couple of dollars by hand and a lot more inside a loop that retries, and the retry is how a test becomes an invoice nobody approved.
What we do not know yet
4 questions about this ranking have no answer yet, and each one of them would move the order if it had one. They are listed here rather than smoothed over, because each answer is coming within weeks and each one of them will change what the table says.
- The day the Beam weights arrive, which Reflection's post places only at later this month.
- What Beam needs to run, because no file list exists yet, so the memory it takes cannot be worked out.
- Whether Beam's window holds up in a test nobody at Reflection ran, since its own numbers are all we have.
- What a hosted Beam will cost, if anybody hosts it, which decides whether its price cell stays empty.
One more gap belongs to Google rather than to Reflection. Its models page printed no token limit for any current model when we fetched it this morning, so the Gemini window in the table comes from the OpenRouter listing while the rate step for long prompts comes from Google's own pricing page, and those are 2 different sources doing one job. When the Beam weights appear, the first thing to read is not a score on somebody's test but the file list, because that is what decides whether this is a model you can run or only one you can read about.
Questions people ask
What is the best long context LLM in 2026?
For one off work, DeepSeek's flash model reads a full million token window for cents, and for anything repeated, Claude Sonnet holds a million tokens at its standard rate with no long prompt surcharge. Reflection's Beam matches those windows on paper and has no weights to download yet, so it cannot be ranked against them.
How many pages fit in a 1 million token context window?
Google's long context documentation puts a million tokens at about 8 average length English novels, which is a small library inside a single request. The real limit is whether the model can find what you need in all of it, not whether you are allowed to send it.
Is the best long context LLM also the cheapest one?
On these rates it almost is, because the cheapest full window belongs to DeepSeek's flash model, which is also the cheapest model here to run repeatedly. Its lowest published rate applies only in discount hours and to text already cached for you, so a flat rate from Anthropic costs more per window and is far easier to budget.
Can I download Reflection's Beam weights today?
No. Reflection's announcement says the weights, the technical report and the model card come later this month under the Apache licence, and a search of Hugging Face and of the OpenRouter model list this morning returned no Beam entry at all.
Does a long prompt cost more per token on Gemini?
Yes. Google's pricing page lists a higher input rate and a higher output rate for prompts above its long prompt threshold, so a nearly full window costs double what the headline rate suggests. Anthropic publishes the opposite and bills its million token window at the standard rate from end to end.
Which long context model can I run on my own machine?
Neither of the downloadable models in this ranking, because both need server hardware or a rented machine to hold the weights in memory. Smaller open models with long windows do exist and some of them will run on a strong desktop, so the question to ask is how much memory your machine has before you choose a window.
Why is Beam called an open weight model if there are no weights?
Because the licence and the release plan have been announced, not the files. Reflection says it will publish the weights under the Apache licence later this month, which is an open weight release when it happens and an announcement until then.
