One million tokens.
A token is the unit a language model reads, about three-quarters of a word. The pages swirling around you hold one million tokens of real text: 3,000 printed pages of Alice in Wonderland, on repeat. This is what a million-token context window holds, and the six-year race that made it real.
The million tokens, laid out as a library: 3,000 pages, 75 across and 40 high. Every model in this story could hold some slice of these shelves in mind. Scroll in.
This is one token: a single word on a single page. Every context window in this story is measured in these.
GPT-3 read 2,048 tokens at a time: the six pages you can see. Anything you said earlier fell out of memory.
The chatbot that reached 100 million users held 4,096 tokens, twelve pages. Long conversations forgot their own beginnings.
32,768 tokens is 98 pages: a novella. For the first time you could hand the model a whole essay and ask questions about it.
Claude jumped to 100,000 tokens, 300 pages, enough to read The Great Gatsby in one pass.
Gemini 1.5 Pro was first to hold every page in view. Today a million tokens ships in open-weight models you can download and run yourself.
3,000 pages.
Those same pages, stacked: a million tokens is about 750,000 words, and the pile stands thirty centimetres tall. A model with this window reads the whole stack at once.
Five Harry Potters.
Philosopher's Stone through Order of the Phoenix is about 957,000 tokens: the first five novels in a single prompt, with room left for your question. The King James Bible runs about 1.04 million, just over the line.
83 hours of talk.
At 150 words a minute, a million tokens transcribes three and a half days of nonstop conversation, or every word from two weeks of full-time meetings.
75,000 lines of code.
At about 13 tokens a line, a whole production codebase fits in one window: every file and every test, read together. This is why context size changed how AI writes software.
Every stack of books is a model release; its height is the context window (log scale, each gridline is 10×). Amber stacks are open weights, blue are closed APIs. For three years the skyline is flat: open or closed, a language model read two to four thousand tokens.
MosaicML's StoryWriter took open weights to 65K in May; Claude hit 100K the same month. By November the closed frontier sat at 128–200K, and Yi-34B carried 200K into the open.
Gemini 1.5 Pro crossed a million in February and doubled it by June. Meanwhile 128K became the everyday default; Llama 3.1 made it standard for open weights.
Qwen2.5-1M matched the million with weights you could download, and MiniMax-Text-01 shipped four million the same month.
Llama 4 Scout claimed 10,000,000 tokens, a window larger than any public benchmark could fill.
DeepSeek-V4, MiniMax M3 and Kimi K3 all ship million-token windows as open weights. In 2024 one closed model served a million tokens; in 2026 you can download three that do.
Hover any stack for its numbers, or filter to open weights to see how fast the gap closed.
From 2,048
to 10,000,000.
In 2020 the frontier model held about six pages in mind. Six years later, an open-weights model you can download claims ten million tokens, a 4,883× expansion of machine memory.
The frontier is open.
Kimi K3, DeepSeek V4 and MiniMax M3, the three stacks that close the story, are open weights with million-token windows. All three run on Together AI, where this page was built. One window holds all 3,000 of these pages in a single prompt.