Tech Antenna logo

The AI Model Release Timeline: Every Major Launch From OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and Perplexity

I tracked every frontier model release from January 2025 to September 2026. Here is the chronological record, with capabilities, prices, and benchmarks.

By Jordan Lee
The AI Model Release Timeline: Every Major Launch From OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and Perplexity

I started keeping this timeline because I got tired of losing arguments about what shipped when. Somebody would claim a model had been out "for ages" and I would go looking for the announcement post and find it was six weeks old. The release cadence in this industry has compressed to the point where institutional memory lasts about a quarter.

So we built the record properly. What follows is a chronological account of the seven labs I watch most closely: OpenAI, Anthropic, Google DeepMind, xAI, DeepSeek, Moonshot AI, and Perplexity. Each section runs in release order, with the capability that mattered at each step. I have linked primary sources wherever they exist, and I have flagged the places where a lab published a number that deserves a second look.

One statistic frames everything below. According to PromptZone's release tracker, 2025 alone saw more than 25 notable releases across OpenAI, Anthropic, Google, Meta, DeepSeek, Alibaba, Mistral, and xAI. In 2026, that pace increased. Google shipped three Flash-tier models in six weeks. xAI shipped two flagships 35 days apart. If you are building on any of these APIs, the model landscape at your production date will not be the one you evaluated against.

 


 

The short version, in order

Date

Lab

Release

Why it mattered

Jan 20, 2025

DeepSeek

DeepSeek-R1

Open-weights reasoning at o1 level

Feb 17, 2025

xAI

Grok 3

200k-GPU training run, Think mode

Feb 24, 2025

Anthropic

Claude 3.7 Sonnet

First hybrid-reasoning Claude

Mar 25, 2025

Google

Gemini 2.5 Pro

Thinking by default, 1M context

May 22, 2025

Anthropic

Claude Opus 4, Sonnet 4

Long-horizon agent generation

Jul 9, 2025

xAI

Grok 4

Topped Humanity's Last Exam

Aug 7, 2025

OpenAI

GPT-5

Router blending fast and reasoning models

Sep 29, 2025

Anthropic

Claude Sonnet 4.5

30-hour autonomous agent runs

Nov 18, 2025

Google

Gemini 3 Pro

Deep Think mode

Jan 2026

Moonshot

Kimi K2.5

1T params, native vision

Apr 24, 2026

DeepSeek

DeepSeek V4

1.6T open weights, 1M context

May 28, 2026

Anthropic

Claude Opus 4.8

Effort control, cheaper fast mode

Jun 9, 2026

Anthropic

Fable 5, Mythos 5

New tier above Opus

Jul 9, 2026

OpenAI

GPT-5.6 Sol, Terra, Luna

Named capability tiers

Jul 16, 2026

Moonshot

Kimi K3

Largest open-weights model ever

Jul 24, 2026

Anthropic

Claude Opus 5

Near-Fable intelligence at half price

Aug 12, 2026

xAI

Grok 4.6

500k context, xhigh reasoning

Sep 1, 2026

Anthropic

Fable 5.1, Mythos 5.1

75% cache-read price cut

Sep 2, 2026

Google

Gemini 3.8 Flash

Third Flash release in six weeks

Sep 3, 2026

OpenAI

GPT-6 Astra

Computer use, Critical cyber threshold

 


 

1. OpenAI: from a router to a capability ladder

GPT-5 (August 7, 2025) was the consolidation release. OpenAI had been running two parallel product lines, the GPT series and the o-series reasoning models, and GPT-5 merged them behind a real-time router that decides how much reasoning to spend per request. It rolled out to every ChatGPT tier including free.

The point releases came fast. GPT-5.1 landed in November 2025 with adaptive reasoning depth. GPT-5.2 followed in December. GPT-5.4 arrived March 5, 2026 in Thinking and Pro variants, with mini and nano versions on March 17. GPT-5.5 shipped April 23, 2026, codenamed Spud, reporting 82.7% on Terminal-Bench 2.0 and 51.7% on FrontierMath Tiers 1 through 3.

GPT-5.6 (July 9, 2026) is where the naming changed permanently. OpenAI dropped the Instant, Thinking, and Pro scheme and replaced it with three durable capability tiers: Sol at the frontier, Terra balancing intelligence against cost, and Luna as the fastest and cheapest. All three carry a 1,050,000-token context window. On the API they price at $5/$30, $1.25/$7.50, and $0.50/$3.00 per million input and output tokens respectively. Luna became the ChatGPT free-tier default on August 6, 2026.

The release also introduced something I did not expect: a limited preview period gated by government restriction. GPT-5.6 went to a small group of trusted partners on June 26 before the public launch two weeks later.

GPT-6 Astra (September 3, 2026) is four days old as I write this, and it is the most consequential release in the set. OpenAI's vice president of research Aidan Clark told reporters that Astra involved by far the company's largest training run, saying "It's the first time we've pretrained on more than 100,000 GPUs" at the Stargate site in Texas.

The headline capability is computer use. Greg Brockman told reporters that "computer use is a particularly important part of what's new", describing a model that navigates spreadsheets, forms, and web pages at what OpenAI characterizes as superhuman speed. The specs: 1,050,000-token context, 128K max output, text and image input, knowledge cutoff April 30, 2026, and pricing at $10/$50 per million tokens, which is 2.5 times GPT-5.6 Sol's promotional rate.

The part that will define the next year is safety classification. OpenAI says Astra meets the Critical cybersecurity threshold under its Preparedness Framework, meaning it can find previously unknown vulnerabilities and work out how to exploit them without step-by-step direction. The public model refuses proof-of-concept exploit work; looser access goes to vetted organizations through the Daybreak program. Nvidia CEO Jensen Huang responded to the launch by declaring "AGI has arrived", which tells you more about GPU demand than about the model.

2. Anthropic: the fastest point-release cadence in the field

Anthropic's 4.x line ran on roughly a 2.3-month Opus cycle: Opus 4.5 (November 24, 2025), Opus 4.6 (February 5, 2026), Opus 4.7 (April 16, 2026), and Opus 4.8 (May 28, 2026). The Opus 4.8 announcement added user-facing effort control, dynamic workflows in Claude Code, and a fast mode running at 2.5 times speed for three times less cost than previous generations.

Then the ladder got a new top rung. Claude Fable 5 and Claude Mythos 5 (June 9, 2026) introduced a Mythos-class tier sitting above Opus. Both are the same underlying model. Fable ships with classifiers that route flagged cybersecurity, biology, chemistry, and distillation requests down to Opus; Mythos runs with those safeguards lifted and is restricted to approved Project Glasswing partners.

The rollout hit a wall immediately. On June 12, the US Department of Commerce prohibited access for non-US nationals regardless of location, and Anthropic revoked access to both models for all customers. Access was restored July 1, 2026. I include this because it is the first time an export-control action has pulled a commercially released frontier model off the market, and it will not be the last.

Claude Sonnet 5 (June 30, 2026) brought near-Opus capability to the mid tier at $2/$10 per million tokens with a 1M context window, scoring 85.2% on SWE-bench Verified.

Claude Opus 5 (July 24, 2026) is the release I recommend most often to teams watching cost. Priced identically to Opus 4.8 at $5/$25, it more than doubled its predecessor on Frontier-Bench v0.1, scoring 43.3% against Opus 4.8's 18.7% and GPT-5.6 Sol's 34.4%. On organic chemistry evaluations it gained 10.2 percentage points over Opus 4.8, and on protein-related tasks 7.7 points.

Fable 5.1 and Mythos 5.1 (September 1, 2026) arrived twelve weeks after their predecessors. Anthropic called them the "world's most advanced models for coding and knowledge work". Sticker price held at $10/$50, but cache-read pricing dropped 75%, which for agentic workloads dominated by cache reads is the number that actually moves budgets. The biology classifiers now fire 85% less often on benign medical questions. Fable 5's own SWE-bench Verified score of 95.0% remains the high-water mark on that benchmark.

3. Google DeepMind: a Flash treadmill and an empty Pro tier

Google's chronology is the most awkward of the seven, and I think it is the most instructive.

Gemini 2.5 Pro (March 25, 2025) took the top LMArena spot with thinking enabled by default and a 1M-token context window. Gemini 3 Pro (November 18, 2025) followed with Deep Think mode and generative UI in the Gemini app. Gemini 3 Deep Think shipped February 12, 2026 and Gemini 3.1 Pro on February 19, 2026.

Then the Pro tier stopped. At Google I/O on May 19, 2026, Sundar Pichai told developers that Gemini 3.5 Pro was a month away, asking them to "Give us until next month to get it to you". Bloomberg reported on July 16, citing ten current and former employees, that coding performance had fallen short of internal targets and that a late-June training data update made results worse rather than better. Google reportedly concluded the base model could not be patched and started pretraining a new flagship, Gemini 4.

As of early September 2026, Gemini 3.1 Pro is still the newest Pro-tier model in Google's own API listings. The Pro tier has been empty for more than three months past its announced date.

What Google shipped instead was Flash, relentlessly. Gemini 3.5 Flash-Lite and 3.6 Flash (July 21, 2026), Gemini 3.7 Flash (August 13, 2026), and Gemini 3.8 Flash (September 2, 2026), three Flash releases in six weeks. The 3.8 Flash release holds 3.7 Flash pricing at $0.75/$3.75 per million tokens through December 31, 2026, then doubles on January 1, 2027. That scheduled increase is buried in most coverage and belongs in any forecast that outlives this year.

Gemini 3.8 Flash supports a 1M-token context, 64,000 max output, and tunable thinking at low, medium, and high. Google shipped Gemini 3.8 Flash Cyber alongside it for vetted defenders through the new Fairwind Program. The Chrome Security team found it produced 2.6 times more correct vulnerability patches than the best commercial models, which are considerably larger. On CWE-Bench it hits 47.2% pass@1 against a frontier model's 47.8%, at a fraction of the cost.

Leadership changed in the middle of this. DeepMind CEO Demis Hassabis stepped down to become chairman in early August 2026, with Koray Kavukcuoglu taking over as SVP.

4. xAI: cadence over disclosure

Grok 3 (February 17, 2025) was trained on a 200,000-GPU cluster and shipped Think mode and DeepSearch. Grok 4 (July 9, 2025) trained on the Colossus cluster and topped Humanity's Last Exam at launch. Grok 4.1 Fast followed on November 19, 2025.

Grok 4.5 (July 8, 2026) brought a 500,000-token context window at $2/$6 per million tokens, positioned as a coding-focused successor. Grok 4.6 (August 12, 2026) landed just 35 days later at identical pricing and context, adding an xhigh reasoning level and scoring 61 on the Artificial Analysis Intelligence Index, fourth behind Claude Opus 5. On CursorBench v3.2 it scored 69.9%, up from Grok 4.5's 66.7% and slightly ahead of GPT-5.6 Sol Max at 67.2%, with Fable 5 Max leading at 70.5%.

Two caveats I would not skip. First, xAI published no architecture details for 4.6: no parameter count, no mixture-of-experts disclosure, no training compute, and no system card. Second, the release is explicitly developer-first, with the announcement covering Grok Build, Cursor, and the API without mentioning the consumer apps. Grok 4.7 is expected around September 12, 2026.

5. DeepSeek: the open-weights price floor

DeepSeek-R1 (January 20, 2025) matched OpenAI o1 on reasoning at a fraction of the cost and triggered roughly a $600B Nvidia sell-off. DeepSeek-V3.2-Exp (September 29, 2025) introduced DeepSeek Sparse Attention and halved API prices.

DeepSeek V4 (previewed April 24, 2026) arrived as two MoE variants. V4-Pro carries 1.6 trillion total parameters with 49B active per token; V4-Flash carries 284B total with 13B active. Both are MIT licensed with weights on Hugging Face. DeepSeek's own release notes put it plainly: "1M context is now the default across all official DeepSeek services".

Both models went generally available later, V4-Flash-0731 on July 31 and V4-Pro-0813 on August 13, 2026. The Flash release is the counterintuitive one: the 13B-active model beat the 49B-active V4-Pro preview on all nine agentic benchmarks DeepSeek publishes. Off-peak API rates run $0.66/$1.98 per million for Pro and $0.22/$0.66 for Flash, doubling at weekday peak. Per output token off-peak, V4-Pro is 12.6 times cheaper than Claude Opus 4.8 and 15.2 times cheaper than GPT-5.5.

Independent assessment tempers the claims. DeepSeek says it trails closed frontier models by three to six months. NIST's Center for AI Standards and Innovation evaluated V4 Pro in April 2026 and put the lag closer to eight months.

6. Moonshot AI: scale as a strategy

Kimi K2 (July 2025) established Moonshot as a serious open-weights coding lab. Kimi K2.5 (January 2026) brought 1 trillion parameters with 32B active and native vision through a 400-million-parameter encoder called MoonViT. K2.5 became the base for US models including Cursor's Composer 2 and Thinking Machines Lab's Inkling. Kimi K2.6 followed in April 2026 and Kimi K2.7 Code in June.

Kimi K3 (July 16, 2026) is the largest open-weights model ever released at 2.8 trillion parameters, roughly 75% larger than DeepSeek V4 Pro. Only 16 of its 896 experts activate per token. The architecture introduces Kimi Delta Attention and Attention Residuals, with a 1M-token context and native vision. Weights followed on July 27 under a modified MIT license that requires revenue sharing of up to 30% for inference providers earning over $20 million annually, which is not what most people mean by open.

Artificial Analysis scored K3 at 57 on its Intelligence Index and 1668 Elo on GDPval-AA v2, leading AutomationBench-AA at 53% at review time. The same evaluation flagged a 51% hallucination rate, so read the benchmark story as strong but qualified. The market reaction was immediate: on July 17, Z.ai fell roughly 28% and MiniMax about 16%, while the Nasdaq lost 1.4% and Nvidia 2.2%.

7. Perplexity: the orchestrator, not the trainer

Perplexity belongs on this list for a different reason. It is not primarily training frontier models; it is orchestrating everyone else's over its own search index, which makes its product timeline a useful read on how the rest of the field gets consumed.

Comet, the AI-native browser, launched July 9, 2025 for Windows and macOS behind the $200/month Max plan. Perplexity dropped the paywall on October 2, 2025, added Android on November 20, 2025, and iOS on March 18, 2026, where it hit #3 overall on the US App Store within 48 hours.

Model Council (February 5, 2026) dispatches one query to three frontier models simultaneously and synthesizes the outputs, a Max-tier feature. In February 2026 the company removed ads from answers and pivoted to subscription-first. Perplexity Computer (May 2026) arrived on Max as a cloud agent coordinating over 19 different models. In July 2026 the Sonar API was folded into an Agent API, and third-party trackers logged a lineup trim from six SKUs to four on July 5.

Sonar itself runs from $1/$1 per million tokens on the base model to $3/$15 on Sonar Pro, with a per-1,000-request fee of $5 to $14 by context level.

 


 

What the chronology actually shows

Four patterns hold across all seven labs.

Price and capability decoupled. Claude Opus 5 matched near-frontier performance at half the price of the tier above it. Gemini 3.8 Flash beat 3.7 Flash on every published benchmark at identical cost. DeepSeek V4 Flash beat the larger V4 Pro preview on every agentic benchmark. The assumption that better costs more stopped holding somewhere around mid-2026.

Context converged at one million tokens. GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4, and Kimi K3 all sit at or just above 1M. Grok is the outlier at 500K. Long context is no longer a differentiator; it is table stakes.

Cybersecurity became the gating function. GPT-6 Astra crossed OpenAI's Critical threshold. Gemini 3.8 Flash Cyber went to vetted testers only. Mythos 5.1 is invite-only under Project Glasswing. Three separate labs converged on the same structure within weeks of each other: one public model with classifiers, one gated model without them.

Open weights closed to within months, not years. Kimi K3 at 2.8T parameters and DeepSeek V4 Pro at 1.6T are shipping weights that trail the closed frontier by a gap measured in months. CAISI puts DeepSeek at eight months. Moonshot's license terms and China-hosted API introduce compliance questions that have nothing to do with capability, which is increasingly where the real decision lives.

If I had to give one piece of advice from building this timeline, it is architectural rather than analytical: keep an OpenAI-compatible endpoint you can repoint. Between July and September 2026 alone, six of these seven labs shipped a new flagship. Whatever you pick today, you will be re-evaluating it by December.