Frontier Models
Current State
The frontier model landscape as of late March 2026 is being reshaped by a dramatic user migration and intensifying multi-lab competition. The #QuitGPT movement β a coordinated wave of 2.5M users abandoning ChatGPT β pushed Anthropic's Claude to #1 on the App Store with a reported 70% enterprise win rate (meaning Claude is chosen over competitors in 7 out of 10 enterprise evaluation processes). This is the first time OpenAI has lost its dominant distribution position, and it signals that model quality and trust are overtaking brand loyalty as the primary selection criteria.
Google's Gemini 3.1 Pro arrived with benchmark dominance across 13 of 16 standard evaluations, including a 77.1% score on ARC-AGI-2 (a reasoning benchmark designed to test genuine abstraction and generalization, not just pattern matching) and 94.3% on GPQA Diamond (a graduate-level science Q&A benchmark where questions are written and verified by domain PhD holders). At $2/M input tokens and $12/M output tokens, Gemini 3.1 Pro is priced competitively for its capability tier. Meanwhile, xAI's Grok 4.20 Beta introduced a four-agent architecture (using specialized sub-agents named Harper, Benjamin, and Lucas that divide reasoning tasks internally) β a novel approach where a single user query is decomposed and processed by multiple coordinating agents before returning a unified response. This is architecturally distinct from standard single-pass inference and may indicate a shift toward multi-agent-inside-the-model designs.
OpenAI is responding to market pressure with aggressive financial moves, offering private equity firms 17.5% guaranteed returns plus early access to unreleased models β an unusual fundraising structure that ties investment returns to model capabilities, suggesting OpenAI views its model pipeline as its strongest asset. Anthropic's Sonnet 4.6 demonstrated extreme cost efficiency, scoring 100% on a practical 38-task coding benchmark at just $0.20 total cost, reinforcing the quality-per-dollar narrative driving the #QuitGPT migration.
The competitive picture has shifted from "who has the biggest model" to a three-way race on different axes: Google leading benchmarks, Anthropic leading user trust and coding reliability, and xAI experimenting with novel architectures. OpenAI's fundraising moves suggest it is preparing a major response. An accidental CMS leak revealed Anthropic's "Claude Mythos" β described internally as a tier above Opus with dramatically higher coding, reasoning, and cybersecurity scores, and notably flagged by Anthropic's own documentation as posing unprecedented cyber capabilities. This is the first known case of a frontier lab's own pre-release materials characterizing a model as a cybersecurity risk. Meanwhile, Apple's announcement that iOS 27 will allow users to replace Siri's AI backend with Claude, Gemini, or other models signals a shift toward the smartphone as a multi-model orchestration layer β transforming distribution dynamics by decoupling the AI provider from the device maker.
Update 2026-04-22: Three new vectors this window. (1) Open-weight frontier parity on real engineering: Moonshot Kimi K2.6 (1T MoE, open weights, $0.60/M) beats GPT-5.4 and Claude Opus 4.6 on SWE-Bench Pro (58.6% vs 57.7% vs 53.4%) β first open-weight model to beat closed frontier on a non-saturated coding benchmark. Combined with DeepSeek V4 arriving on Huawei Ascend, the "closed-model benchmark moat" is gone for coding tasks. (2) OpenAI Cloud-Next counter-salvo: ChatGPT Images 2.0 retakes image-gen lead at 1,512 ELO (+242 above prior); Workspace Agents launched for Business/Enterprise/Edu, directly competing with Google's Gemini Enterprise Agent Platform unveiled same day; Codex Labs signed partnerships with Infosys/Accenture/Capgemini/Cognizant/PwC/TCS (Codex 4M+ WAU, +33% in a week). (3) Google neutral-compute positioning: Gemini 3.1 Pro powers Deep Research + Deep Research Max (autonomous research agents); Gemini Enterprise Agent Platform hosts Claude alongside Gemini; multi-billion TML cloud deal on GB300 β Google now supplies compute to Anthropic, TML, and internally.
Update 2026-04-17: Two material shifts in the past 72 hours. First, Claude Opus 4.7 shipped on April 16 and retook the coding lead from GPT-5.4 β SWE-bench Pro jumped +10.9pp (53.4% β 64.3%), visual acuity leaped from 54.5% to 98.5% (collapsing a prior agent blocker), and multi-step agent workflows improved 14% while emitting one-third the tool errors. Second, the Mythos story moved from leak to confirmed gating: Bloomberg's in-depth feature reports Anthropic withheld Mythos from public release after red-teamers (Nicholas Carlini + Logan Graham's Frontier Red Team) confirmed the model autonomously generates full cyber intrusion toolchains against Linux and other widely-deployed systems. The White House OMB is coordinating controlled access for federal agencies, and ECB President Lagarde and BoE Governor Bailey publicly raised Mythos cyber risk at IMF/World Bank meetings β the first time an unreleased frontier model has been flagged as a systemic financial-stability category. Anthropic simultaneously launched Project Glasswing (industry coalition on AI-cyber threats) and began rolling out mandatory KYC (passport/license + facial scan) on Claude consumer accounts. The combined picture: Anthropic is establishing a "gated dual-use" deployment pattern where offensive-cyber-capable models ship to vetted defenders rather than open APIs.
Key Players
| Player | Current Model | Notable |
|---|---|---|
| Anthropic | Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, Mythos (gated) | Opus 4.7 SWE-bench Pro 64.3%, XBOW 98.5%; Mythos withheld on cyber risk |
| OpenAI | GPT-5 series | Broadest ecosystem, ChatGPT distribution |
| Google DeepMind | Gemini 3.1 Pro/Flash | Benchmark leader (13/16), best multimodal |
| xAI | Grok 4.20 Beta | Four-agent architecture, fast iteration |
| Meta | Muse Spark (closed), Llama 4 (open) | First closed model from MSL/Wang, strategic pivot |
Recent Signals
| Date | Signal | Significance | Source |
|---|---|---|---|
| 2026-04-27 | Microsoft and OpenAI end exclusive revenue-sharing deal β Bloomberg confirms Microsoft stops sharing revenue with OpenAI; OpenAI publishes "Next phase of Microsoft partnership." The foundational exclusive API resale + revenue-share structure from 2019 ends. Other hyperscalers can now offer OpenAI APIs on comparable commercial terms. β Azure loses OpenAI pricing moat; frontier-model API market shifts to multi-cloud competition. | significant | Bloomberg Β· OpenAI |
| 2026-04-26 | Claude 4.7 identifies writer Kelsey Piper from 125 words of unpublished text β Zero-shot stylometric fingerprinting. All alternative identification channels (account, browser, IP, topic) were eliminated by test design; the only remaining channel was prose structure. ChatGPT and Gemini failed the same test. β Frontier models appear to have absorbed enough public literary record to identify writers from ~100-word samples. Users with traceable writing histories cannot interact anonymously with Claude 4.7. | notable | The Argument Magazine |
| 2026-04-26 | OpenAI abandons SWE-bench Verified as primary frontier coding eval β "No longer used to evaluate frontier coding capability." Forces every shop to commit to a new eval. The cycle from Kimi K2.6 saturation to OpenAI abandons is 6 days β benchmark replacement compressed an order of magnitude. | significant | openai.com |
| 2026-04-25 | Cohere β Aleph Alpha merger β Schwarz Group ~β¬500M structured financing; combined entity $20B β First billion-dollar sovereign-AI consolidation. Targets defense/energy/finance/healthcare/manufacturing/telecom. Aidan Gomez stays Cohere CEO. Operates on STACKIT (Schwarz sovereign cloud). Cohere $240M ARR; Aleph Alpha minimal revenue. | significant | techcrunch.com |
| 2026-04-25 | Amateur (23) cracks 60-year-old ErdΕs conjecture using GPT-5.4 Pro β Liam Price proved ErdΕs conjecture about primitive sets. AI applied formula from related math area no human researcher had tried. First widely-reported amateur + frontier LLM case in mathematics. | significant | Scientific American |
| 2026-04-24 | Google β Anthropic up to $40B + 5GW TPU; Anthropic discloses $30B+ ARR (tripled QoQ, surpasses OpenAI per CNBC) β $10B at $350B + $30B contingent. 5GW TPU on top of 3.5GW Broadcom-TPU = 8.5GW total Google TPU. Anthropic now triple-stacked: Amazon ($13B/$100B) + Google ($10-40B + 8.5GW TPU) + CoreWeave-Meta ($21B). VCs offer $800B+ secondary. Reportedly plans IPO October 2026. | significant | TechCrunch |
| 2026-04-24 | Anthropic-NEC Japan deal β Claude (Opus 4.7, Code, Cowork) deployed to all 30,000 NEC worldwide employees β First Japan-anchor enterprise deal at frontier-lab scale. NEC BluStellar integration (finance/manufacturing/cyber/gov). | significant | Anthropic |
| 2026-04-23 | OpenAI GPT-5.5 + GPT-5.5 Pro launched β first fully retrained base since GPT-4.5; MCP-native; computer use, hosted shell, apply patch, Skills built-in β Natively omnimodal (text/image/audio/video). 1M-token context. Pricing $5/$30 per Mtok ($30/$180 for Pro) β frontier price went UP vs GPT-5.4 ($2.50/$10). Performance: Terminal-Bench 2.0 82.7%, Expert-SWE 73.1%, OSWorld 78.7%, FrontierMath Tier 1-3 51.7%; trails Opus 4.7 on SWE-Bench Pro (58.6 vs 64.3). NVIDIA confirms production on GB200 NVL72. Codex now powered by GPT-5.5 with marketplace + persistent memory + MCP support (4M+ WAU). | significant | openai.com Β· NVIDIA |
| 2026-04-23 | XBOW: GPT-5.5 reaches Mythos-like cyber capability β but freely available β Vulnerability miss-rate: GPT-5 = 40%, Opus 4.6 = 18%, GPT-5.5 = 10%. 97.5% visual acuity. Black-box outperforms GPT-5 with full source code access. Anthropic gated Mythos behind clearance criteria; OpenAI shipping equivalent capability open-to-all undermines the gating framework. | significant | xbow.com |
| 2026-04-23 | Anthropic publishes 7-week Claude Code regression postmortem β Three compounding silent-default bugs from March 4 β April 20: reasoning_effort highβmedium; prompt-caching mis-fire dropping reasoning state; verbosity prompt β€25 words. Resolved in v2.1.116. Anthropic resets usage limits as compensation. Lands during GPT-5.5 + DeepSeek V4 launch week β worst possible timing. | significant | Anthropic Engineering |
| 2026-04-22 | Google Cloud Next 2026 β Gemini on Blackwell preview; 200+ models on Gemini Enterprise β Google announced Gemini on Blackwell/Blackwell Ultra in preview for Google Distributed Cloud (sovereign/on-prem); 200+ models in Model Garden (incl. Claude Opus/Sonnet/Haiku, Gemma 4, Nano Banana 2); Gemini 3.1 Flash Image, Lyria 3. Also announced: multi-billion TML (Murati) cloud deal on GB300. β Google positioning as neutral compute substrate β running its own frontier (Gemini), a competitor's (Claude), and custom-for-lab (TML) on the same infrastructure. | significant | blog.google |
| 2026-04-21 | SpaceX/xAI Colossus + Cursor $10B deal with $60B acquisition option β Cursor will train Composer (its agentic coding model) on xAI's Colossus supercomputer. $10B compute-for-equity + $60B acquisition option exercisable later 2026. Cursor eng leads Milich + Ginsberg joined xAI reporting to Musk. Conflicts with April 17 NVIDIA/$50B Thrive+a16z round β needs watching. β Musk vertically integrating a coding-AI layer; if exercised, removes Cursor as Claude/GPT-neutral client. | significant | cursor.com Β· techcrunch.com |
| 2026-04-21 | OpenAI ChatGPT Images 2.0 (GPT-Image-2) β 1,512 ELO on LMArena, +242 above prior leader β New image-gen model with thinking-level intelligence, sharper editing, in-image text rendering fidelity. GPT-Image-2 hits 1,512 on Artificial Analysis Arena text-to-image leaderboard β 242 points above the cluster of all prior leaders (Nano Banana + GPT-Image were tied near 1,270). β Largest single-model jump on any frontier image benchmark in 2026 so far; OpenAI retaking the image-gen lead it lost to Nano Banana in Q1. | significant | openai.com Β· x.com |
| 2026-04-21 | Google DeepMind Deep Research + Deep Research Max (Gemini 3.1 Pro) β Autonomous research agents that navigate web + custom data (internal docs, specialized financial info) and produce fully cited professional reports. Built on Gemini 3.1 Pro. β Directly targets the "deep research" use case that OpenAI and Perplexity productized first; extends Google's agent push at Cloud Next. | significant | x.com |
| 2026-04-20 | Moonshot Kimi K2.6 β 1T MoE open-weight beats GPT-5.4 and Opus 4.6 on SWE-Bench Pro β 1T params (32B active), 262K context, native vision. SWE-Bench Pro 58.6% (GPT-5.4 57.7%, Opus 4.6 53.4%). SWE-Bench Verified 80.2%. Agent swarm 300 sub-agents Γ 4,000 steps. $0.60/M input β ~10Γ cheaper than closed frontier. Modified MIT license on HuggingFace. β First open-weight model to cleanly beat GPT-5.4 and Opus 4.6 on a non-saturated real-engineering benchmark. | significant | huggingface.co |
| 2026-04-16 | Claude Opus 4.7 GA β SWE-bench Pro 53.4% β 64.3% (+10.9pp, beats GPT-5.4 at 57.7%); SWE-bench Verified 87.6%; CursorBench 70% (from 58%); XBOW Visual Acuity 54.5% β 98.5%; GPQA Diamond 94.2%. 14% better multi-step agentic workflows with 1/3 the tool errors. Vision to 2,576px (~3.75MP, 3Γ prior). New xhigh effort level. Pricing unchanged at $5/$25 per M tokens. On API, Bedrock, Vertex, Foundry, Copilot (7.5Γ premium promo through April 30). β First major coding benchmark where Anthropic decisively retakes the lead from GPT-5.4; visual acuity jump collapses a major agent-blocking weakness. | breakthrough | anthropic.com |
| 2026-04-16 | Mythos too dangerous for release β Anthropic's unreleased frontier model can autonomously (not just assist) discover critical/high-severity vulnerabilities in software including Linux, generate working intrusion tools, and chain end-to-end cyberattacks. External red-teamer Carlini and Frontier Red Team (15 staff, Logan Graham) confirmed in hours. Limited release only to technology firms, financial firms, select partners for defensive use after CISA + Center for AI Standards and Innovation briefings. β First documented case where a frontier lab's own safety team successfully gated its flagship model from public release on capability (not policy) grounds. Establishes "autonomous offensive cyber" as the first new frontier-lab refusal class since the post-GPT-4 era. | breakthrough | bloomberg.com |
| 2026-04-16 | Opus 4.7 regressions flagged on day one β r/ClaudeAI reports 50% higher effective cost with context regression vs 4.6 (285 pts); MRCR long-context benchmark materially worse than 4.6 (285 pts). Ethan Mollick: adaptive thinking router flags non-math/code tasks as "low effort" with no manual override, producing worse results. β Day-one sentiment shift that advantages Opus 4.6 holdouts on long-context tasks; tests Anthropic's quality-control narrative at the release boundary. | notable | reddit.com / x.com |
| 2026-04-16 | Epoch AI poll: Claude US weekly usage up >40% β Epoch survey: Claude usage in the US rose >40% last month, only service with a clear upward trend. Implies "several million new weekly users" but ChatGPT ~30% share still far larger. β First quantified post-#QuitGPT share shift to Anthropic. | notable | x.com |
| 2026-04-15 | Anthropic KYC β identity verification rolling out β Mandatory verification via passport/driver's license + facial recognition scan. Published on official support page. Heavy backlash on r/ClaudeAI (935+ pts) and r/LocalLLaMA. β First major consumer frontier lab to require biometric identity verification; timing is consistent with Mythos-class capability gating. Materially changes developer/privacy-sensitive usage. | significant | support.claude.com |
| 2026-04-11 | Musk vs OpenAI "legal ambush" β changes strategy to seek Altman ouster, trial April 27 | notable | bloomberg.com |
| 2026-04-10 | Claude behavioral shift toward sycophancy β 403pts Reddit post β A highly upvoted Reddit post (403 points) documenting observable behavioral changes in Claude toward sycophancy (agreeing with users rather than providing accurate pushback). β User-observed behavioral regression in a frontier model; follows the Science journal sycophancy study from March 30. Raises questions about whether RLHF optimization is drifting Claude toward user-pleasing over accuracy. | notable | reddit.com |
| 2026-04-09 | Anthropic Advisor Strategy β tiered model collaboration API, Sonnet+Opus 74.8% SWE-bench β Anthropic introduced an "Advisor" strategy where Sonnet generates code and Opus reviews/corrects it in a tiered collaboration. This multi-model workflow achieved 74.8% on SWE-bench (a benchmark measuring real-world software engineering task completion). β Demonstrates that orchestrating cheaper and more expensive models together outperforms either alone; 74.8% SWE-bench is a new high for multi-model strategies and suggests the future of frontier performance is model composition, not just single-model scaling. | significant | anthropic.com |
| 2026-04-08 | Meta launches Muse Spark β first closed model from Meta Superintelligence Labs β MSL (led by Alexandr Wang, ~100 direct reports) built "Avocado" over 9 months. Closed-source (major pivot from open-source Llama strategy). Multimodal input (voice, text, image), text output. Competitive but not SOTA β company acknowledges coding gap vs Claude/GPT. Trained using Qwen and other open-source models. First in "Muse" model family. Meta stock +6%. Considering API access and subscription fees. β Signals Meta's strategic pivot to closed models under Wang, ending the era when Meta was purely the "open-source AI company." The hybrid strategy (open Llama + closed Muse) mirrors how cloud companies offer both open-source and premium tiers. | significant | bloomberg.com |
| 2026-04-08 | Claude outage April 8 β second consecutive day of service disruptions β Sonnet 4.6 exhibited elevated error rate between 23:00-01:50 PT. ~1.5 hour duration. Chat and Code affected (~80% of reports), API unaffected. 450+ user reports. β Platform stability concerns amid rapidly growing demand; second day suggests infrastructure scaling challenges rather than one-off incident. | notable | ibtimes.com.au |
| 2026-03-30 | Intercom Fin Apex vertical model outperforms GPT-5.4 + Claude at customer service β A domain-specific model (fine-tuned exclusively on customer service conversations and data) surpasses frontier general-purpose models on customer service benchmarks. Deployed by Intercom as the backbone of their Fin product. β Demonstrates that vertical fine-tuning (training a specialized model on domain-specific data) can overtake the most capable general models at targeted tasks; enterprises don't always need frontier models if they're willing to invest in domain-specific training. | significant | intercom.com |
| 2026-03-30 | Claude subscriptions doubled in two months β Anthropic's Claude subscription base (paid users across all tiers) doubled over the two-month period following the #QuitGPT migration. β Quantifies the revenue impact of the user migration signal and reinforces the Q4 2026 IPO thesis. | notable | bloomberg.com |
| 2026-03-30 | OpenAI ChatGPT App Store results lag Apple App Store β ChatGPT's own app marketplace is showing significantly lower traction than expected compared to the iOS App Store. β Signals OpenAI's platform ambitions are running behind its model ambitions; distribution via Apple (iOS 27) matters more than building a proprietary app store. | notable | bloomberg.com |
| 2026-03-30 | Cohere Transcribe β 2B parameter ASR tops open leaderboard, 14 languages β ASR (Automatic Speech Recognition β converting audio speech to text) model from Cohere using only 2 billion parameters reaches the top of the open leaderboard across 14 languages. β Demonstrates parameter efficiency: a 2B model can lead its category, contrasting with the "more parameters = better" assumption that dominated early LLM development. | notable | cohere.com |
| 2026-03-27 | Anthropic "Claude Mythos" leaked via CMS misconfiguration β Draft blog describes a model tier above Opus, "the most capable we've built to date." Dramatically higher scores on coding, academic reasoning, and cybersecurity vs Opus 4.6. Anthropic's own leaked language warns of unprecedented cyber capabilities. Currently in early access testing. β First frontier model leak where the lab's own docs describe it as a cybersecurity risk before release. | significant | Fortune |
| 2026-03-26 | Apple plans to open Siri to rival AI assistants in iOS 27 β Bloomberg reports Extensions in iOS 27 will let users select Claude, Gemini, etc. as default AI handler. WWDC 2026 (June 8) announcement expected. β Apple positioning iPhone as multi-model AI orchestrator. | significant | Bloomberg |
| 2026-03-24 | OpenAI shuts down Sora + Disney $1B deal dies β Sora app and API discontinued after 6 months (downloads fell 45%). Disney licensing deal terminated, no money changed hands. Compute redirected to core LLM products. β First major product death from a frontier lab; video generation deprioritized for text/agent capabilities. | significant | cnbc.com |
| 2026-03-24 | Anthropic launches Claude Computer Use for Mac β desktop control via connectors + screen control fallback. Includes Dispatch (iPhone-to-Mac task delegation). Permission-first safety. β Competitive with Meta Manus; positions Anthropic as first frontier lab with consumer desktop agent product. | significant | cnbc.com |
| 2026-03-23 | OpenAI offers PE firms 17.5% guaranteed returns + early model access β unusual fundraising structure tying investment to model pipeline, signaling OpenAI views unreleased models as its strongest competitive asset | significant | reuters via llm-stats.com |
| 2026-03-21 | #QuitGPT: 2.5M users abandon ChatGPT; Claude hits #1 App Store, 70% enterprise win rate β first major user migration away from OpenAI, driven by quality and trust rather than pricing | significant | fortune.com |
| 2026-03-20 | Gemini 3.1 Pro dominates 13/16 benchmarks β ARC-AGI-2: 77.1% (abstract reasoning), GPQA Diamond: 94.3% (PhD-level science Q&A), priced at $2/$12M tokens | significant | medium.com |
| 2026-03-20 | Sonnet 4.6 scores 100% on practical 38-task coding benchmark at $0.20 total cost β demonstrates extreme cost-efficiency for real-world coding tasks | notable | ianlpaterson.com |
| 2026-03-08 | Grok 4.20 Beta β four-agent architecture (sub-agents Harper, Benjamin, Lucas) where a single query is decomposed across specialized internal agents before returning a unified response | significant | labla.org |
| 2026-03-19 | Xiaomi MiMo-V2-Pro 1T revealed as mystery "Hunter Alpha" | significant | technology.org |
| 2026-03-17 | GPT-5.4 mini + nano launched ($0.20/M for nano) | significant | openai.com |
| 2026-03-15 | Zhipu AI previews GLM-5-Turbo, open-source release planned | notable | x.com |
| 2026-03-13 | Claude 1M context GA at standard pricing (Opus 4.6, Sonnet 4.6) | significant | claude.com |
| 2026-03-09 | DeepSeek V4 imminent β 1T params, native multimodal | significant | technode.com |
| 2026-03-05 | GPT-5.4 launched β 1M context, computer use, 33% fewer errors | significant | openai.com |
| 2026-03-03 | Google Gemini 3.1 Flash-Lite β 2.5x faster, $0.25/M tokens | notable | blog.google |
30-Day Trend
Accelerating sharply. Five frontier-class model launches or announcements in two weeks across OpenAI, Anthropic, Google, Xiaomi, and DeepSeek. Pricing race intensifying β GPT-5.4 nano at $0.20/M and Gemini 3.1 Flash-Lite at $0.25/M signal commoditization of inference. Chinese labs (Xiaomi, DeepSeek, Zhipu AI) emerging as serious frontier contenders.
What to Watch For
- GPT-5 successor / GPT-6 announcements from OpenAI
- Claude 5 or new model family from Anthropic
- Gemini 3.0 from Google
- Benchmark convergence β are models becoming commoditized?
- Pricing moves β aggressive cuts signal commoditization
- New entrants (Amazon, Apple, Samsung) reaching frontier level
- Regulation impact on model releases
Builder's Notes
(To be filled by daily scan β Phase 5)
Related Nodes
Source: nodes/frontier-models.md