Claude Sonnet 5: Default Model Now, Beats Opus 4.8 on Two Benchmarks
Claude Sonnet 5 launched June 30 at $3/M input tokens and became the default for all free and Pro users July 1. It beats Opus 4.8 on Terminal-Bench and GDPval.
Topic
7 articles
Claude Sonnet 5 launched June 30 at $3/M input tokens and became the default for all free and Pro users July 1. It beats Opus 4.8 on Terminal-Bench and GDPval.
GPT-6 expected May-July 2026. Claude 5 "Fennec" targets May-September. Llama 4 is overdue. Here's what each model means for developers and what to prepare for now.
GPT-5.4 scores 80% on SWE-Bench at $2.50/1M input. Claude Opus 4.6 hits 81.4% SWE-Bench at $5/1M. Gemini 3.1 Pro leads reasoning at $2/1M. Full breakdown for developers.
Cursor launched Composer 2 on March 19: beats Claude Opus 4.6 on coding benchmarks at $0.50/1M tokens. Built on Kimi K2.5. Moonshot AI is now accusing Cursor of license violation.
DeepSeek's next flagship model is imminent — 1 trillion parameter MoE architecture, multimodal support, 1M token context, trained on Huawei Ascend. Here's what it means for developers and the widening US-China AI stack split.
OpenAI released GPT-5.4 on March 5, 2026 with native computer use — AI agents that operate desktop and web apps without wrapper code. 1 million token context, 33% fewer errors. Here is what this means for every developer building AI agents.
DeepSeek V4 launch: 1 million token context, multimodal, coding-first. Benchmarks vs GPT-4o and Claude, API pricing, and what developers actually get in 2026.