Kimi K3: China's 2.8T Open-Weight Model Beats GPT-5.6 Sol in Code
Moonshot AI released Kimi K3 on July 16 — a 2.8T sparse MoE model ranking first globally in Frontend Code Arena at 1,679 points. Full weights drop July 27.
Topic
8 articles
Moonshot AI released Kimi K3 on July 16 — a 2.8T sparse MoE model ranking first globally in Frontend Code Arena at 1,679 points. Full weights drop July 27.
xAI launched Grok 4.5 on July 8 at $2/$6 per million tokens — 60% cheaper than Opus 4.8. SpaceX acquired Cursor for $60B in June to build it. Benchmarks and developer guide inside.
OpenAI launched GPT-5.6 Sol, Terra, and Luna June 26 to 20 vetted partners. Sol is $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens. General access expected mid-July.
Gemini 3.5 Pro hits GA in late June 2026 with a 2M token context and Deep Think mode. Specs, pricing, and comparison to Fable 5 and GPT-5 explained.
Microsoft unveiled 7 in-house MAI models at Build 2026: MAI-Thinking-1 (35B active params, no OpenAI data) and MAI-Code-1-Flash now live in GitHub Copilot and VS Code.
Google's Gemini 3 Deep Think scored 84.6% on ARC-AGI-2, 48.4% on Humanity's Last Exam, and 3455 Elo on Codeforces. Gemini 3.1 Pro is now in preview. Here is what the benchmarks actually mean.
On March 11 a mystery 1-trillion-parameter model appeared on OpenRouter. The AI community burned 500 billion tokens assuming it was DeepSeek V4. On March 19 Xiaomi revealed it was theirs.
Most RAG tutorials show you how to build a demo. This post covers what breaks in production: chunking at 512 tokens beats semantic splitting, embedding costs range from $0.02 to $0.18 per million tokens, re-ranking boosts precision by 18–42%, and agentic RAG is now the 2026 standard. A practical guide for developers shipping RAG to real users.