Cursor released Composer 2.5 on May 18, built on Moonshot's Kimi K2.5 checkpoint with 85% of compute spent on Cursor's own post-training pipeline. It scores 63.2% on CursorBench v3.1, edging out both Opus 4.7 and GPT-5.5 at a fraction of the inference cost.
Google launched Gemini 3.5 Flash today at I/O 2026. It posts 78% on SWE-Bench Verified, runs four times faster than other frontier models, and costs $0.50 input / $3 output per million tokens. Here's the launch in detail.
Moonshot AI released Kimi K2.6 on April 20. The open-weight model scores 80.2% on SWE-Bench Verified and 66.7% on Terminal-Bench 2.0, with pricing of $0.60 per million input tokens on the official API. It's a meaningful upgrade over K2.5 on every benchmark that matters for coding agents.
Xiaomi released MiMo-V2.5-Pro on April 27: an open-source 1-trillion-parameter MoE model that scores comparably to Claude Opus 4.6 on coding benchmarks while using 40–60% fewer tokens per task. Weights are freely available on Hugging Face.
Anthropic's latest flagship model accepts images at more than 3x the previous resolution, adds a new xhigh effort level between high and max, and ships file system memory that persists across sessions. The per-token price is unchanged, but a new tokenizer maps the same input to up to 35% more tokens.
The Information reports Anthropic will release Claude Opus 4.7 and an AI design tool for websites and presentations this week. Figma fell 6% on the news.
Meta Superintelligence Labs unveiled Muse Spark, a multimodal reasoning model with three thinking modes, 1,000-physician-curated health training, and a pretraining stack that needs an order of magnitude less compute than Llama 4 Maverick. Full breakdown of benchmarks, modes, and what it means.
Google DeepMind's Gemma 4 ships four models from 2B to 31B parameters, all under Apache 2.0. The 31B Dense model ranks #3 among open models globally. Here's everything you need to know.
OpenAI releases GPT-5.4 Mini and GPT-5.4 Nano with 400K context, computer use support, and pricing starting at $0.20 per million tokens. Here's what developers actually get.
OpenAI releases GPT-5.4 with native computer-use capabilities, 1M token context, scalable tool search, and best-in-class agentic coding. Full breakdown of benchmarks, pricing, and what it means for developers.
Google DeepMind releases Gemini 3.1 Flash-Lite with $0.25/M input pricing, 363 tok/s output speed, and 1M token context. Here's the full breakdown: benchmarks, pricing, and what it means for developers.
Alibaba released the Qwen 3.5 Medium open-weight models on February 24, 2026. The 35B-A3B MoE model hits 111 tokens/sec on an RTX 3090 and handles 1M+ context on 32GB VRAM. Here's a full hardware guide with real tok/sec numbers.
Google releases Gemini 3.1 Pro with a 77.1% ARC-AGI-2 score, top marks on 13 of 16 benchmarks, and the same pricing as its predecessor. Here's what it does, where it leads, and where it doesn't.
Anthropic's Claude Sonnet 4.6 matches flagship-class performance on coding, computer use, and agent tasks while costing 5x less than Opus. Here's what the benchmarks say, what developers think, and what it means for your workflow.
MiniMax launches M2.5, calling it the first production-level model built natively for agent scenarios. Here's what the benchmarks show, what it costs, and how to actually use it.
Anthropic's Claude Opus 4.6 launches with 1M token context, agent teams, and adaptive thinking. Here's what changed, what matters, and what to actually use it for.