Claude Opus 5 arrived July 24 with the same pricing as Opus 4.8, a 1 million-token context window, and benchmark numbers that push past Claude Fable 5 on agentic and computer-use tasks. It scores 96% on SWE-bench Verified, solves every problem on IMO 2026, and includes a new effort dial that lets you trade cost for capability per request.
Three new Gemini models landed on July 21, led by Gemini 3.6 Flash — cheaper, faster, and more accurate than its predecessor. Meanwhile Google confirms Gemini 4 pretraining has started, and 3.5 Pro is still stuck in partner testing.
Gemini 2.5 Pro and Gemini 3 Flash are being retired from all Copilot experiences on July 31. Kimi K2.7 Code, the first open-weight model in Copilot, is now available for Business and Enterprise plans.
Thinking Machines Lab released Inkling, a 975B-parameter open-weight model trained on 45 trillion tokens across text, image, audio, and video. Former OpenAI CTO Mira Murati is betting enterprises want AI they can customize, not just rent.
Kimi K3 has 2.8 trillion total parameters, a 1M-token context window, and native vision. It launched July 16 on the Kimi app and API. Open weights arrive by July 27.
Weeks after the $60 billion acquisition closed, SpaceXAI and Cursor shipped Grok 4.5 — a model built for software engineering, legal work, and finance that's now available across all Cursor plans.
DeepSeek V4 brings two model tiers, 1M-token context by default, and a hard cutoff on July 24 for legacy API names that millions of developers still use.
Google DeepMind delayed Gemini 3.5 Pro from June to July 17 after abandoning the 2.5 Pro foundation model. The rebuilt version adds a 2M token context window, Deep Think reasoning, and improved math.
Moonshot AI's open-weight Kimi K2.7 Code model expanded to Copilot Business and Enterprise plans on July 7, but it's off by default and requires admin action to enable.
Moonshot AI's Kimi K2.7 Code reached general availability in GitHub Copilot on July 1, marking the first open-weight model available in the platform's model picker.
OpenAI's next model family breaks into three tiers: Sol (flagship), Terra (balanced), and Luna (cheap and fast). A US government request limited the initial rollout to around 20 vetted partners. General availability is expected in coming weeks.
Microsoft's 5B-parameter in-house coding model MAI-Code-1-Flash reached general availability for GitHub Copilot Business and Copilot Enterprise on June 26, with admin policy controls required before users can access it.
The 13-day free window for Claude Fable 5 on paid plans closed on June 22. Starting June 23, all Claude subscribers need usage credits to access the model. Here's what that means in practice and what Anthropic has said about restoring it.
Two model changes hit GitHub Copilot on June 18: Microsoft's MAI-Code-1-Flash expanded from VS Code to eight more surfaces, and GitHub announced Opus 4.6 (fast) will be deprecated across all Copilot experiences on June 29.
Cohere released North Mini Code on June 9, 2026 under Apache 2.0. The 30B/3B mixture-of-experts model targets enterprise teams who want a capable agentic coding model they can run on-premises without vendor dependency.
Zhipu AI released GLM-5.2 on June 13, 2026 with a usable 1-million-token context window, two thinking-effort levels, and a promise of MIT open weights the following week. The release came two days after the US government ordered Anthropic to cut foreign access to its Fable 5 models.
Moonshot AI released Kimi K2.7-Code on June 12: a trillion-parameter open-source coding model that cuts reasoning token usage by 30% compared to K2.6. Every benchmark number comes from Moonshot's own proprietary evals.
On June 9, 2026, Anthropic made a Mythos-class model generally available for the first time. Claude Fable 5 tops frontier coding benchmarks, ships at $10/$50 per million tokens, and routes its most dangerous capabilities to an older model through new safeguards. Here is what launched, what the benchmarks show, and how the safety system works.
MAI-Code-1-Flash is Microsoft's first coding AI trained entirely in-house, without OpenAI's data or technology. It's rolling out to all Copilot plans now, and its benchmark numbers suggest it was built for token efficiency rather than headline scores.
MAI-Code-1-Flash, Microsoft's first purpose-built coding model for GitHub Copilot, started rolling out June 2 across Free, Pro, Pro+, Max, and Student plans. It's designed for fast, efficient responses at lower credit cost.
DeepSeek's 75% promotional discount on V4-Pro expired May 31, and the company made it permanent instead of reverting. Output tokens are now $0.87 per million.
Anthropic shipped Opus 4.8 forty-one days after Opus 4.7. The headline numbers are a five-point jump on agentic coding, a fast mode that costs a third of what it used to, and a new Claude Code feature that runs hundreds of subagents in parallel for codebase-scale work.
Between May 18-20, GitHub added cheap cloud agent models, brought Gemini 3.5 Flash to IDEs, launched auto model selection in VS Code with a 10% discount, and then stripped all Gemini models from web chat.
ESC
Start typing to search across tools, news articles, and reviews
No results found
Optional analytics and social embeds help us improve the site. It works fully without them.
Privacy
Privacy Settings
Choose which optional features you want to enable.
Your browser sent a global privacy opt-out signal. We respect that by default.