Skip to content

MiniMax quietly ships M3.1-Flash-Preview, a coding-only model for MiniMax Code

MiniMax's official account confirmed M3.1-Flash-Preview went live inside MiniMax Code on September 27, a day after users spotted it in the product, with no benchmarks and no published price.

By VibecodedThis 3 min read
MiniMax's official launch graphic for M3.1-Flash-Preview showing the model picker inside MiniMax Code
MiniMax

MiniMax shipped a new coding model this weekend and told almost nobody about it first.

M3.1-Flash-Preview appeared in the model picker inside MiniMax Code, the company's coding assistant, where Chinese AI watchers on X spotted it roughly a day before MiniMax's official account confirmed the launch on September 27. The confirmation was brief: the model is fast, reliable, and "ready for real work, from quick bug fixes to full features." No model card. No benchmark table. No published price. Startup Fortune

A gated launch, kept inside MiniMax's own product

What makes this release unusual is how little of it is open. The API endpoint for M3.1-Flash-Preview is gated, so the only way to try the model is inside MiniMax Code itself. Reasoning levels run from low up through a new "max" tier, sitting alongside the existing M3, M2.7, and M2.7-highspeed options.

Alongside the launch, MiniMax is running a promotion from September 28 to October 7 (UTC+8): new and existing MiniMax Code users get double free credits from daily check-ins, usable across supported models, including M3.1-Flash-Preview for coding and H3 and H3 Max for video. Token plan quotas reset when the model goes live, with more resets promised during the promotion period. Selnoviktech

That is a sharp contrast with how MiniMax shipped M3 in June: architecture details, SWE-bench scores, and a pricing sheet that undercut rivals by roughly 90 percent.

The numbers we do have, and what is still a claim

To calibrate what MiniMax can build when it wants attention: the company describes M3 as a 428-billion-parameter mixture-of-experts model with about 23 billion active parameters per token, a million-token context window, and native image and video input. MiniMax reports 80.5 percent on SWE-bench Verified and 59.0 percent on the harder SWE-Bench Pro, which it says puts it ahead of GPT-5.5 and Gemini 3.1 Pro on that test. Those are the company's own numbers on its own infrastructure, so treat them as a claim rather than an independent verdict. The pricing is independently visible, though: roughly $0.30 per million input tokens and $1.20 per million output through MiniMax directly, or about $0.23 and $0.96 through OpenRouter. That is around 5 to 10 percent of what comparable proprietary models charge.

Flash is positioned as the version built for routine coding work, the quick fixes and small features a developer runs dozens of times a day, rather than the hard agentic tasks reserved for the big model. OpenAI and Google split their lineups the same way. What stands out is that MiniMax is doing it around an open-weights flagship while keeping the faster variant inside its own tool for now.

The shelf is crowded

MiniMax is not shipping into a quiet market. Alibaba released Qwen3.8-Max on August 3, a 2.4-trillion-parameter flagship priced at $2 per million input tokens, then put out the first open-weights version of a Max-tier Qwen on August 12. Zhipu's GLM-5.3 landed August 14 with a million-token context window of its own. DeepSeek and Kimi are running the same playbook: ship often, price low, publish weights.

And developers are responding. Per CNBC's analysis of OpenRouter data, Chinese models accounted for 57 to 67 percent of token usage on the platform for the week including September 14, up from 6 to 13 percent in February. Vercel saw Chinese models' share rise to 55 percent in August from 11 percent in January. For the agentic coding and customer-service workloads that now dominate enterprise AI budgets, Chinese open models run 60 to 90 percent cheaper than the leading U.S. alternatives.

Washington is watching that shift with alarm. Daniel Remler of the Center for a New American Security told CNBC that Chinese AI poses "real economic and security risks for the United States."

Whether M3.1-Flash-Preview becomes a sustained product line or stays an experiment MiniMax lets the internet discover, the pattern is the same: near-constant releases designed to keep developers evaluating Chinese models instead of settling on one Western default.