📑 Table of contents

MiniMax M3.1 Flash Preview: China continues its war of small, fast models

LLM & Modèles 🟢 Beginner ⏱️ 15 min read 📅 2026-09-29

MiniMax M3.1 Flash Preview: China continues its war of fast small models

🔎 A discreet launch at the heart of an avalanche

On September 27, 2026, MiniMax put M3.1-Flash-Preview into production. No keynote, no model card, no press release. A model id appeared in the documentation, placed at the top of the model table, at the precise moment when the market is trying to digest 46 models released in 30 days by 19 different providers (LLM MarketCap, September 2026).

The discretion is all the more striking given that the segment is saturated. Gemini 3.8 Flash at Google, GLM 5.3 Flash at Z.AI, Qwen3.8 Flash at Alibaba: the "Flash" category — fast, multimodal, cheap — has become the main battleground of autumn 2026. Every lab has deployed at least one model there, and the price war has been raging for weeks.

Why care about a model with no benchmark and no public pricing? Because M3.1 Flash sums up the current Chinese strategy: target high-frequency agents — those programs that call an API dozens of times per task — and monetize usage through subscriptions rather than tokens. This choice — breaking with M3's open-weight approach — says more about the state of the market than any ranking.

We went over what we know, what we don't know, and what it concretely changes for your agents.


Key Takeaways

  • MiniMax M3.1-Flash-Preview went live on September 27, 2026, hot on the heels of the MiniMax Agent launch, accessible only via the Token Plan (subscription) and MiniMax Code — no pay-as-you-go, no published per-token pricing.
  • One million token context, text/image/video inputs, five effort levels (default: max) — but zero independent benchmarks on launch day (0/486 on BenchLM).
  • Thinking is mandatory: any attempt to disable it returns an HTTP 400 error, whereas MiniMax-M3 still allowed it as of June 2026.
  • Against Gemini 3.8 Flash ($0.75/M input, $3.75/M output on OpenRouter, September 2026), M3.1 Flash can't be compared on either price or performance: too many unknowns.
  • The strategy is clear: lock in usage through subscriptions, quotas, and credits rather than compete on a public rate card.

Tool Main use Pricing (September 2026) Best for
MiniMax Code Test M3.1-Flash-Preview on code Token Plan subscription (pricing not published, check minimax.io) Developers already sold on the M family
OpenRouter Compare Space Bunny Alpha (free) and Gemini 3.8 Flash Free to $3.75/M output In-house multi-model benchmarks
BenchLM Track the arrival of M3.1 Flash benchmarks Free Verify before you buy
Hostinger Host your agent orchestration From a few €/month (check hostinger.com) High-frequency API call infrastructure

M3.1 Flash Preview: what exactly is it?

The direct successor to MiniMax-M3, launched on September 27, 2026 in the wake of the MiniMax Agent rollout, and positioned by the vendor as "fast and reliable, from quick bug fixes to full feature work". A model for code and agents, not for show. It sits at the top of the model table in the MiniMax documentation, above M3 (Orcarouter).

The announced specs are aggressive: a one-million-token context window, native multimodality (text, image, and video input), five effort levels ranging from low to max — with max as the default. The endpoints come in Anthropic style (marked "recommended" by MiniMax) and OpenAI style, on the same model id (Cocoloop).

According to First Financial, as reported by Gate News, the model is explicitly development-oriented: bug fixing, feature building, test verification. The MiniMax Code positioning confirms this focus — Token Plan credits can be used on M3.1-Flash-Preview for code and on the H3 and H3 Max models for video (promptblueprints). To place this type of model in the broader market landscape, our comparison of the best LLMs for coding is updated every month.

A telling detail about the fuzziness surrounding the launch: the September 30 date put forward in some calendars is a community prediction, not an announcement — the model had already been live since the 27th (eyestech).

What we don't know (and the list is long)

No model card. No evaluation report. No per-token pricing. No downloadable weights — only a restricted Hugging Face repo, MiniMax-M3.1-preview-private, of about 250 GB, spotted but inaccessible (Cocoloop).

Even the speed is merely a claim: "more than 100 tokens/s" is a vendor figure, with no public test harness. In 2026, launching a model without benchmarks or pricing is not an oversight. It's a strategic choice.


Zero independent benchmarks on launch day

None. As of September 27, 2026, BenchLM lists 0 benchmarks out of 486 for M3.1 Flash Preview: unranked, no overall score, no listed context window (BenchLM).

To situate the family, you have to go back a generation. MiniMax M3 scores 54.86 on BenchLM, M2.5 caps out at 49.89 and M2.7 at 48.08. The release cadence is telling: three major iterations in a few months, each promising the performance leap the previous one didn't always deliver publicly.

If M3.1 Flash follows the trajectory, it should cross the 55-point mark. But "should" is not data. It's a bet you're making with your API budget.

The worst part is that the absence of a benchmark has become a signal in itself. In an avalanche of 46 models in 30 days, launching without independent evaluation means counting on the noise to exist. Teams building agents don't have time to test 46 models: they look at the rankings, then the rate card. M3.1 Flash appears in neither. Our advice: keep an eye on the BenchLM page before making any budget commitment.


Against Gemini 3.8 Flash, GLM 5.3 Flash and Qwen3.8 Flash

There is no way to separate these models on the numbers: M3.1 Flash publishes neither per-token pricing nor benchmarks. The comparison can only cover access, announced specs, and the level of transparency.

Model Context Input / output price Access Thinking
MiniMax M3.1 Flash Preview 1M tokens (announced) Not published (27/09/2026) Token Plan / MiniMax Code only Mandatory, effort low→max
MiniMax M3 (June 2026) — $0.30 / $1.20 (cache $0.06) Open-weight + API Toggleable, binary effort
Gemini 3.8 Flash 1,048,576 tokens $0.75 / $3.75 API, OpenRouter —
Space Bunny Alpha 1M tokens Free on OpenRouter OpenRouter Adjustable

Sources: OpenRouter, Orcarouter, September 2026.

Against Gemini 3.8 Flash, the duel is lopsided on transparency. Google displays a clear rate card — $0.75/M input, $3.75/M output for 1,048,576 tokens of context (OpenRouter, September 2026). MiniMax answers with a subscription and no per-unit pricing. On paper, the two models share roughly one million tokens of context and multimodality. On the bill, only one of the two is predictable.

Against GLM 5.3 Flash and Qwen3.8 Flash, the balance of power is asymmetric: the two Chinese models went head-to-head on pricing from launch, on the very same day, as we analyzed in our article on GLM-5.3-Flash and Qwen3.8-Flash-Next. M3.1 Flash refuses to play on that field: no public pricing, therefore no price war to wage.

The only comparative data available are uncontrolled hands-on tests, to be treated as documented anecdotes rather than benchmarks: GLM 5.3 Flash Max did better than Space Bunny on an animated SVG test, and DeepSeek V4.1 Flash produced a more "walkable" Minecraft than a Space Bunny attempt (eyestech). No public harness, no shared methodology — but a density signal: the Chinese Flash generation is already so plentiful that developers are comparing four models on niche use cases.

Then there's the strange Space Bunny Alpha: free on OpenRouter, one million tokens of context, text/image/video, tool calling, adjustable reasoning, no named developer. The tokenizer fingerprint — 50 out of 50 strings matching MiniMax models, tested twice in two passes — ties it to the MiniMax family without identifying the checkpoint (eyestech). Could MiniMax have released two models on the same day, one paid and locked down, the other free to capture the leaderboards? A tempting hypothesis, strictly unconfirmed. In the meantime, the free alternatives that are actually documented are in our selection of the best free LLMs.


Mandatory thinking, a risky bet for agents

You won't be able to turn off M3.1 Flash's reasoning, and this choice weighs directly on your costs and latency.

The technical detail is a treat: any attempt to disable thinking returns an HTTP 400 error with an explicit message — "message requires adaptive thinking". MiniMax-M3, launched on June 1, 2026, allowed you to disable it. M3.1 Flash does not. The vendor's official answer: lower the effort rather than switch it off (Orcarouter).

The problem isn't philosophical, it's accounting. The five effort levels come with no quantified consumption figures: nobody knows what a pass at low effort costs versus max (Cocoloop). Yet an agent that chains 30 calls per task multiplies that unknown cost by 30. The most important variable in your agent economics — the cost of reasoning — is a black box.

For coding use cases, mandatory thinking holds up: refactoring tasks gain reliability with a model that always reasons. For extraction, routing, or high-frequency classification, it's a structural handicap compared to models that offer a direct mode. Agent frameworks will have to adapt — or look elsewhere. Our selection of the best LLMs for agents details the alternatives that leave this setting in your hands.


The real strategy: subscription vs. rate card

MiniMax doesn't want to fight on per-token price. It wants subscribers.

The contrast with M3 is striking. M3, on June 1, 2026: open-weight, public rate card — $0.30/M input, $1.20/M output, cache at $0.06/M — optional thinking, binary effort. M3.1 Flash, three and a half months later: subscription only, no downloadable weights, no unit pricing, mandatory reasoning. Two opposing philosophies within the same family (Orcarouter).

The Token Plan mechanics deserve a closer look. Two types of non-interchangeable credentials: a Token Plan Subscription Key on one side, a pay-as-you-go API key on the other. And an ongoing documentation inconsistency: the footnote on the Token Plan pricing page still lists M3, M2.7, and the image/speech models — without M3.1-Flash-Preview — while the model pages claim it's available on Token Plan and nothing else (Orcarouter). When the documentation contradicts itself on launch day, that's rarely a good sign for the maturity of the offering.

Add the acquisition promo: double check-in credits from September 28 to October 7 (UTC+8) and Token Plan quota resets for all users. Quota resets, double credits, a subscription-exclusive model: the goal is to build the habit before the benchmarks arrive.

This strategy isn't isolated, but its variant differs from the Western one. OpenAI launched GPT-6 Sol and Luna at half price, 90 minutes after Anthropic's competing release — an open price war, with published rates (our analysis). MiniMax, for its part, hides its pricing behind a flat-rate plan. Same target, opposite tactic: not selling the token cheaper — selling quota peace of mind.


46 models in 30 days: China industrializes Flash

This launch is not an event, it's a production line.

The figures from LLM MarketCap give the scale: 46 models in 30 days, 19 active providers. In late September alone, China rolled out GLM-5.3-Flash and Qwen3.8-Flash-Back on the same day, DeepSeek V4 released in Pro and Flash versions (our analysis), not to mention the parallel battle over image between Tencent Hy Image 3.5 Preview and Alibaba's opening of Apsara (our coverage).

The common target is identifiable: high-frequency agents. An agent that calls an LLM 50 times per task is 50 times more sensitive to per-token pricing and latency than a human sitting in front of a chatbot. This is exactly the terrain where Chinese models, structurally cheaper, have an advantage — and it's where the Flash category was born.

The paradox is that this war is unfolding while open-weight is winning the infrastructure battle: open models already run on 56% of Vercel's tokens (our analysis of the silent migration). By closing M3.1 Flash, MiniMax is swimming against the current of its own history — M3 was open — and against the broader movement. The underlying bet: value would no longer lie in the weights, but in the quota.


Testing M3.1 Flash Without Falling Into Traps

Grab the Token Plan during the promo, but document your costs yourself: no one will do it for you.

Three precautions before you start. One: don't confuse the Token Plan Subscription Key with a pay-as-you-go API key — they are not interchangeable, and the documentation still contradicts itself. Two: thinking cannot be disabled; adjust the effort level instead of looking for a switch that doesn't exist. Three: the promo window (double credits from September 28 to October 7, UTC+8) is the right time to test at reduced credit — but read the quota reset terms before committing a production pipeline.

To situate M3.1 Flash, our protocol comes down to three measurements, taken on the same battery of tasks: tokens consumed per complete task (not per call), end-to-end latency, and agent failure rate. We run the same protocol on Space Bunny Alpha (free) and Gemini 3.8 Flash via OpenRouter:

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "google/gemini-3.8-flash", "messages": [{"role": "user", "content": "Votre tâche de test"}]}'

Without these three numbers, any comparison between Flash models is just marketing. With them, you'll know within an hour whether the Token Plan beats your current bill.

Last point: M3.1 Flash is not downloadable. If your constraint is data sovereignty, head over to our guide to installing a local LLM and our selection of the best local LLMs — M3, for its part, remains open-weight.


❌ Common Mistakes

Mistake 1: confusing the two Token Plan credentials

The Token Plan Subscription Key and the pay-as-you-go API key are not interchangeable. Wrong key = incomprehensible authentication errors and hours wasted. Solution: check the credential type in the console before any integration, and be wary of the footnote on the pricing page, which is not yet updated with M3.1-Flash-Preview.

Mistake 2: trying to disable thinking

Any attempt returns an HTTP 400 "message requires adaptive thinking". Solution: adjust the effort level (low to max) and measure the consumption of each level — the vendor publishes no figures on this point, so the measurement is up to you.

Mistake 3: buying based on the advertised 100+ tokens/s

This is a vendor claim, not an independent measurement, and BenchLM shows 0/486 benchmarks as of September 27, 2026. Solution: wait for BenchLM coverage or run your own harness before making any serious budget commitment.

Mistake 4: comparing the "price" of M3.1 Flash to that of M3

M3 had a public rate card ($0.30/M input, $1.20/M output); M3.1 Flash has no published per-token price. The comparison is methodologically impossible. Solution: estimate the total cost of a month of Token Plan against your measured M3 consumption on an identical workload.

Mistake 5: building on the assumption "Space Bunny = free MiniMax"

The tokenizer points to MiniMax, but nothing is confirmed and the checkpoint has not been identified. Building an architecture on this assumption is an unnecessary risk. Solution: treat Space Bunny Alpha and M3.1 Flash as distinct products until MiniMax announces anything.


❓ Frequently Asked Questions

How much does MiniMax M3.1 Flash Preview cost?

No per-token pricing has been published as of September 27, 2026. The model is accessible only via the Token Plan (subscription) and MiniMax Code. A double check-in credits promo runs from September 28 to October 7 (UTC+8). For the exact subscription price, check MiniMax's pricing page — the documentation is still being updated.

Is M3.1 Flash Preview open-weight?

No. Unlike MiniMax-M3 (open-weight, June 2026), M3.1 Flash has no downloadable weights. A restricted Hugging Face repo of about 250 GB (MiniMax-M3.1-preview-private) has been spotted but remains inaccessible. No announcement of an open release has been made to date.

What context and modalities are supported?

A one-million-token context window, with text, image, and video inputs (native multimodality, according to First Financial). BenchLM does not yet officially list the window. Endpoints are available in Anthropic style (recommended by MiniMax) and OpenAI style, on the same model id.

Can you disable thinking on M3.1 Flash?

No. Any attempt returns an HTTP 400 error ("message requires adaptive thinking"). MiniMax recommends lowering the effort — five levels, from low to max, defaulting to max — rather than turning it off. This is a notable change from M3, which allowed disabling thinking.

Is M3.1 Flash better than Gemini 3.8 Flash?

Impossible to say as of September 27, 2026: no independent benchmark exists for M3.1 Flash (0/486 on BenchLM). Gemini 3.8 Flash has a clear rate card — $0.75/M input, $3.75/M output, 1,048,576 tokens of context on OpenRouter. Until measurements are available, the choice comes down to transparency, and Gemini wins.

Is Space Bunny a MiniMax model?

Unconfirmed, but the clues are strong. A tokenizer fingerprint test (50/50 strings, two passes) associates Space Bunny with the MiniMax family, and OpenRouter lists it without a named developer. Space Bunny Alpha is free, with 1M context, text/image/video, tool calling, and adjustable reasoning. MiniMax has confirmed nothing.


✅ Conclusion

M3.1 Flash Preview is the most opaque model of a generation obsessed with price transparency — and this bet on subscriptions makes it the barometer to watch for fall 2026. We will update this comparison as soon as the first benchmarks come out; in the meantime, our monthly roundup of the best LLMs lists the models that do publish their numbers.