MiniMax M3: The Open-Weight Model You Can Actually Run Yourself
While Kimi K3 requires 64+ accelerators and Qwen3.8-Max's weights still hadn't landed, one open-weight model stood out as actually practical to self-host in August 2026. But the landscape has since shifted.
🆕 Update, Aug 9: Qwen3.8-27B is out now!
Alibaba has released Qwen3.8-27B as open weights on Hugging Face. This is the "smaller" sibling of Qwen3.8-Max (2.4T), but still a powerful model that can run on far more modest hardware.
| Model | Status | VRAM requirement (INT4) | Use case |
|---|---|---|---|
| Qwen3.8-27B | ✅ Out now (HF) | ~14 GB | Small teams/enthusiasts – runs on 1-2× consumer GPU |
| MiniMax M3 | ✅ Out, #1 on BenchLM | ~230 GB | Enterprise teams with a GPU pool |
| Kimi K3 (2.8T) | ✅ Out, with license | 64+ accelerators | Hyperscalers/states |
| Qwen3.8-Max (2.4T) | ⏳ Expected soon | TBD (~500+ GB) | Still unknown |
What is Qwen3.8-27B good at?
- ✅ Code generation and understanding
- ✅ Multilingual tasks (Danish included)
- ✅ Mathematical reasoning
- ✅ Agentic workflows (Alibaba specifically named "agentic workloads" as a focus area)
With only around 14 GB of VRAM needed at INT4 quantization, a single team — or even a hobbyist — can run it on consumer hardware (2× RTX 4090 24GB) or a couple of used A100/H100 cards. That makes it the perfect "sweet spot" between performance and practical usability.
Where MiniMax M3 excels at general reasoning and long context (256K tokens), Qwen3.8-27B particularly shines at coding, math, and agents.
📊 What is MiniMax M3?
| Property | Detail |
|---|---|
| Model type | MoE (Mixture of Experts) – hybrid architecture |
| Total parameters | ~456B (estimated) |
| Active per token | ~45B (about 10% of total weights) |
| Context window | 256K tokens |
| License | Commercial use permitted (not OSI open source) |
August 2026: MiniMax just topped BenchLM's August ranking as the best open-weight model — confirmed by two independent sources with a score of 68.8.
🧩 Where do they fit?
The open-weights movement has come of age — we now have models covering the entire spectrum:
Want maximum performance AND have access to a supernode? → Kimi K3 or Qwen3.* awaits
Want strong enterprise-grade performance at reasonable requirements? → MiniMax M3 is ready
Want something you can run locally without selling your house? → Qwen*.*B just shipped
That's exactly why the managed self-hosting concept makes sense — not everyone should or wants to own the hardware themselves, but everyone deserves access to the best open models under responsible conditions.
When "open" isn't quite open
Hidden license terms hit the big cloud providers hardest — but small EU hosting providers can navigate them for you.
The EU trilemma
Chinese model? Expensive hardware? Or a dumbed-down, lower-quality solution? With Liviate you don't have to choose — we host the best open models on European infrastructure with full compliance.
Alibaba vs. DeepSeek: China's two-front AI war
While Alibaba is betting broad with the Qwen family, from the small 27B up to the giant Max, DeepSeek is focusing on efficiency and reasoning with its V4 Pro series.
Mistral Shieldstral: a safety guard for images and text
Mistral has launched Shieldstral — a safety filter designed specifically for open-weight models in enterprise environments.
GPU hoarding
Four companies are stockpiling billions in Nvidia GPU chips — Meta alone owns more than 500,000 H100-equivalents.
Conclusion
The open-weights movement has reached a level of maturity where the discussion is no longer about whether open models can compete with closed frontier models — the answer is yes.
The question now is: who can actually run them responsibly in practice?
The way forward:
- Models large enough for serious workloads
- Models small enough for practical use
- A combined, neutral host that ensures compliance, isolation, and support — so customers can skip the multi-million-dollar investments
Freedom + control + infrastructure = Liviate.