Mistral opened a public preview of Mistral Large 4 on 6 October, a 1.05-trillion-parameter mixture-of-experts model with 49 billion active parameters and a one-million-token context window. Full open weights are scheduled for later this month. For anyone paying per-token to OpenAI, Anthropic, or Google to run a chatbot, internal search tool, or document pipeline, this is the first real checkpoint this year to re-run the self-hosted math.
What is Mistral Large 4?#
Mistral Large 4 is a multimodal mixture-of-experts model — 1.05 trillion total parameters, 49 billion active per inference, a 1.6-billion-parameter vision encoder, and a context window of one million tokens, according to Mistral's own model documentation, as reported by Let's Data Science. It supports structured outputs, function calling, document question-answering, and built-in tools, and Mistral is pitching it at coding, agent, and enterprise workloads.
Here's the part worth slowing down on: what's live right now is a hosted preview. The downloadable, self-hostable weights aren't out yet — Mistral says later in October. At WebEpex we've built client chatbot and automation infrastructure on self-hosted VPS stacks for businesses across the GCC and Europe since before this release, which is exactly why we read every "open" model announcement the week it drops instead of the month after. A model launches with the word "open" in the headline, and the actual weights show up weeks later. Until Mistral Large 4's weights are sitting on Hugging Face, it's a preview, not an open-weight model.
Who does this actually matter to?#
If you're sending under a few hundred thousand tokens a month through a chatbot or internal tool, this changes nothing. Hosted pricing from OpenAI, Anthropic, or Google is still cheaper and far less work than running your own inference box, and you can stop reading here.
If you're above that — a WhatsApp bot handling thousands of conversations a month, or a document pipeline chewing through long contracts and reports — Large 4's million-token window is worth modelling against your current SaaS build cost once real weights and real GPU pricing both exist. Not before. A preview API with no published pricing tells you nothing about what self-hosting will actually cost in GPU time, and we've gotten that math wrong once already on a different model, which is why we wait for the real numbers now.
The real trade-offs of self-hosting a model this size#
A trillion-parameter MoE model, even with 49 billion active parameters, isn't a laptop job. The honest trade-offs:
- You need GPU infrastructure capable of serving a model this size, which most SMBs don't keep sitting idle
- Data stays inside your own infrastructure — the actual draw for GCC and European clients with residency terms in their contracts
- No per-token bill that scales against you as volume grows
- You own the ops burden: updates, scaling, uptime, and security patching, same as any self-hosted agent setup
We don't think this replaces hosted APIs for most businesses. It's a tool for the narrow slice that's both high-volume and genuinely particular about where their data sits.
How we're pricing this into client builds#
At WebEpex we build client WhatsApp and chatbot automation on self-hosted infrastructure — PM2, Nginx, PostgreSQL, n8n — for clients across the GCC and Europe who care where their data lives. We read Large 4's spec sheet the day the preview went live, and our rule hasn't changed: we don't move a client's production flow onto a model that only exists as a hosted preview, however good the numbers look on paper.
The actual decision point isn't 6 October. It's whenever the real, downloadable weights land and someone can finally price a GPU box against an OpenAI invoice. That's the comparison we're waiting to run, and we'll run it the week the weights are out.
What I'd tell a client asking about this this week#
Do nothing yet. There's no published self-hosted pricing and the weights aren't out, so any number you'd model today is a guess. If you're already paying four figures a month or more for a hosted AI API and your contract has data-residency terms, put a reminder on your calendar for late October and ask for the comparison then. If you're under that spend, this isn't worth a meeting yet.
If you want a second opinion on what you're actually paying for AI right now versus what you could be, send me what you're running and I'll tell you straight — takes two minutes, no pitch. [cal.com/webepex/growth-review]