How to Run Mistral Large 4 on Ollama: Cloud Setup Guide (2026)
Mistral Large 4 (Le Chonk) isn't a local Ollama pull yet. Learn to run mistral-large-4:cloud, set up API access, and see when open weights ship.

Mistral Large 4 is Mistral AI's new flagship model, nicknamed "Le Chonk," announced on October 6, 2026. It's a mixture-of-experts design with roughly 1 trillion total parameters and 52 billion active per token, a 1,048,576 token context window, and a 1.6 billion parameter vision encoder for native image input. Open weights are promised by the end of October 2026, but as of this guide's publish date they are not out yet, which means Ollama currently serves Mistral Large 4 as a cloud-only model. `ollama run mistral-large-4:cloud` sends your prompt to Mistral's hosted infrastructure through Ollama's servers instead of loading anything on your machine.
That cloud-only status is worth sitting with for a second, because it's different from models like GLM 5.2 or Kimi K2.6, where a cloud tag exists because the weights are too large for ordinary hardware even though open weights are already public. With Mistral Large 4, there is currently no local alternative at all, licensed or not, since Mistral has not released any weights yet. The `:cloud` tag is the only way to use this model through Ollama right now.
This guide covers the full cloud setup: installing Ollama, signing in, running `mistral-large-4:cloud` from the terminal, and generating an API key for scripts and agents. The alternatives section near the end covers Mistral Medium 3.5, which you can run locally today, along with GLM 5.2 and DeepSeek R1 for other cloud and local options.
Prerequisites
- Ollama 0.12 or later, installed on Linux, macOS, or Windows (no GPU or high-RAM machine required)
- A free account at ollama.com for the `ollama signin` step
- A stable internet connection, since inference runs on Mistral and Ollama's servers, not your hardware
- Basic terminal familiarity for running `ollama run` and `curl` commands
- (Optional) An API key from ollama.com/settings/keys if you plan to call Mistral Large 4 from your own scripts or agents
In This Guide
What Mistral Large 4 Is and Why Ollama Runs It in the Cloud
Mistral Large 4 is Mistral AI's newest flagship large language model, built as a mixture-of-experts with approximately 1 trillion total parameters and 52 billion active per token. Mistral trained it from scratch on 3,800 Nvidia Grace Blackwell GPUs across its own European data centers, on training data spanning more than 160 languages. The model is multimodal, pairing its language backbone with a 1.6 billion parameter vision encoder that handles image input alongside text, and Mistral positions it for coding, long-running agentic workflows, and enterprise domains including cybersecurity, finance, and legal work.
| Spec | Value |
|---|---|
| Total parameters | ~1 trillion (MoE) |
| Active parameters | 52 billion per token |
| Context window | 1,048,576 tokens |
| Max output | 262,144 tokens |
| Vision encoder | 1.6B parameters |
| Training hardware | 3,800 Nvidia Grace Blackwell GPUs |
| Ollama tag | `mistral-large-4:cloud` (only option) |
On early third-party benchmarking from Artificial Analysis, the preview build scored 38 on the Intelligence Index, behind GLM 5.3 (45) and Kimi K3 (44), while posting a stronger 50 on the Cyber Index with 82% on a reproduce-then-patch security test and 93% on Cybench. Treat these as early preview numbers rather than final scores, since Mistral is still actively tuning the model ahead of the full release.
During the preview period, API pricing runs at $0.68 per million input tokens and $2.09 per million output tokens, half of the eventual list price of $1.36 input and $4.18 output. That API pricing applies to Mistral's own endpoints directly; the Ollama Cloud route covered in this guide works within Ollama's own free usage limits, separate from Mistral's direct API billing.
Set Up Ollama Cloud and Run Mistral Large 4
Running Mistral Large 4 through Ollama takes three steps: install Ollama, sign in, and run the model. Nothing here downloads a large file, since inference happens on Mistral's infrastructure rather than your disk.
Step 1: Install Ollama
# Linux and macOS, one-command installer
curl -fsSL https://ollama.com/install.sh | shOn Windows, download the installer from ollama.com/download, or use winget:
winget install Ollama.OllamaVerify the installation:
ollama --versionollama version is 0.12.3Step 2: Sign In to Ollama Cloud
ollama signinThis prints a sign-in URL and opens your browser. Create a free account at ollama.com, or log in if you already have one, then approve the device.
Signing in to ollama.com...
Signed in as your-usernameStep 3: Run Mistral Large 4 from the Terminal
ollama run mistral-large-4:cloudOllama fetches a small manifest rather than model weights, then drops you into a prompt:
pulling manifest
pulling 66c77f50c28c... 100% ▕████████████████▏ 4.1 KB
success
>>> Send a message (/? for help)Test it with a prompt:
>>> Write a function that validates an email address with a regex, and explain the pattern.The first response after signing in can take 10-20 seconds while Ollama establishes the cloud session. Responses after that stream back at normal speed.
Step 4: Verify the Model
ollama list`mistral-large-4:cloud` shows up at a few KB, not hundreds of gigabytes. That's expected for a cloud passthrough entry.
Use Mistral Large 4 in Your Own Scripts and Agents
Beyond the interactive `ollama run` session, Mistral Large 4's cloud tag is reachable through Ollama's REST API, so any tool that already talks to a local Ollama instance can use it with a one-line model name change.
Generate an API Key
Visit ollama.com/settings/keys while signed in, click "Create API key," and copy the value.
export OLLAMA_API_KEY=your_api_key_hereCall Mistral Large 4 from the Local Endpoint
curl http://localhost:11434/api/chat -d '{
"model": "mistral-large-4:cloud",
"messages": [
{ "role": "user", "content": "Summarize the OWASP Top 10 in five bullet points." }
],
"stream": false
}'Python Example
from ollama import Client
client = Client(host="https://ollama.com", headers={"Authorization": "Bearer " + api_key})
response = client.chat(
model="mistral-large-4:cloud",
messages=[{"role": "user", "content": "Plan a migration from REST to GraphQL for a mid-size API."}],
)
print(response["message"]["content"])OpenAI-Compatible Endpoint for Existing Agent Tools
Ollama exposes an OpenAI-compatible layer at `http://localhost:11434/v1`, the same endpoint used in the Hermes Agent and OpenClaw setups. Point that configuration at `mistral-large-4:cloud` instead of a local model name:
model:
default: mistral-large-4:cloud
provider: custom
base_url: http://localhost:11434/v1
context_length: 1048576The agent then runs on Mistral Large 4's full 1M-plus context window without any other configuration changes.
Troubleshooting
`ollama pull mistral-large-4` fails with "pull model manifest: file does not exist"
Cause: Ollama's official library only hosts `mistral-large-4:cloud`, with no plain or quantized local tag
Fix: Use `ollama run mistral-large-4:cloud` with the `:cloud` suffix included. There is currently no local tag to pull for this model.
`ollama run mistral-large-4:cloud` returns "model not found"
Cause: The installed Ollama version predates cloud model support
Fix: Update Ollama by re-running the install command, then retry. Cloud models require Ollama 0.6.x or later.
"unauthorized" error or repeated sign-in prompts
Cause: The machine is not signed in, or the session expired
Fix: Run `ollama signin` again and complete the browser approval. Check ollama.com/settings/connections to confirm the device is listed as connected.
First response takes 10-20 seconds or longer
Cause: Cold start while Ollama establishes a session with Mistral's cloud infrastructure
Fix: This is normal for the first request after signing in or after an idle period. Subsequent requests in the same session stream back at normal speed.
API requests to `https://ollama.com/api` return 401
Cause: Missing or invalid `OLLAMA_API_KEY`
Fix: Generate a new key at ollama.com/settings/keys and re-export the environment variable.
Image input in a prompt is ignored
Cause: The file path was not detected, often from a typo or an unsupported image format
Fix: Use an absolute path to a JPEG or PNG file and confirm the file exists before sending the prompt.
Alternatives to Consider
| Tool | Type | Price | Best For |
|---|---|---|---|
| Mistral Medium 3.5 | Local (Ollama) | Free, local install | A 128B dense model you can actually download and run today, if you have 80GB-plus of combined VRAM and RAM. |
| GLM 5.2 via Ollama Cloud | Cloud (Ollama) | Free within Ollama Cloud limits | A larger 744B cloud model with a 1M context window and effort-level control, also cloud-only through Ollama. |
| Kimi K2.6 via Ollama Cloud | Cloud (Ollama) | Free within Ollama Cloud limits | Another cloud-only flagship, with strong multi-agent and tool-use performance. |
| DeepSeek R1 | Local (Ollama) or VPS | Free | A true local reasoning model on hardware from 4GB (distilled) up to 64GB or more, with no cloud dependency. |
Frequently Asked Questions
Can I run Mistral Large 4 locally with Ollama?
Not yet. Mistral has not released open weights for Large 4 as of this guide's publish date, so there is no local tag to pull through Ollama or any other tool. Mistral has said open weights are coming by the end of October 2026, at which point a local option may appear.
Until then, `mistral-large-4:cloud` through Ollama is the only way to use this model with Ollama's command syntax and ecosystem. If you need something you can run on your own hardware today, see Mistral Medium 3.5 in the alternatives section.
Is Mistral Large 4 free to use through Ollama?
Yes, within Ollama's free usage limits as of October 2026. `ollama signin` does not require payment information, and `ollama run mistral-large-4:cloud` works immediately after signing in.
Mistral's own direct API is separately priced at $0.68 per million input tokens and $2.09 per million output tokens during the preview period, half of the eventual list price of $1.36 input and $4.18 output. That direct API pricing is not required for the Ollama Cloud setup in this guide.
What does the "Le Chonk" nickname mean?
"Le Chonk" is Mistral's own informal nickname for Large 4, a reference to its size, roughly 1 trillion total parameters. It's not an official model name you need to use in commands; `mistral-large-4` is the actual identifier Ollama and Mistral's API use.
How does Mistral Large 4 compare to Mistral Medium 3.5?
Mistral Large 4 is far larger, roughly 1 trillion total parameters against Medium 3.5's 128 billion, and uses a mixture-of-experts design (52B active per token) rather than Medium 3.5's fully dense architecture. Large 4 also has a much bigger context window, just over 1 million tokens against Medium 3.5's 256K.
The tradeoff is availability. Medium 3.5 has public weights you can download and run locally through Ollama today. Large 4 is cloud-only until Mistral ships open weights, expected by the end of October 2026.
When will Mistral Large 4 open weights be available?
Mistral has said open weights for Large 4 are coming by the end of October 2026. As of this guide's publish date, no weights or license terms have been published, so the exact date and license are not yet confirmed.
Check Mistral's official announcements and Hugging Face page for the actual release once it happens. This guide will be updated with local installation steps once weights are public.
How does Mistral Large 4 compare to GLM 5.3 and Kimi K3 on benchmarks?
On Artificial Analysis's early preview testing, Mistral Large 4 scored 38 on the Intelligence Index, behind GLM 5.3 (45) and Kimi K3 (44). It performed more competitively on security-focused testing, scoring 50 on the Cyber Index with 82% on a reproduce-then-patch test and 93% on Cybench.
These are preview-stage numbers, and Mistral is still actively tuning the model ahead of its full release, so treat them as a starting point rather than a final comparison.
Does Mistral Large 4 support image input?
Yes. Mistral Large 4 ships with a 1.6 billion parameter vision encoder for native multimodal input alongside text. Through Ollama, drop an image file path into your prompt and Ollama attaches it automatically, the same mechanism used by any vision-capable model.
Can I use Mistral Large 4 with an agent like Hermes Agent or OpenClaw?
Yes. Both Hermes Agent and OpenClaw connect to Ollama's OpenAI-compatible endpoint at `http://localhost:11434/v1`. Point the agent's model configuration at `mistral-large-4:cloud` and set `context_length` to 1048576.
The agent then runs on Mistral Large 4's full context window through your existing Ollama setup, with no other configuration changes needed.
Do I need a VPS or GPU to use Mistral Large 4 with Ollama?
No. Inference for the `:cloud` tag runs on Mistral's infrastructure regardless of where you run `ollama`, so a laptop or desktop with no GPU is enough. There is currently no local install path for this model at all, since open weights are not yet public.
Related Guides
How to Run Mistral Medium 3.5 Locally with Ollama (2026 Guide)
How to Run GLM 5.2 on Ollama: Cloud Setup Guide (2026)
How to Run Kimi K2 on Ollama: Cloud Setup Guide (2026)
How to Run DeepSeek R1 Locally with Ollama (2026 Guide)
Best Local LLM Models to Run in 2026 (Benchmarks + Use Cases)
How to Use Ollama with Python: API Integration Tutorial (2026)
How to Install Hermes Agent with Ollama Local Models (2026)