Tool DiscoveryTool Discovery
Local AIBeginner15 min to complete12 min read

How to Run Mistral Large 4 on Ollama: Cloud Setup Guide (2026)

Mistral Large 4 (Le Chonk) isn't a local Ollama pull yet. Learn to run mistral-large-4:cloud, set up API access, and see when open weights ship.

AmaraBy Amara|Updated 10 October 2026
Terminal running ollama run mistral-large-4:cloud next to a Mistral AI badge and a cloud icon

Mistral Large 4 is Mistral AI's new flagship model, nicknamed "Le Chonk," announced on October 6, 2026. It's a mixture-of-experts design with roughly 1 trillion total parameters and 52 billion active per token, a 1,048,576 token context window, and a 1.6 billion parameter vision encoder for native image input. Open weights are promised by the end of October 2026, but as of this guide's publish date they are not out yet, which means Ollama currently serves Mistral Large 4 as a cloud-only model. `ollama run mistral-large-4:cloud` sends your prompt to Mistral's hosted infrastructure through Ollama's servers instead of loading anything on your machine.

That cloud-only status is worth sitting with for a second, because it's different from models like GLM 5.2 or Kimi K2.6, where a cloud tag exists because the weights are too large for ordinary hardware even though open weights are already public. With Mistral Large 4, there is currently no local alternative at all, licensed or not, since Mistral has not released any weights yet. The `:cloud` tag is the only way to use this model through Ollama right now.

This guide covers the full cloud setup: installing Ollama, signing in, running `mistral-large-4:cloud` from the terminal, and generating an API key for scripts and agents. The alternatives section near the end covers Mistral Medium 3.5, which you can run locally today, along with GLM 5.2 and DeepSeek R1 for other cloud and local options.

Prerequisites

  • Ollama 0.12 or later, installed on Linux, macOS, or Windows (no GPU or high-RAM machine required)
  • A free account at ollama.com for the `ollama signin` step
  • A stable internet connection, since inference runs on Mistral and Ollama's servers, not your hardware
  • Basic terminal familiarity for running `ollama run` and `curl` commands
  • (Optional) An API key from ollama.com/settings/keys if you plan to call Mistral Large 4 from your own scripts or agents

What Mistral Large 4 Is and Why Ollama Runs It in the Cloud

Mistral Large 4 is Mistral AI's newest flagship large language model, built as a mixture-of-experts with approximately 1 trillion total parameters and 52 billion active per token. Mistral trained it from scratch on 3,800 Nvidia Grace Blackwell GPUs across its own European data centers, on training data spanning more than 160 languages. The model is multimodal, pairing its language backbone with a 1.6 billion parameter vision encoder that handles image input alongside text, and Mistral positions it for coding, long-running agentic workflows, and enterprise domains including cybersecurity, finance, and legal work.

SpecValue
Total parameters~1 trillion (MoE)
Active parameters52 billion per token
Context window1,048,576 tokens
Max output262,144 tokens
Vision encoder1.6B parameters
Training hardware3,800 Nvidia Grace Blackwell GPUs
Ollama tag`mistral-large-4:cloud` (only option)
ℹ️
Note:Mistral has promised open weights for Large 4 by the end of October 2026, but as of this guide's publish date, no weights or license terms are public. Unlike GLM 5.2 or Kimi K2.6, where cloud tags exist because the published weights are too large for consumer hardware, there is currently no local version of Mistral Large 4 to run anywhere, through Ollama or otherwise.

On early third-party benchmarking from Artificial Analysis, the preview build scored 38 on the Intelligence Index, behind GLM 5.3 (45) and Kimi K3 (44), while posting a stronger 50 on the Cyber Index with 82% on a reproduce-then-patch security test and 93% on Cybench. Treat these as early preview numbers rather than final scores, since Mistral is still actively tuning the model ahead of the full release.

During the preview period, API pricing runs at $0.68 per million input tokens and $2.09 per million output tokens, half of the eventual list price of $1.36 input and $4.18 output. That API pricing applies to Mistral's own endpoints directly; the Ollama Cloud route covered in this guide works within Ollama's own free usage limits, separate from Mistral's direct API billing.

Set Up Ollama Cloud and Run Mistral Large 4

Running Mistral Large 4 through Ollama takes three steps: install Ollama, sign in, and run the model. Nothing here downloads a large file, since inference happens on Mistral's infrastructure rather than your disk.

Step 1: Install Ollama

# Linux and macOS, one-command installer
curl -fsSL https://ollama.com/install.sh | sh

On Windows, download the installer from ollama.com/download, or use winget:

powershell
winget install Ollama.Ollama

Verify the installation:

ollama --version
ollama version is 0.12.3
⚠️
Warning:Cloud models require Ollama 0.6.x or later. If your version is older, re-run the install command above to update before continuing.

Step 2: Sign In to Ollama Cloud

ollama signin

This prints a sign-in URL and opens your browser. Create a free account at ollama.com, or log in if you already have one, then approve the device.

Signing in to ollama.com...
Signed in as your-username
ℹ️
Note:No payment information is required for cloud models within Ollama's free usage limits as of October 2026. Check ollama.com/settings for current limits, since these change over time.

Step 3: Run Mistral Large 4 from the Terminal

ollama run mistral-large-4:cloud

Ollama fetches a small manifest rather than model weights, then drops you into a prompt:

pulling manifest
pulling 66c77f50c28c... 100% ▕████████████████▏  4.1 KB
success
>>> Send a message (/? for help)

Test it with a prompt:

>>> Write a function that validates an email address with a regex, and explain the pattern.

The first response after signing in can take 10-20 seconds while Ollama establishes the cloud session. Responses after that stream back at normal speed.

Step 4: Verify the Model

ollama list

`mistral-large-4:cloud` shows up at a few KB, not hundreds of gigabytes. That's expected for a cloud passthrough entry.

💡
Tip:Mistral Large 4 accepts image input too. Drop a file path into your prompt, like `>>> What's shown in this chart? ./chart.png`, and Ollama attaches it automatically through the model's 1.6B vision encoder.

Use Mistral Large 4 in Your Own Scripts and Agents

Beyond the interactive `ollama run` session, Mistral Large 4's cloud tag is reachable through Ollama's REST API, so any tool that already talks to a local Ollama instance can use it with a one-line model name change.

Generate an API Key

Visit ollama.com/settings/keys while signed in, click "Create API key," and copy the value.

export OLLAMA_API_KEY=your_api_key_here
💡
Tip:An API key is only needed for direct requests to `https://ollama.com/api`. If your application talks to `localhost:11434`, the standard local Ollama server, `ollama signin` already authenticated that machine and no separate key is required.

Call Mistral Large 4 from the Local Endpoint

curl http://localhost:11434/api/chat -d '{
  "model": "mistral-large-4:cloud",
  "messages": [
    { "role": "user", "content": "Summarize the OWASP Top 10 in five bullet points." }
  ],
  "stream": false
}'

Python Example

python
from ollama import Client

client = Client(host="https://ollama.com", headers={"Authorization": "Bearer " + api_key})

response = client.chat(
    model="mistral-large-4:cloud",
    messages=[{"role": "user", "content": "Plan a migration from REST to GraphQL for a mid-size API."}],
)
print(response["message"]["content"])

OpenAI-Compatible Endpoint for Existing Agent Tools

Ollama exposes an OpenAI-compatible layer at `http://localhost:11434/v1`, the same endpoint used in the Hermes Agent and OpenClaw setups. Point that configuration at `mistral-large-4:cloud` instead of a local model name:

yaml
model:
  default: mistral-large-4:cloud
  provider: custom
  base_url: http://localhost:11434/v1
  context_length: 1048576

The agent then runs on Mistral Large 4's full 1M-plus context window without any other configuration changes.

Troubleshooting

`ollama pull mistral-large-4` fails with "pull model manifest: file does not exist"

Cause: Ollama's official library only hosts `mistral-large-4:cloud`, with no plain or quantized local tag

Fix: Use `ollama run mistral-large-4:cloud` with the `:cloud` suffix included. There is currently no local tag to pull for this model.

`ollama run mistral-large-4:cloud` returns "model not found"

Cause: The installed Ollama version predates cloud model support

Fix: Update Ollama by re-running the install command, then retry. Cloud models require Ollama 0.6.x or later.

"unauthorized" error or repeated sign-in prompts

Cause: The machine is not signed in, or the session expired

Fix: Run `ollama signin` again and complete the browser approval. Check ollama.com/settings/connections to confirm the device is listed as connected.

First response takes 10-20 seconds or longer

Cause: Cold start while Ollama establishes a session with Mistral's cloud infrastructure

Fix: This is normal for the first request after signing in or after an idle period. Subsequent requests in the same session stream back at normal speed.

API requests to `https://ollama.com/api` return 401

Cause: Missing or invalid `OLLAMA_API_KEY`

Fix: Generate a new key at ollama.com/settings/keys and re-export the environment variable.

Image input in a prompt is ignored

Cause: The file path was not detected, often from a typo or an unsupported image format

Fix: Use an absolute path to a JPEG or PNG file and confirm the file exists before sending the prompt.

Alternatives to Consider

ToolTypePriceBest For
Mistral Medium 3.5Local (Ollama)Free, local installA 128B dense model you can actually download and run today, if you have 80GB-plus of combined VRAM and RAM.
GLM 5.2 via Ollama CloudCloud (Ollama)Free within Ollama Cloud limitsA larger 744B cloud model with a 1M context window and effort-level control, also cloud-only through Ollama.
Kimi K2.6 via Ollama CloudCloud (Ollama)Free within Ollama Cloud limitsAnother cloud-only flagship, with strong multi-agent and tool-use performance.
DeepSeek R1Local (Ollama) or VPSFreeA true local reasoning model on hardware from 4GB (distilled) up to 64GB or more, with no cloud dependency.

Frequently Asked Questions

Can I run Mistral Large 4 locally with Ollama?

Not yet. Mistral has not released open weights for Large 4 as of this guide's publish date, so there is no local tag to pull through Ollama or any other tool. Mistral has said open weights are coming by the end of October 2026, at which point a local option may appear.

Until then, `mistral-large-4:cloud` through Ollama is the only way to use this model with Ollama's command syntax and ecosystem. If you need something you can run on your own hardware today, see Mistral Medium 3.5 in the alternatives section.

Is Mistral Large 4 free to use through Ollama?

Yes, within Ollama's free usage limits as of October 2026. `ollama signin` does not require payment information, and `ollama run mistral-large-4:cloud` works immediately after signing in.

Mistral's own direct API is separately priced at $0.68 per million input tokens and $2.09 per million output tokens during the preview period, half of the eventual list price of $1.36 input and $4.18 output. That direct API pricing is not required for the Ollama Cloud setup in this guide.

What does the "Le Chonk" nickname mean?

"Le Chonk" is Mistral's own informal nickname for Large 4, a reference to its size, roughly 1 trillion total parameters. It's not an official model name you need to use in commands; `mistral-large-4` is the actual identifier Ollama and Mistral's API use.

How does Mistral Large 4 compare to Mistral Medium 3.5?

Mistral Large 4 is far larger, roughly 1 trillion total parameters against Medium 3.5's 128 billion, and uses a mixture-of-experts design (52B active per token) rather than Medium 3.5's fully dense architecture. Large 4 also has a much bigger context window, just over 1 million tokens against Medium 3.5's 256K.

The tradeoff is availability. Medium 3.5 has public weights you can download and run locally through Ollama today. Large 4 is cloud-only until Mistral ships open weights, expected by the end of October 2026.

When will Mistral Large 4 open weights be available?

Mistral has said open weights for Large 4 are coming by the end of October 2026. As of this guide's publish date, no weights or license terms have been published, so the exact date and license are not yet confirmed.

Check Mistral's official announcements and Hugging Face page for the actual release once it happens. This guide will be updated with local installation steps once weights are public.

How does Mistral Large 4 compare to GLM 5.3 and Kimi K3 on benchmarks?

On Artificial Analysis's early preview testing, Mistral Large 4 scored 38 on the Intelligence Index, behind GLM 5.3 (45) and Kimi K3 (44). It performed more competitively on security-focused testing, scoring 50 on the Cyber Index with 82% on a reproduce-then-patch test and 93% on Cybench.

These are preview-stage numbers, and Mistral is still actively tuning the model ahead of its full release, so treat them as a starting point rather than a final comparison.

Does Mistral Large 4 support image input?

Yes. Mistral Large 4 ships with a 1.6 billion parameter vision encoder for native multimodal input alongside text. Through Ollama, drop an image file path into your prompt and Ollama attaches it automatically, the same mechanism used by any vision-capable model.

Can I use Mistral Large 4 with an agent like Hermes Agent or OpenClaw?

Yes. Both Hermes Agent and OpenClaw connect to Ollama's OpenAI-compatible endpoint at `http://localhost:11434/v1`. Point the agent's model configuration at `mistral-large-4:cloud` and set `context_length` to 1048576.

The agent then runs on Mistral Large 4's full context window through your existing Ollama setup, with no other configuration changes needed.

Do I need a VPS or GPU to use Mistral Large 4 with Ollama?

No. Inference for the `:cloud` tag runs on Mistral's infrastructure regardless of where you run `ollama`, so a laptop or desktop with no GPU is enough. There is currently no local install path for this model at all, since open weights are not yet public.

Related Guides