Mistral Large 4 is live on CloudCode.ONE — and it reads your screenshots itself
Mistral released Mistral Large 4 on October 6, and it's on CloudCode.ONE from today as mistral-large-4. It's already enabled on every key — no new key, no waitlist, nothing to switch on. Name it in your config and go.
The short version: a trillion-parameter model that reads your screenshots itself, answers without a long wait, and costs a fraction of the other big model on our list that can do the same.
What Mistral says it can do
Mistral calls it "our largest and most capable model to date": 1 trillion parameters with 52 billion active per token, natively multimodal, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters. It's open-weight — the API preview is live now, and Mistral says the weights follow by the end of October.
On agentic coding, Mistral reports:
| Benchmark |
Mistral Large 4 |
| DeepSWE v1.1 |
61.7% |
| SWE-Atlas-QnA |
59.4% |
| Terminal-Bench 4 |
28.3% |
| Coding Agent Index (combined) |
49.8% |
That Coding Agent Index score, Mistral says, puts it ahead of DeepSeek V4 Pro (0813) and Qwen3.8 Max.
The number we'd look at first is the least glamorous one. Mistral had professional annotators rate coding output blind, on a 1–5 scale, with model names hidden. Large 4 came second of five with 3.74 — ahead of GLM-5.3 (3.60) and Kimi K3 (3.59), behind only Claude Opus 5 (4.22). Both of those run on CloudCode.ONE today, so it's a comparison you can check against your own repo rather than take on trust.
Two more claims stand out:
- Security work. On a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, Mistral reports 82%, the highest of any model, plus 93% of Cybench's 40 competition challenges.
- Vision. Mistral calls Large 4 "a step change" in how its models understand images, and reports it edging past GPT-6-Astra on the Dense 200 visual-grounding benchmark (42% vs 41%).
These are Mistral's numbers, not ours, and Mistral doesn't say which reasoning setting it used. Treat them as a reason to try it, not as a verdict.
What it's like on CloudCode.ONE
We ran it through real Claude Code sessions before switching it on. Here's what to expect.
It reads images itself. Most strong open coding models are text-only, which is why the gateway quietly hands image requests for glm-5.3 or deepseek-v4-pro to a vision model. Mistral Large 4 doesn't need that: the screenshot of the broken layout, the error dialog, the chart you want read back as data all go straight into the model that writes the fix. The usual small per-image surcharge applies.
It answers straight away. Left to itself, Large 4 thinks before every answer — on some prompts for minutes, with every token of it billed as output. So on CloudCode.ONE it runs with reasoning off by default, and answers start within a second or two. When you want it to think first, send reasoning_effort: "high" on the OpenAI-compatible endpoint; the thinking bills as output, like any reasoning model.
It handles the agent loop. Write a file, run it, read the output, fix, repeat — including several tool calls in one turn: multi-step tool use worked cleanly in our Claude Code tests.
Long sessions run mostly from cache. An agent resends the whole conversation on every turn, so caching decides what a long session costs. We route each account's requests so its context stays warm in the cache — nothing to configure — and cached input bills at about a tenth of the normal rate. In our Claude Code test sessions, every turn after the first read 94–97% of its prompt from cache: the first turn of a 15K-token session cost about 2 cents, the turns after it about a third of a cent each.
Price, and how it compares
| Model |
Input / 1M |
Cached / 1M |
Output / 1M |
Context |
Images |
mistral-large-4 |
$1.36 |
$0.14 |
$4.18 |
512K |
Native |
glm-5.3 |
$1.26 |
$0.234 |
$3.96 |
1M |
Via a vision model |
deepseek-v4-pro |
$1.705 |
$0.0568 |
$5.115 |
1M |
Via a vision model |
kimi-k3 |
$3.90 |
$0.39 |
$19.50 |
1M |
Native |
Against Kimi K3, the other big model here that sees for itself, it's roughly a third of the price on input and a fifth on output. Against GLM-5.3 it costs a few percent more per token, and 40% less per cached token.
The trade-off is context. Mistral's model card lists 1M tokens, but the preview serves 512K, so that's the number we quote.
Getting started
Same key, same endpoints. In Claude Code, put this in ~/.claude/settings.json:
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.cloudcode.one",
"ANTHROPIC_AUTH_TOKEN": "<your CloudCode.ONE API key>",
"ANTHROPIC_MODEL": "mistral-large-4",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "mistral-large-4",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "mistral-large-4",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash",
"CLAUDE_CODE_SUBAGENT_MODEL": "glm-5.3-flash",
"CLAUDE_CODE_EFFORT_LEVEL": "max"
}
}
Three details. Use the plain mistral-large-4, without the [1m] suffix our GLM instructions use: that suffix tells Claude Code to plan for a 1M window, and without it Claude Code compacts long sessions well inside Large 4's 512K. Keeping the haiku tier and subagents on glm-5.3-flash keeps Claude Code's background work cheap. And Claude Code's effort level doesn't switch Large 4's reasoning on, so you won't see thinking blocks from it — that's the fast default, not a fault.
For Cline, Roo Code, Continue, OpenCode or Pi, choose "OpenAI Compatible", set the base URL to https://api.cloudcode.one/v1, and use model mistral-large-4. Or call it straight from code — with a screenshot, since that's the point:
import base64
from openai import OpenAI
client = OpenAI(api_key="<your CloudCode.ONE API key>", base_url="https://api.cloudcode.one/v1")
with open("broken-layout.png", "rb") as f:
screenshot = base64.b64encode(f.read()).decode()
reply = client.chat.completions.create(
model="mistral-large-4",
# reasoning_effort="high", # let it think first; thinking bills as output
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "The sidebar overlaps the header on mobile. What's causing it?"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot}"}},
],
}],
)
print(reply.choices[0].message.content)
The same key works in Harness, our desktop agent for Windows and Ubuntu.
Good to know
- It's a preview. Mistral says Large 4 "continues to improve rapidly as we refine it". We pin the exact 4.0 release, so a newer version won't be swapped in under you — but expect 4.0 itself to keep getting better during the preview.
- Reasoning is off unless you ask. Send
reasoning_effort: "high" to turn it on; thinking tokens bill as output.
- Nothing else changed. Pay-as-you-go, no subscription, credit never expires, and the free tier is untouched.
The full price list is always on the front page.
New here? Top up from $2 and a key is provisioned instantly.