Can I run Magistral Small 1.2?
Magistral Small 1.2 by Mistral AI needs around 24 GB of RAM at the recommended 4-bit quantization (14.6 GB download). Your hardware is checked below β instantly, nothing leaves your browser. Expect roughly ~24 tok/s on a Apple M-series Max.
Reading your hardware signalsβ¦
Real-world notes
Magistral Small 1.2 is Mistral's 24B reasoning model, and it is aimed at people who want a local assistant that actually thinks through multi-step problems rather than just chatting. It also handles vision and general conversation, but reasoning is the reason to reach for it. The footprint is the first thing to plan around: at a 4-bit quant it lands near 14.6 GB, and you want roughly 24 GB of memory to run it comfortably. That rules out a 12 GB card like an RTX 3060, where it simply does not fit, and points you toward a 24 GB GPU or a higher-memory Apple Silicon Mac.
On an RTX 4090 it streams at around 59 tokens per second, which is quick enough that its step-by-step reasoning never feels like waiting. On an M-series Max you are closer to 24 tokens per second, still perfectly usable for interactive work, and CPU-only on DDR5 drops to about 4 tokens per second, fine for batch jobs but not live chat. The 128K context is genuine, but it is memory-hungry: fill it and total usage climbs to roughly 42 GB, well past what a single 24 GB card holds, so keep working context modest unless you have the headroom.
Against its siblings, Mistral Nemo 12B is the lighter, faster pick if you mainly want chat and cannot spare the memory, while Gemma 4 26B A4B generally competes more directly on reasoning, coding, and vision at a similar size. Magistral's standout trait is that reasoning focus in a model you can fully own: the Apache 2.0 license means you can use it commercially and in production with no provider strings attached, which is rare for a capable 24B reasoner. If you have the 24 GB to feed it, it is one of the more serious local thinking models available.
Specifications
Size by quantization
| Quantization | Bits/weight | Download | Min RAM | Quality |
|---|---|---|---|---|
| Q2_K | 3.35 | 10.1 GB | 16 GB | Noticeable loss |
| Q4_K_MRecommended | 4.85 | 14.6 GB | 24 GB | Recommended |
| Q5_K_M | 5.65 | 17.0 GB | 24 GB | High |
| Q8_0 | 8.5 | 25.5 GB | 48 GB | Near-original |
| F16 | 16 | 48.0 GB | 64 GB | Original |
Sizes are estimates from parameter count Γ bits per weight; real GGUF builds vary slightly. Β· Data updated: 2026-06-11 Β· How we calculate these numbers β
Memory needed by context length
| Context | KV cache (est.) | Total memory (Q4) |
|---|---|---|
| 4K tokens | ~0.9 GB | ~15.5 GB |
| 8K tokens | ~1.7 GB | ~16.3 GB |
| 32K tokens | ~6.9 GB | ~21.5 GB |
| 128K tokens | ~27.5 GB | ~42.1 GB |
The KV cache grows with context length β a model that fits at 4K can run out of memory at 32K. Estimates assume an FP16 cache with grouped-query attention; actual usage varies by runtime.
Estimated speed by hardware
| Hardware | Bandwidth | ~Speed |
|---|---|---|
| NVIDIA RTX 3060 12GB | 360 GB/s | Won't fit in VRAM |
| NVIDIA RTX 4090 24GB | 1008 GB/s | ~59 tok/s |
| Apple M-series (base) | 100 GB/s | ~6 tok/s |
| Apple M-series Pro | 270 GB/s | ~16 tok/s |
| Apple M-series Max | 410 GB/s | ~24 tok/s |
| CPU only (dual-channel DDR5) | 60 GB/s | ~4 tok/s |
Token generation is memory-bandwidth bound: tok/s β bandwidth Γ 0.85 Γ· model size at Q4. Real-world numbers vary by runtime and context length.