Can I run Gemma 3 12B on the Apple M4?
Runs wellQ4_K_M fits in unified memory (8.4 of 11.2 GB usable by the GPU): Runs well, est. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%).
16 GB unified memory · 120.0 GB/s · FP16 8.5 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Verdict by memory size
The page uses the 16 GB configuration; here is every size.
- 16 GBRuns wellest. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%)
- 24 GBRuns wellest. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%)
- 32 GBRuns wellest. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%)
Every quant of Gemma 3 12B on the Apple M4
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 4.77 GB | Runs well | est. 13.6 tok/s10.9–16.3calibrated estimate ±20% | 5.9 / 11.2 GB | Unified memory | — |
| IQ4_XS | 6.55 GB | Runs well | est. 10.2 tok/s8.1–12.2calibrated estimate ±20% | 7.7 / 11.2 GB | Unified memory | — |
| Q4_K_Mbaseline | 7.30 GB | Runs well | est. 9.2 tok/s7.3–11.0calibrated estimate ±20% | 8.4 / 11.2 GB | Unified memory | — |
| Q8_0 | 12.51 GB | Won't run | — | Needs about 13.6 GB; 12.0 GB of memory is free (11.2 GB usable by the GPU) | — | — |
Where the memory goes at Q4_K_M
- Weights
- 7.3 GB
- KV cache
- 0.5 GB
- Compute buffer
- 0.6 GB
- Free
- 2.8 GB
- Not usable by the GPU
- 4.8 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 8.1 GBRuns well est. 9.5 tok/s7.6–11.4calibrated estimate ±20% | 8.0 GBRuns well est. 9.7 tok/s7.7–11.6calibrated estimate ±20% | 7.9 GBRuns well est. 9.8 tok/s7.8–11.7calibrated estimate ±20% |
| 8k | 8.4 GBRuns well est. 9.2 tok/s7.3–11.0calibrated estimate ±20% | 8.2 GBRuns well est. 9.5 tok/s7.6–11.4calibrated estimate ±20% | 8.0 GBRuns well est. 9.7 tok/s7.7–11.6calibrated estimate ±20% |
| 16k | 9.0 GBRuns well est. 8.6 tok/s6.9–10.3calibrated estimate ±20% | 8.5 GBRuns well est. 9.1 tok/s7.3–11.0calibrated estimate ±20% | 8.3 GBRuns well est. 9.5 tok/s7.6–11.4calibrated estimate ±20% |
| 32k | 10.2 GBRuns slowly est. 7.6 tok/s6.1–9.1calibrated estimate ±20% | 9.2 GBRuns well est. 8.5 tok/s6.8–10.2calibrated estimate ±20% | 8.7 GBRuns well est. 9.1 tok/s7.3–10.9calibrated estimate ±20% |
| 64k | 11.6 GBRuns slowly est. 3.6 tok/s2.5–4.7theoretical estimate ±30% | 10.7 GBRuns slowly est. 7.5 tok/s6.0–9.0calibrated estimate ±20% | 9.6 GBRuns well est. 8.4 tok/s6.8–10.1calibrated estimate ±20% |
| 128k | 17.6 GBdoes not fit | 11.9 GBRuns slowly est. 3.5 tok/s2.5–4.6theoretical estimate ±30% | 9.7 GBRuns slowly est. 4.3 tok/s3.0–5.6theoretical estimate ±30% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs well up to 16k context with an f16 KV cache, and up to 32k with q8_0.
Estimated speed
- Generation
- est. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%)
- Prompt processing
- about 383 tok/s (230–536, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
Gemma 3 12B on the Apple M4 Pro
Runs greatQ4_K_M: Runs great, est. 20.9 tok/s (16.7–25.1, calibrated estimate ±20%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: google_gemma-3-12b-it-Q4_K_M.gguf (7.30 GB) from bartowski/google_gemma-3-12b-it-GGUF.
llama-server -hf bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.29 GB at 8k context.
Terminal:
OLLAMA_CONTEXT_LENGTH=8192 ollama serve
ollama run hf.co/bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: Running local LLMs on a Mac (M4, M5): how much unified memory?
Other models on the Apple M4
Gemma 3 12B on other hardware
Frequently asked questions
Can I run Gemma 3 12B on the Apple M4?
Q4_K_M fits in unified memory (8.4 of 11.2 GB usable by the GPU): Runs well, est. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%). Q4_K_M needs 8.4 GB at 8k context; the GPU can use 11.2 GB of the unified memory.
How much context can Gemma 3 12B use on the Apple M4?
At Q4_K_M the verdict stays Runs well up to 16k context with an f16 KV cache (1.1 GB of KV) and up to 32k with a q8_0 KV cache (1.1 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
Q4_K_M (7.30 GB) is the recommended balance: Runs well, est. 9.2 tok/s (7.3–11.0, calibrated estimate ±20%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.