Can I run Llama 3.1 8B on the Apple M4 Pro?
Runs greatQ4_K_M fits in unified memory (6.6 of 16.8 GB usable by the GPU): Runs great, est. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%).
24 GB unified memory · 273.0 GB/s · FP16 17.0 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Verdict by memory size
The page uses the 24 GB configuration; here is every size.
- 24 GBRuns greatest. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%)
- 48 GBRuns greatest. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%)
- 64 GBRuns greatest. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%)
Every quant of Llama 3.1 8B on the Apple M4 Pro
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 3.18 GB | Runs great | est. 38.5 tok/s30.8–46.2calibrated estimate ±20% | 4.8 / 16.8 GB | Unified memory | — |
| IQ4_XS | 4.45 GB | Runs great | est. 29.7 tok/s23.7–35.6calibrated estimate ±20% | 6.1 / 16.8 GB | Unified memory | — |
| Q4_K_Mbaseline | 4.92 GB | Runs great | est. 27.3 tok/s21.9–32.8calibrated estimate ±20% | 6.6 / 16.8 GB | Unified memory | — |
| Q8_0 | 8.54 GB | Runs well | est. 17.0 tok/s13.6–20.4calibrated estimate ±20% | 10.2 / 16.8 GB | Unified memory | — |
Where the memory goes at Q4_K_M
- Weights
- 4.9 GB
- KV cache
- 1.1 GB
- Compute buffer
- 0.6 GB
- Free
- 10.2 GB
- Not usable by the GPU
- 7.2 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 6.0 GBRuns great est. 30.0 tok/s24.0–36.0calibrated estimate ±20% | 5.7 GBRuns great est. 31.5 tok/s25.2–37.7calibrated estimate ±20% | 5.6 GBRuns great est. 32.3 tok/s25.8–38.7calibrated estimate ±20% |
| 8k | 6.6 GBRuns great est. 27.3 tok/s21.9–32.8calibrated estimate ±20% | 6.1 GBRuns great est. 29.8 tok/s23.8–35.8calibrated estimate ±20% | 5.8 GBRuns great est. 31.3 tok/s25.1–37.6calibrated estimate ±20% |
| 16k | 7.7 GBRuns great est. 23.2 tok/s18.5–27.8calibrated estimate ±20% | 6.7 GBRuns great est. 27.0 tok/s21.6–32.4calibrated estimate ±20% | 6.2 GBRuns great est. 29.6 tok/s23.7–35.5calibrated estimate ±20% |
| 32k | 10.0 GBRuns well est. 17.8 tok/s14.2–21.3calibrated estimate ±20% | 8.0 GBRuns great est. 22.7 tok/s18.2–27.2calibrated estimate ±20% | 6.9 GBRuns great est. 26.7 tok/s21.3–32.0calibrated estimate ±20% |
| 64k | 14.6 GBRuns well est. 12.1 tok/s9.7–14.5calibrated estimate ±20% | 10.6 GBRuns well est. 17.2 tok/s13.8–20.7calibrated estimate ±20% | 8.5 GBRuns great est. 22.2 tok/s17.8–26.7calibrated estimate ±20% |
| 128k | 23.8 GBdoes not fit | 15.8 GBRuns well est. 11.6 tok/s9.3–13.9calibrated estimate ±20% | 11.5 GBRuns well est. 16.7 tok/s13.3–20.0calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs great up to 16k context with an f16 KV cache, and up to 32k with q8_0.
Estimated speed
- Generation
- est. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%)
- Prompt processing
- about 765 tok/s (459–1,071, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A larger model that also runs well on the Apple M4 Pro
Runs wellMistral Small 3.2 24B: Runs well, est. 10.5 tok/s (8.4–12.5, calibrated estimate ±20%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf (4.92 GB) from bartowski/Meta-Llama-3.1-8B-Instruct-GGUF.
llama-server -hf bartowski/Meta-Llama-3.1-8B-Instruct-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.57 GB at 8k context.
Terminal:
OLLAMA_CONTEXT_LENGTH=8192 ollama serve
ollama run hf.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: Running local LLMs on a Mac (M4, M5): how much unified memory?
Other models on the Apple M4 Pro
Llama 3.1 8B on other hardware
Frequently asked questions
Can I run Llama 3.1 8B on the Apple M4 Pro?
Q4_K_M fits in unified memory (6.6 of 16.8 GB usable by the GPU): Runs great, est. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%). Q4_K_M needs 6.6 GB at 8k context; the GPU can use 16.8 GB of the unified memory.
How much context can Llama 3.1 8B use on the Apple M4 Pro?
At Q4_K_M the verdict stays Runs great up to 16k context with an f16 KV cache (2.1 GB of KV) and up to 32k with a q8_0 KV cache (2.3 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
Q4_K_M (4.92 GB) is the recommended balance: Runs great, est. 27.3 tok/s (21.9–32.8, calibrated estimate ±20%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.