Running local LLMs on a Mac (M4, M5): how much unified memory?

By CanRun · Updated September 30, 2026

In CanRun's calculation, a Mac's GPU can use about 70% of its unified memory, so a 16GB Mac runs Llama 3.1 8B and Gemma 3 12B at Q4_K_M fully on the GPU. To keep gpt-oss-20b on the GPU, you need at least the 24GB configuration. Among M4 and M5 Macs, gpt-oss-120b needs a 128GB M4 Max or M5 Max with the 40-core GPU, and no 64GB Mac holds it. Memory size decides what fits on the GPU; the chip's memory bandwidth decides how fast it runs there.

How much unified memory the GPU can use

A Mac has no separate graphics memory: the CPU and GPU share one pool, unified memory, and macOS limits how much of it the GPU may take. Apple's Metal documentation reports a related value, recommendedMaxWorkingSetSize: an approximation of how much memory the GPU can allocate without affecting performance. CanRun uses one figure: about 70% of unified memory is usable by the GPU.

Take a 16GB M4. Qwen3 14B at Q4_K_M, the 4-bit quant CanRun uses as its baseline, needs 10.9 GB at an 8k context with an f16 (unquantized 16-bit) KV cache. That covers the model file, the KV cache (the memory holding the conversation so far) and a compute buffer, the runtime's working memory.

Where the memory goes: Qwen3 14B (Q4_K_M) on the Apple M4

Weights, KV cache and compute buffer are the model; the OS reservation is the display driver.
Unified memory
Weights
9.0 GB
KV cache
1.3 GB
Compute buffer
0.6 GB
Free
0.3 GB
Not usable by the GPU
4.8 GB

The bar shows the 16GB configuration (its caption names only the chip); its last segment is the memory the GPU cannot use. The need sits just inside about 70% of 16GB, so the model fits with little margin — check ollama ps for 100% GPU. At the 8k baseline, Qwen3 14B on the Apple M4 Runs slowly.

Parameter count alone does not decide whether a model fits. Phi-4 has slightly fewer parameters than Qwen3 14B, 14.7B against 14.8B, but its KV cache is larger, 1.7 GB against 1.3 GB, so at Q4_K_M and 8k it needs 11.3 GB, just over the GPU share of a 16GB M4. CanRun does not model a GPU/CPU split on a Mac, so it moves the model off the GPU entirely: CPU-only, a theoretical estimate capped at Runs slowly. Ollama itself may split the layers instead; PROCESSOR in ollama ps shows what it did.

The share for every memory size of six chips:

How many of our 29 catalog models each configuration runs, by verdict (8k context, f16 KV cache)
HardwareMemoryBandwidthUsable by the GPURuns greatRuns wellRuns slowlyWon't run
Apple M416 GB120 GB/s11.2 GB27416
Apple M424 GB120 GB/s16.8 GB29810
Apple M432 GB120 GB/s22.4 GB21197
Apple M516 GB154 GB/s11.2 GB28316
Apple M524 GB154 GB/s16.8 GB210710
Apple M532 GB154 GB/s22.4 GB21296
Apple M4 Pro24 GB273 GB/s16.8 GB10559
Apple M4 Pro48 GB273 GB/s33.6 GB13745
Apple M4 Pro64 GB273 GB/s44.8 GB13754
Apple M5 Pro24 GB307 GB/s16.8 GB11459
Apple M5 Pro48 GB307 GB/s33.6 GB14735
Apple M5 Pro64 GB307 GB/s44.8 GB14744
Apple M4 Max (40-core GPU)48 GB546 GB/s33.6 GB18605
Apple M4 Max (40-core GPU)64 GB546 GB/s44.8 GB18614
Apple M4 Max (40-core GPU)128 GB546 GB/s89.6 GB20711
Apple M5 Max (40-core GPU)48 GB614 GB/s33.6 GB20405
Apple M5 Max (40-core GPU)64 GB614 GB/s44.8 GB20414
Apple M5 Max (40-core GPU)128 GB614 GB/s89.6 GB23501

Compare a model's need with the "Usable by the GPU" column, not the memory size: a 16GB Mac gives its GPU about as much as a GeForce RTX 3060 12GB can use on Windows, 11.4 GB. The verdict columns count catalog models per verdict; the Apple M4 page and the other chip pages list them.

When macOS keeps back more or less

The 70% figure is CanRun's single approximation, not a macOS rule. According to code quoted in llama.cpp discussion #2182, macOS keeps back about a third of unified memory on Macs with up to 32GB and a quarter on larger ones. Apple's memory sizes are binary, so a 16GB Mac holds about 7% more bytes than the 16 decimal GB CanRun counts. Measured that way, the default gives the GPU slightly more than CanRun's figure on 16GB to 32GB Macs and noticeably more on larger ones, so CanRun leans conservative. Other open apps still take memory from the same pool, so check ollama ps.

The same discussion describes raising the limit in macOS Terminal with sudo sysctl iogpu.wired_limit_mb= plus a value in megabytes; it resets at restart unless added to /etc/sysctl.conf, and participants warn that leaving macOS too little memory causes heavy swapping. CanRun's tables assume the default limit.

Which Mac runs which model

The table lists the six models on CanRun's Mac grid at 8k with an f16 KV cache, each at its default quant: Q4_K_M, or MXFP4, the 4-bit format the gpt-oss models ship in.

Memory a model needs at 8k context: file of the default quant + f16 KV cache + compute buffer
ModelParametersDefault quant fileKV cache at 8kTotal at 8k
Llama 3.1 8B8.03BQ4_K_M · 4.9 GB1.1 GB6.6 GB
Gemma 3 12B12.2BQ4_K_M · 7.3 GB0.5 GB8.4 GB
Qwen3 14B14.8BQ4_K_M · 9.0 GB1.3 GB10.9 GB
gpt-oss-20b20.9B (3.6B active)MXFP4 · 12.1 GB0.2 GB12.9 GB
Qwen3 30B-A3B (2507)30.5B (3.3B active)Q4_K_M · 18.6 GB0.8 GB19.9 GB
gpt-oss-120b116.8B (5.1B active)MXFP4 · 63.4 GB0.3 GB64.3 GB

On a graphics card the OS also keeps some VRAM (about 0.6 GB on a Windows display GPU); on a Mac the GPU can use about 70% of unified memory by default.

Compare each total with the "Usable by the GPU" column above. In the grid, each Mac cell names the smallest memory size ("with N GB") that reaches the highest verdict that chip gets for the model, and the row header lists every size. A Won't run cell names the largest size, meaning even that one does not hold the model.

Verdicts at 8k context with an f16 KV cache, assuming 32 GB of DDR5 system RAM and Windows on PCs (Macs: the unified memory size shown in the cell). Speeds are estimates with an error band and a confidence label.
HardwareLlama 3.1 8BGemma 3 12BQwen3 14Bgpt-oss-20bQwen3 30B-A3B (2507)gpt-oss-120b
Apple M416/24/32 GB
Runs well
est. 12.0 tok/s9.6–14.4calibrated estimate ±20%
with 16 GBDetails
Runs well
est. 9.2 tok/s7.3–11.0calibrated estimate ±20%
with 16 GBDetails
Runs slowly
est. 7.0 tok/s5.6–8.4calibrated estimate ±20%
with 16 GBDetails
Runs well
est. 15.7 tok/s12.6–18.9calibrated estimate ±20%
with 24 GBDetails
Runs well
est. 12.8 tok/s10.2–15.4calibrated estimate ±20%
with 32 GBDetails
Won't run—with 32 GBDetails
Apple M4 Pro24/48/64 GB
Runs great
est. 27.3 tok/s21.9–32.8calibrated estimate ±20%
with 24 GBDetails
Runs great
est. 20.9 tok/s16.7–25.1calibrated estimate ±20%
with 24 GBDetails
Runs well
est. 15.8 tok/s12.7–19.0calibrated estimate ±20%
with 24 GBDetails
Runs great
est. 35.8 tok/s28.6–43.0calibrated estimate ±20%
with 24 GBDetails
Runs great
est. 29.1 tok/s23.3–34.9calibrated estimate ±20%
with 48 GBDetails
Won't run—with 64 GBDetails
Apple M4 Max (40-core GPU)48/64/128 GB
Runs great
est. 54.7 tok/s43.7–65.6calibrated estimate ±20%
with 48 GB
Runs great
est. 41.8 tok/s33.4–50.2calibrated estimate ±20%
with 48 GB
Runs great
est. 31.7 tok/s25.3–38.0calibrated estimate ±20%
with 48 GB
Runs great
92.4 tok/s8k estimate 71.6 tok/s (57.3–85.9, calibrated estimate ±20%)measured (1 run, 2k context)
with 48 GB
Runs great
est. 58.2 tok/s46.6–69.9calibrated estimate ±20%
with 48 GB
Runs great
est. 53.4 tok/s42.7–64.0calibrated estimate ±20%
with 128 GB
Apple M5 Pro24/48/64 GB
Runs great
est. 30.7 tok/s24.6–36.9calibrated estimate ±20%
with 24 GB
Runs great
est. 23.5 tok/s18.8–28.2calibrated estimate ±20%
with 24 GB
Runs well
est. 17.8 tok/s14.2–21.4calibrated estimate ±20%
with 24 GB
Runs great
est. 40.3 tok/s32.2–48.3calibrated estimate ±20%
with 24 GB
Runs great
est. 32.7 tok/s26.2–39.3calibrated estimate ±20%
with 48 GB
Won't run—with 64 GB

OpenAI's gpt-oss-20b needs 12.9 GB at 8k, and its MXFP4 file alone, 12.1 GB, is larger than a 16GB Mac's GPU share, so 24GB is the smallest configuration that holds it. With 24GB, gpt-oss-20b on the Apple M4 Runs well, and on the 24GB M4 Pro and M5 Pro it Runs great. The model card says it runs within 16GB of memory; a 16GB Mac gives its GPU only part of that, so in CanRun's calculation the file alone does not fit.

Qwen3 30B-A3B (2507) is a mixture-of-experts (MoE) model: each token uses only about 3B of its parameters, but all of them must sit in memory, 19.9 GB at 8k. It Runs well on the 32GB M4 and Runs great on the 48GB M4 Pro and M5 Pro. At 24GB it exceeds the GPU share, so CanRun moves it to CPU-only, a theoretical estimate capped at Runs slowly; the grid omits that cell, but the chip pages show it. Our page for Qwen3 30B-A3B (2507) on the Apple M4 Pro covers the 48GB configuration.

The largest model here needs 64.3 GB, more than the 44.8 GB the GPU of a 64GB Mac can use, so gpt-oss-120b on the Apple M4 Pro Won't run, and neither will the 64GB M5 Pro. On the 128GB M4 Max with the 40-core GPU, the GPU can use 89.6 GB, and the model Runs great fully on the GPU, as it does on the 128GB M5 Max with the 40-core GPU, not in the grid.

Each chip's page covers every memory size: Apple M4 Pro, Apple M5 Pro, Apple M4 Max (40-core GPU) and Apple M5 Max (40-core GPU).

Why memory bandwidth, not memory size, sets the speed

Each generated token reads the model's active weights and the whole KV cache from memory, so CanRun estimates generation speed as efficiency times memory bandwidth, divided by the bytes read per token. The public llama.cpp thread on Apple silicon performance (#4167) tracks the same relationship.

Bandwidth varies widely: the M4 has 120 GB/s, the M5 154 GB/s, the M4 Pro 273 GB/s and the M5 Pro 307 GB/s, and with the 40-core GPU the M4 Max has 546 GB/s and the M5 Max 614 GB/s. The M4 Pro has more than twice the M4's bandwidth, and the 40-core M4 Max twice the M4 Pro's.

A faster chip can lift the verdict for the same model. Llama 3.1 8B on the Apple M4 Pro Runs great, while the same model Runs well on the base M4.

Qwen3 14B Runs slowly on a 16GB Apple M4 but Runs well on a 16GB Apple M5, with little margin on both (check ollama ps), and it also Runs well on the Apple M4 Pro.

In CanRun's calculation, once a model fits fully on the GPU, every memory size of that chip gets the same estimate. More memory changes which models and contexts fit on the GPU, not the speed of a model already there, and Apple lists a single bandwidth for the M4 Pro at every memory size.

The exception: the M4 Max and M5 Max each come with a 32-core or a 40-core GPU of different bandwidth, 410 GB/s against 546 GB/s for the M4 Max and 460 GB/s against 614 GB/s for the M5 Max. Apple sells the 36GB configuration only with the 32-core GPU; this guide uses the 40-core versions.

MoE and dense models on a Mac

A dense model reads all of its weights for every token; an MoE model reads only the parameters it activates for that token. On a 48GB M4 Pro, Qwen3 30B-A3B (2507) Runs great, while the dense Qwen3 32B, which needs a similar 22.5 GB against 19.9 GB, also fits fully but Runs slowly. At 24GB on the M4 Pro, gpt-oss-20b Runs great and the dense Qwen3 14B Runs well, even though gpt-oss-20b has the larger file and, as a reasoning model, faces speed thresholds 1.5 times stricter.

A newer generation at the same tier

In the grid, the M5 Pro gets the same verdicts as the M4 Pro for all six models, but not for every model in the catalog. Qwen3 32B at 48GB Runs slowly on the M4 Pro and Runs well on the M5 Pro. Check your model on the chip's page before comparing generations.

Longer contexts on a Mac

The KV cache grows with the context and comes out of the same GPU share as the weights. A Mac has no second memory pool to spill into, so when the total passes the share, CanRun's only fallback is CPU-only.

The table is Llama 3.1 8B on the Apple M4 at Q4_K_M in the 16GB configuration; the caption names only the chip. At 8k with an f16 KV cache it needs 6.6 GB and Runs well, and 16k with f16 still fits with plenty of room.

Llama 3.1 8B on the Apple M4 (Q4_K_M): memory and verdict by context length and KV cache type
ContextKV f16KV q8_0KV q4_0
4k
6.0 GBRuns well
est. 13.2 tok/s10.6–15.8calibrated estimate ±20%
5.7 GBRuns well
est. 13.8 tok/s11.1–16.6calibrated estimate ±20%
5.6 GBRuns well
est. 14.2 tok/s11.4–17.0calibrated estimate ±20%
8k
6.6 GBRuns well
est. 12.0 tok/s9.6–14.4calibrated estimate ±20%
6.1 GBRuns well
est. 13.1 tok/s10.5–15.7calibrated estimate ±20%
5.8 GBRuns well
est. 13.8 tok/s11.0–16.5calibrated estimate ±20%
16k
7.7 GBRuns well
est. 10.2 tok/s8.2–12.2calibrated estimate ±20%
6.7 GBRuns well
est. 11.9 tok/s9.5–14.2calibrated estimate ±20%
6.2 GBRuns well
est. 13.0 tok/s10.4–15.6calibrated estimate ±20%
32k
10.0 GBRuns slowly
est. 7.8 tok/s6.3–9.4calibrated estimate ±20%
8.0 GBRuns well
est. 10.0 tok/s8.0–12.0calibrated estimate ±20%
6.9 GBRuns well
est. 11.7 tok/s9.4–14.1calibrated estimate ±20%
64k
14.6 GBdoes not fit
10.6 GBRuns slowly
est. 7.6 tok/s6.1–9.1calibrated estimate ±20%
8.5 GBRuns well
est. 9.8 tok/s7.8–11.7calibrated estimate ±20%
128k
23.8 GBdoes not fit
15.8 GBdoes not fit
9.8 GBRuns slowly
est. 4.3 tok/s3.0–5.5theoretical estimate ±30%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At 32k with f16, the KV cache alone is 4.3 GB. The total still fits, but the cell drops a verdict level because the whole KV cache is read for every token; with a q8_0 (8-bit) KV cache, 32k shows no drop. At 64k with f16, the KV cache is 8.6 GB and the cell does not fit at all, not even CPU-only. A q8_0 KV cache brings the 64k need down to 10.6 GB, and both the q8_0 and q4_0 (4-bit) cells stay on the GPU, the q8_0 one with little margin — check ollama ps.

At 128k only the q4_0 cell runs, and only CPU-only: a theoretical estimate (±30%), with -ngl 0 (no layers on the GPU) in the combo page's command. CPU-only cells show weights plus KV cache without the compute buffer, so that figure looks as if it fits the GPU share; the buffer pushes it past. Don't compare CPU-only cells with GPU cells directly.

Context costs differ by model. At 32k with f16, the KV cache of gpt-oss-20b is 0.8 GB against 4.3 GB for Llama 3.1 8B, because its sliding-window layers keep only recent tokens. For the mechanism, see KV cache explained: how context length eats VRAM.

Setting the context and KV cache type on a Mac

Three ways to set the context in Ollama on a Mac:

  • Ollama app: the context length slider in Settings.
  • Mac app, until the Mac restarts: in macOS Terminal, run launchctl setenv OLLAMA_CONTEXT_LENGTH 16384, then restart the Ollama app. Run it again after each reboot.
  • One-off server: in macOS Terminal, with the app closed, run OLLAMA_CONTEXT_LENGTH=16384 ollama serve.

For the KV cache type: Ollama quantizes the KV cache only when the server starts with OLLAMA_KV_CACHE_TYPE (f16, q8_0 or q4_0) and flash attention on; the setting applies to every model the server runs (Ollama docs, checked 2026-09-29). In CanRun's tables, q8_0 and q4_0 take about 54% and 29% of the f16 size. With the Mac app, set both variables in macOS Terminal, then restart the app:

launchctl setenv OLLAMA_FLASH_ATTENTION 1
launchctl setenv OLLAMA_KV_CACHE_TYPE q8_0

Like the context setting, these reset when the Mac restarts.

llama.cpp's macOS builds use Metal, Apple's GPU interface, by default. In macOS Terminal, the command from our page for Llama 3.1 8B on the Apple M4, set to a 16k context, looks like this; for a q8_0 KV cache, add --cache-type-k q8_0 --cache-type-v q8_0 alongside -fa on.

llama-server -hf bartowski/Meta-Llama-3.1-8B-Instruct-GGUF:Q4_K_M -c 16384 -ngl all -fa on

If you leave the context alone, Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). We have not verified which tier a Mac lands in, so check CONTEXT in ollama ps while a model is loaded; PROCESSOR should read 100% GPU. For more on these settings, see Ollama slow or forgetting your prompt? Context length defaults.

Where to go next:

How sure are these numbers?

Memory figures are calculated with a formula: real GGUF file sizes, KV cache sizes from each model's configuration, and a compute-buffer formula. The approximation is the GPU share: 70% stands in for macOS's default limit, which varies with memory size and sits at or above CanRun's figure (see the first section).

Speeds on a Mac running fully on the GPU are calibrated estimates with a ±20% band unless a cell is labelled measured; the ±12% band applies only to dense models on NVIDIA and AMD cards. CPU-only cells on a Mac are theoretical (±30%), never checked against measurements.

CanRun applies the same efficiency factors (one for dense models, one for MoE) to every base, Pro and Max chip, so its estimates rise in step with bandwidth. In the public llama.cpp Apple silicon thread, the M4 Max doubles the M4 Pro's bandwidth but lifts measured generation speed by clearly less than double, so Max-chip estimates may run high; read them with the band in mind.

Not modelled: models for MLX, Apple's machine-learning framework (CanRun's estimates are for GGUF files, calibrated on llama.cpp results), a GPU/CPU layer split, other open apps' memory, laptop thermal limits, a raised GPU memory limit, and more than one model or request at a time.

On your own Mac, ollama ps is the ground truth: SIZE, PROCESSOR (100% GPU) and CONTEXT. For llama.cpp, read the load log.

Verdict names in this guide's text refer to the 8k context, f16 KV baseline. The thresholds:

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

FAQ

Is 16GB of unified memory enough for local LLMs?

For 8B to 12B models at Q4_K_M, yes, in CanRun's calculation: Llama 3.1 8B and Gemma 3 12B fit fully on the GPU of a 16GB M4. Qwen3 14B fits with little margin, Phi-4 does not fit on the GPU, and gpt-oss-20b needs 24GB to stay on the GPU. Longer contexts shrink the headroom, so check ollama ps for 100% GPU.

Can I let the GPU use more than about 70% of unified memory?

Yes, with sudo sysctl iogpu.wired_limit_mb= as described in the first section, but leaving macOS too little memory causes heavy swapping. CanRun's tables assume the default limit and sit at or below it at every memory size; other open apps share the same memory, so close them and check ollama ps.

Does more unified memory make a model run faster?

No, not once the model fits fully on the GPU. In CanRun's calculation, speed then depends on the chip's memory bandwidth, and every memory size of that chip gets the same estimate. More memory can still move a model from CPU-only onto the GPU, which lifts its verdict because CanRun caps CPU-only estimates at Runs slowly: Qwen3 30B-A3B (2507) on the M4 Pro goes from Runs slowly at 24GB to Runs great at 48GB.

How does a Mac's unified memory compare with a graphics card's VRAM?

In CanRun's calculation, a 16GB Mac's GPU gets about as much memory as a GeForce RTX 3060 12GB can use on Windows, though the card has about three times the base M4's memory bandwidth. Where Macs pull ahead is size: the 128GB M4 Max with the 40-core GPU runs gpt-oss-120b fully on the GPU, while no graphics card in CanRun's catalog has more than 32GB.

Sources

Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.