How Much Mac Memory Do You Need to Run Local LLMs?
Unified memory decides which models load. Bandwidth decides how fast they answer. Here is the arithmetic for every 2026 Mac, up to the 512GB Mac Studio.
Apple's new Mac Studio tops out at 512GB of unified memory. The press framing is that this puts frontier-scale models on a desk. That is broadly true, and it is also the answer to a question most people buying a Mac are not asking.
The question they are asking is narrower and more useful: given what I want to run, how much memory do I actually need to buy? Unified memory is soldered to the chip package. You get one chance to answer this, at checkout, for the life of the machine.
The good news is that it is arithmetic, not a matter of opinion. This article does the arithmetic.
Key Takeaways
- Two numbers decide everything. Memory capacity decides whether a model loads at all. Memory bandwidth decides how fast it generates text. They are independent, and a machine can be strong in one and weak in the other.
- Estimate weights as parameters × bytes-per-parameter. At 4-bit quantization that is roughly 0.55 bytes per parameter, so a 70B model needs about 38GB before context.
- macOS does not let the GPU use all your RAM. There is an undocumented ceiling controlled by
iogpu.wired_limit_mb. On a 32GB Mac you do not get 32GB for the model. - Apple's headline AI numbers describe prompt processing, not text generation. Those have different bottlenecks. Read the wording carefully.
- The M5 Ultra's capacity is not new. The M3 Ultra already offered 512GB in 2025. What changed is bandwidth: 819GB/s to 1.2TB/s.
The Two Numbers
Almost every bad Mac-for-AI buying decision comes from collapsing these into one.
Capacity is a wall. If the model plus its context does not fit in memory, it does not run. There is no graceful degradation on Apple silicon the way there is with a discrete GPU that can spill to system RAM — on a Mac, unified memory is the system RAM, and once you exceed what macOS will hand to the GPU you either fail to load or fall back to something much slower.
Bandwidth is a speed limit. To generate one token, a dense model reads its entire weight set from memory. Every token. That means the theoretical ceiling on generation speed is memory bandwidth divided by model size, and no amount of GPU cores changes it.
This is why a machine can hold a 200GB model and still feel unusable, and why a small model on a modest machine can feel instant.
Capacity: What Fits
Weights are the dominant cost, and they are predictable:
memory for weights ≈ parameters × bytes per parameter
Bytes per parameter depends on quantization:
| Precision | Bytes per parameter | Typical use |
|---|---|---|
| FP16 / BF16 | 2.0 | Training, maximum fidelity |
| 8-bit (Q8) | ~1.0 | Near-lossless inference |
| 4-bit (Q4_K_M) | ~0.55 | The practical default |
| 3-bit | ~0.42 | Noticeable quality loss |
Four-bit K-quants average a little above 4 bits once you include the scaling metadata, which is why 0.55 rather than 0.50 is the honest figure. Applying it:
| Model size | 4-bit | 8-bit | FP16 |
|---|---|---|---|
| 8B | 4.4 GB | 8 GB | 16 GB |
| 14B | 7.7 GB | 14 GB | 28 GB |
| 32B | 17.6 GB | 32 GB | 64 GB |
| 70B | 38.5 GB | 70 GB | 140 GB |
| 120B | 66 GB | 120 GB | 240 GB |
| 235B | 129 GB | 235 GB | — |
| 400B | 220 GB | 400 GB | — |
| 671B | 369 GB | 671 GB | — |
Then add context
The KV cache holds the attention keys and values for every token in the conversation, and it grows linearly with context length. Its size per token varies a lot by architecture — models using grouped-query attention are far cheaper here than older designs — but a workable range is 0.1 MB to 0.5 MB per token.
At 32,000 tokens of context that is roughly 3GB to 16GB on top of the weights. At 128,000 tokens it can exceed the model itself.
This is the single most common reason a model that "should fit" does not. If you plan to feed it long documents or hold long conversations, budget for it explicitly rather than sizing to the weights and hoping.
The practical rule
memory you need ≈ (weights) + (KV cache for your context) + 2–3 GB overhead
And then check that against what macOS will actually give you, which is not the number on the box.
The Ceiling Nobody Mentions
macOS reserves unified memory for the system and caps how much the GPU may wire down. The knob is a sysctl:
$ sysctl iogpu.wired_limit_mb
iogpu.wired_limit_mb: 0
Zero means automatic. Apple does not document what automatic resolves to, and I am not going to invent a formula — but community measurement consistently puts it somewhere around 65–75% of total memory, with larger machines allowed a larger share.
The practical consequence, on the machines you might buy:
| Installed memory | Roughly available to the GPU |
|---|---|
| 16 GB | ~10–12 GB |
| 32 GB | ~21–24 GB |
| 64 GB | ~42–48 GB |
| 128 GB | ~85–96 GB |
| 512 GB | ~340–384 GB |
You can raise it. It takes effect immediately and does not survive a reboot:
# Allow the GPU to wire 28GB on a 32GB Mac
sudo sysctl iogpu.wired_limit_mb=28672
# Back to automatic
sudo sysctl iogpu.wired_limit_mb=0
Be careful with this. Everything you take, you take from the operating system and every other application. Push it too close to your total and you will hit swapping, beachballs, or an out-of-memory kill — and on a Mac with a fast SSD, swapping under memory pressure is also a sustained write load. If you have seen the System has run out of application memory dialog, this is one way to get there.
A reasonable ceiling is total memory minus 6–8GB for macOS and your other apps. Set it, run your model, and watch memory pressure in Activity Monitor. If you are seeing pressure spikes generally, our Mac memory management guide covers the diagnostics.

Bandwidth: How Fast It Answers
Here is the estimate that matters, and here is exactly what it is:
tokens per second ≈ (memory bandwidth × efficiency) ÷ weight size in bytes
Efficiency accounts for the gap between peak theoretical bandwidth and what an inference engine achieves — on Apple silicon with a well-optimized runtime, roughly 0.6 to 0.8. I use 0.7 below.
These are computed estimates from published bandwidth figures, not benchmarks I ran. Treat them as the right order of magnitude and the correct relative ordering, not as measurements. Real numbers vary with the runtime, the quantization, the context length, and thermal behavior.
Generation speed for a dense model at 4-bit:
| M6 (170GB/s) | M5 Pro (307GB/s) | M5 Max (614GB/s) | M5 Ultra (1.2TB/s) | |
|---|---|---|---|---|
| 8B (4.4GB) | ~27 tok/s | ~49 tok/s | ~98 tok/s | ~190 tok/s |
| 32B (17.6GB) | ~7 tok/s | ~12 tok/s | ~24 tok/s | ~48 tok/s |
| 70B (38.5GB) | does not fit | ~6 tok/s | ~11 tok/s | ~22 tok/s |
| 120B (66GB) | does not fit | does not fit | ~7 tok/s | ~13 tok/s |
For reference: comfortable reading speed is around 10 tokens per second. Below about 5, most people stop using the thing.
Look at the 32B row. That is one model, and it goes from barely tolerable to genuinely quick across the range — with no change in whether it fits. That is bandwidth, not capacity, and it is the axis people forget when they buy the cheapest machine that technically holds the model.
Why MoE models change the answer
Mixture-of-experts models break the rule above, and this is the whole reason a 512GB machine is interesting.
An MoE model stores many expert subnetworks but activates only a few per token. A 235B model with 22B active parameters must hold 235B worth of weights, but only reads about 22B worth to produce each token.
So:
- Capacity is governed by total parameters — 129GB at 4-bit.
- Speed is governed by active parameters — like a 22B dense model, so roughly 40 tok/s on an M5 Ultra rather than the ~9 tok/s a 235B dense model would manage.
This is the asymmetry that makes large MoE models practical on Apple silicon in a way they are not on a consumer GPU: you need the capacity, which Apple sells in a way nobody else does at this price, and you get speed proportional to a much smaller model.
It is also why "512GB runs frontier models" is true but incomplete. It runs sparse frontier models at usable speed. A hypothetical 400B dense model at 4-bit would occupy 220GB and generate around 4 tokens per second on an M5 Ultra. It fits. You would not enjoy it.
Prompt Processing Is a Different Problem
Read Apple's claims for the new Mac mini precisely:
Up to 4.8x faster LLM prompt processing versus M4 Up to 13.5x faster LLM prompt processing versus M1
Prompt processing — prefill — is the phase where the model ingests what you gave it. Every token of your input is processed in parallel, which makes it compute-bound: it scales with GPU throughput, with the Neural Accelerators now built into each GPU core, and on the M6 with the first dual Neural Engine Apple has shipped.
Token generation — decode — happens one token at a time, each requiring a full pass over the weights. It is bandwidth-bound, and no amount of compute fixes it.
Apple chose the compute-bound half to advertise. That is not misleading; it is genuinely the half that improved most this generation. But it means the marketing numbers describe how quickly the machine reads your 50-page document, not how quickly it writes the summary.
Which one you care about depends on your work:
- Long inputs, short outputs — summarizing documents, classifying, extracting from a codebase — is prefill-dominated. Apple's numbers apply.
- Short inputs, long outputs — drafting, chat, code generation — is decode-dominated. Bandwidth is your number.
- Long inputs and long outputs — agentic workflows over a large context — needs both, and needs capacity for the KV cache on top.
What Actually Changed with the M5 Ultra
Some of the launch coverage implied that holding a frontier model on a desktop is new. It is not, and getting this right changes the upgrade decision for anyone who already owns a Mac Studio.
| M3 Ultra (2025) | M5 Ultra (2026) | |
|---|---|---|
| Max unified memory | 512GB | 512GB |
| Memory bandwidth | 819GB/s | 1.2TB/s |
| GPU cores | up to 80 | up to 80 |
| CPU cores | 32 | up to 36 |
Capacity did not change. GPU core count did not change. Bandwidth went up about 1.47x, and the CPU gained four cores.
This lines up exactly with Apple's own comparisons against the M3 Ultra — "up to 1.3x higher multithreaded performance" and "up to 1.8x faster graphics" are modest numbers, and Apple's big claim, "up to 4.3x the peak AI compute performance," comes from the Neural Accelerators now built into each GPU core rather than from more or bigger cores.
So for local inference specifically:
- Generation speed improves roughly in line with bandwidth — call it 1.4x or so on the same model.
- Prompt processing improves much more, in line with that 4.3x AI compute figure.
- What you can load is unchanged.
If you own an M3 Ultra Mac Studio with 512GB and your constraint is which models fit, this generation does not solve a problem you have. If your constraint is waiting for long prompts to process, it addresses that directly. Our Mac Studio M4 Max and M3 Ultra performance guide covers the previous generation in detail.
One more scheduling note worth knowing before you order: the 512GB configuration does not ship on 22 September with everything else. Apple has it arriving in late October.

Which Machine, By What You Actually Do
16GB — not for this
You have roughly 10–12GB available to the GPU. That runs 8B models at 4-bit with modest context and nothing larger. It is fine for autocomplete-class models and small local assistants, and it will frustrate you at anything else. In 2026, 16GB is a machine that does other things well and runs LLMs incidentally.
32GB (M6 Mac mini, maxed — $1,299) — the entry point
About 21–24GB to work with. This runs 8B and 14B models comfortably and a 32B model at 4-bit with short context, at around 7 tokens per second. That last figure is the honest limit: it fits, but at the edge of pleasant.
Good for: local coding assistants, private document work, learning the tooling. The $400 upgrade from 16GB is not optional if local models are a reason you are buying the machine.
64GB (M5 Pro Mac mini — $1,699+) — the sweet spot for most people
About 42–48GB available, and 307GB/s. A 70B model at 4-bit fits with room for real context, generating around 6 tokens per second — usable for batch work, slow for interactive chat. A 32B model runs at roughly 12 tokens per second, which is comfortable.
This is where most serious local-model use lands, and it is the configuration I would point most people at. It also brings Thunderbolt 5, which the M6 mini does not have.
128GB (M5 Max Mac Studio — $2,499+) — the professional tier
About 85–96GB available at 614GB/s. 70B at 4-bit runs at roughly 11 tokens per second, comfortably interactive. 120B-class models fit. Mid-size MoE models become practical.
This is the configuration for someone whose work depends on local inference daily — the bandwidth roughly doubles the M5 Pro's generation speed on the same model.
512GB (M5 Ultra Mac Studio — $5,499+, over $18,000 loaded) — capacity you cannot buy elsewhere
About 340–384GB available at 1.2TB/s. This holds 400B-class MoE models at 4-bit and generates at the speed of their active parameter count, which is the specific thing no other desktop does at this price.
Buy this if you are capacity-bound in a way nothing else solves. If your models fit in 128GB, the Ultra buys you speed you can get more cheaply.
If you are not buying a new Mac
An existing M1/M2/M3/M4 Mac with adequate memory runs local models fine — bandwidth on the Pro and Max tiers has been good for several generations. Check your own numbers:
# Total memory in GB
echo "$(( $(sysctl -n hw.memsize) / 1073741824 )) GB"
# Chip and current GPU wired limit
sysctl -n machdep.cpu.brand_string
sysctl iogpu.wired_limit_mb
Then apply the tables above. A 64GB M1 Max at 400GB/s is a perfectly good local-inference machine and always was.
A Worked Example
Abstract tables are easy to nod along to and hard to act on. Here is the reasoning applied to one concrete case.
The requirement: a private coding assistant that reads a moderately large codebase, holds a 32,000-token context, and answers interactively. You want a 32B-class model because smaller ones are noticeably worse at code.
Step 1 — weights. 32B at 4-bit: 32 × 0.55 = 17.6GB.
Step 2 — context. 32,000 tokens at roughly 0.2 MB per token for a modern grouped-query-attention model: about 6.4GB. This is the term people forget, and here it is over a third of the weight size.
Step 3 — overhead. Runtime, Metal buffers, tokenizer: 2–3GB.
Total: roughly 26–27GB that must be wired for the GPU.
Step 4 — translate to installed memory. At a ~70% automatic ceiling, 27GB wired needs about 38GB installed. A 32GB Mac gets you 21–24GB automatically — not enough. You could raise iogpu.wired_limit_mb to 27GB on a 32GB machine, which leaves 5GB for macOS and everything else. That will work while nothing else is running and fall apart the moment you open a browser.
Conclusion: 64GB installed. Which lands on the M5 Pro Mac mini, not the M6.
Step 5 — check the speed. 307GB/s × 0.7 ÷ 17.6GB ≈ 12 tokens per second. Comfortable for interactive use.
Now change one variable. Keep everything else and ask for a 128,000-token context: the KV cache goes to roughly 25GB, total demand to about 45GB, and required installed memory to 64GB minimum with the wired limit raised — or 128GB to be comfortable. One setting moved the answer up an entire product tier.
That is the whole method. Weights, context, overhead, divide by the ceiling, then check bandwidth against the weight size.
Runtimes Do Not All Behave the Same
The arithmetic above describes the hardware. What you actually observe depends on the software, and the differences are large enough to change a buying decision.
MLX is Apple's own array framework, built for unified memory. It does not copy between "CPU memory" and "GPU memory" because on Apple silicon there is no such distinction, and it is generally the most memory-efficient option on a Mac. If you are buying hardware specifically for local inference, MLX-based tooling is what makes the hardware look good.
llama.cpp, and the tools built on it, is the most portable option and has excellent Metal support. Its quantization formats — the Q4_K_M family the tables above assume — are the de facto standard, and its memory behavior is predictable and well documented.
Ollama wraps llama.cpp with model management. Convenient, and the thing to know is that it keeps models resident in memory after use for a configurable period. If you are near your ceiling, a model you finished with an hour ago may still be occupying memory. That is a frequent cause of "it worked yesterday."
LM Studio is the graphical option and is honest about memory, showing you what will and will not fit before you load it. It is a good way to develop intuition for these numbers without doing arithmetic.
Two behaviors matter regardless of which you use:
- Whether the KV cache is preallocated. Some runtimes reserve the full context window at load time; others grow it as the conversation does. Preallocation means a model that would run in practice fails to load. If a model refuses to load at 128K but loads at 8K, this is why — and lowering the configured context is the fix, not a smaller model.
- Whether unused models are unloaded. Check your runtime's keep-alive setting before concluding your machine is too small.
Buying Headroom Without Buying Memory
Before spending $400 on a memory tier, several techniques change the arithmetic.
Quantize harder. Dropping from 8-bit to 4-bit halves your memory requirement for modest quality loss. This is almost always a better trade than moving to a smaller model at higher precision — a 32B model at 4-bit generally beats a 14B model at 8-bit, at similar memory.
Quantize the KV cache. Most runtimes can store the KV cache at 8-bit or lower. On long contexts, where the cache rivals the weights, this is the single largest saving available and the quality cost is small.
Right-size the context. A 128K context window that you use 4,000 tokens of still costs full price in preallocating runtimes. Configure the context you actually use.
Prefer MoE models. As covered above, a mixture-of-experts model gives you the quality of a large model at the generation speed of a small one. The cost is capacity, which is exactly what a Mac has more of than the alternatives.
Use speculative decoding. A small draft model proposes several tokens and the large model verifies them in one pass. Because verification is parallel, this attacks the bandwidth bottleneck directly and can meaningfully raise tokens per second — at the cost of holding a second, small model in memory. On a bandwidth-limited machine with capacity to spare, it is a good trade.
Close things. Obvious, and still the answer surprisingly often. Browsers with many tabs, virtual machines, and video editors are all competing for the same pool. There is no separate VRAM to fall back on.
Troubleshooting Common Issues
The model loads but generation is extremely slow
Problem: You have exceeded what macOS will give the GPU, and work is spilling to swap. Symptoms are a first token that takes many seconds and generation that stutters rather than streaming evenly.
Solution: Check memory pressure in Activity Monitor while generating. If it is yellow or red, either use a smaller quantization, reduce your context window, or raise iogpu.wired_limit_mb — leaving at least 6–8GB for the system. If pressure is green and it is still slow, you are simply bandwidth-limited and the fix is a smaller model.
"System has run out of application memory"
Problem: Usually a wired limit set too aggressively, or a large context that grew during a long session.
Solution: Reset with sudo sysctl iogpu.wired_limit_mb=0 and restart the inference process. Reboot if the system stays unhappy — the setting is not persistent, so a restart clears it regardless. Related causes are covered in our memory leak troubleshooting guide.
It fits on paper but fails to load
Problem: You sized for the weights and forgot the KV cache, or the runtime is allocating the full context up front rather than growing it.
Solution: Lower the context length in your runtime's settings and try again. If a 32K context loads and 128K does not, that confirms the diagnosis, and the arithmetic in the capacity section tells you how much you need for the context you want.
The Mac gets hot and slows down over a long session
Problem: Sustained inference is a sustained load on both GPU and memory. Laptops throttle. Desktops mostly do not.
Solution: This is one of the real arguments for a Mac mini or Mac Studio over a MacBook for this workload — sustained thermal headroom. On a laptop, see our MacBook Pro M5 thermal throttling guide.
Fast prompt processing, slow generation
Problem: Nothing is wrong. This is the expected shape of the hardware.
Solution: See the prefill-versus-decode section above. If generation speed is what you need, the fix is a smaller model, a more aggressive quantization, an MoE model with a small active parameter count, or more bandwidth — in that order of cost.
FAQ
How much RAM do I need to run a 70B model?
At 4-bit quantization the weights are about 38.5GB. Add context and overhead and you want 64GB installed as a realistic minimum, because macOS will only hand roughly 42–48GB of that to the GPU. 48GB installed is too tight once you include any meaningful context.
Is 16GB enough for local AI in 2026?
For 8B models at 4-bit with short context, yes. For anything larger, no. If local models are a reason you are buying the machine, treat 32GB as the floor and 64GB as the target.
Does the Neural Engine speed up local LLMs?
It contributes to prompt processing, which is where Apple's "4.8x faster LLM prompt processing" figure comes from, and the M6's dual 16-core Neural Engine is a real step. It does not change the bandwidth limit on token generation. Most popular local-inference runtimes on the Mac lean primarily on the GPU via Metal.
Should I buy the 512GB Mac Studio?
Only if you are capacity-bound. If the models you actually run fit in 128GB, an M5 Max Mac Studio does the same work for a third of the price. Also note the 512GB option does not ship until late October, and the fully configured machine reaches $18,299.
Can I add memory later?
No. Unified memory is part of the chip package on every Apple silicon Mac. This is the specification most worth overbuying, because it is the only one you can never revisit.
Is a Mac better than a PC with a discrete GPU for this?
Different trade-offs. A discrete GPU has far higher memory bandwidth but much less memory — consumer cards top out well below what a Mac Studio offers. A Mac wins decisively on capacity per dollar and on power draw, and loses on raw throughput for models small enough to fit in VRAM. If your models are large, the Mac is often the only single-box option. If they are small, a GPU is faster.
Does quantization hurt quality?
8-bit is close to lossless for most purposes. 4-bit K-quants are the practical default and the degradation is modest for most tasks. Below 4-bit, quality falls off noticeably. Given the memory tables above, dropping from 8-bit to 4-bit roughly halves your memory requirement — usually a better trade than moving to a smaller model.
Conclusion
The 512GB headline is real, and for a small number of people it is decisive. For everyone else the useful version of this article is three lines of arithmetic: weights are parameters times bytes-per-parameter, context adds 0.1–0.5 MB per token on top, and macOS gives the GPU only about 65–75% of what you bought.
Run those numbers against the models you actually intend to use and the answer usually lands on 64GB. Then check the second number — bandwidth — because it decides whether the model that fits is a model you will keep using. The gap between a 32B model at 7 tokens per second and the same model at 24 tokens per second is the gap between a demo and a tool, and no amount of capacity closes it.
And whichever you choose, choose carefully. It is soldered.
Related reading: Mac mini M6 and Mac Studio M5 Ultra: what actually changed, local AI apps on Mac and what they do with your data, and Mac RAM prices are rising in 2026.
