{"code":"var Component=(()=>{var m=Object.create;var r=Object.defineProperty;var u=Object.getOwnPropertyDescriptor;var p=Object.getOwnPropertyNames;var g=Object.getPrototypeOf,y=Object.prototype.hasOwnProperty;var f=(i,e)=>()=>(e||i((e={exports:{}}).exports,e),e.exports),b=(i,e)=>{for(var n in e)r(i,n,{get:e[n],enumerable:!0})},h=(i,e,n,a)=>{if(e&&typeof e==\"object\"||typeof e==\"function\")for(let o of p(e))!y.call(i,o)&&o!==n&&r(i,o,{get:()=>e[o],enumerable:!(a=u(e,o))||a.enumerable});return i};var w=(i,e,n)=>(n=i!=null?m(g(i)):{},h(e||!i||!i.__esModule?r(n,\"default\",{value:i,enumerable:!0}):n,i)),k=i=>h(r({},\"__esModule\",{value:!0}),i);var s=f((G,l)=>{l.exports=_jsx_runtime});var B={};b(B,{default:()=>c});var t=w(s());function d(i){let e={a:\"a\",blockquote:\"blockquote\",code:\"code\",em:\"em\",h2:\"h2\",h3:\"h3\",img:\"img\",li:\"li\",ol:\"ol\",p:\"p\",pre:\"pre\",span:\"span\",strong:\"strong\",table:\"table\",tbody:\"tbody\",td:\"td\",th:\"th\",thead:\"thead\",tr:\"tr\",ul:\"ul\",...i.components};return(0,t.jsxs)(t.Fragment,{children:[(0,t.jsx)(e.p,{children:\"Apple's new Mac Studio tops out at 512GB of unified memory. The press framing is that this puts frontier-scale models on a desk. That is broadly true, and it is also the answer to a question most people buying a Mac are not asking.\"}),`\n`,(0,t.jsxs)(e.p,{children:[\"The question they are asking is narrower and more useful: \",(0,t.jsx)(e.strong,{children:\"given what I want to run, how much memory do I actually need to buy?\"}),\" Unified memory is soldered to the chip package. You get one chance to answer this, at checkout, for the life of the machine.\"]}),`\n`,(0,t.jsx)(e.p,{children:\"The good news is that it is arithmetic, not a matter of opinion. This article does the arithmetic.\"}),`\n`,(0,t.jsx)(e.p,{children:(0,t.jsx)(e.img,{alt:\"Unified memory and local LLMs on Mac\",src:\"/images/blog/mac-unified-memory-local-llm-how-much-ram-m5-ultra-512gb-guide-2026/hero.webp\",width:\"1024\",height:\"541\"})}),`\n`,(0,t.jsx)(e.h2,{id:\"key-takeaways\",children:\"Key Takeaways\"}),`\n`,(0,t.jsxs)(e.ul,{children:[`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Two numbers decide everything.\"}),\" Memory \",(0,t.jsx)(e.em,{children:\"capacity\"}),\" decides whether a model loads at all. Memory \",(0,t.jsx)(e.em,{children:\"bandwidth\"}),\" decides how fast it generates text. They are independent, and a machine can be strong in one and weak in the other.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Estimate weights as parameters \\xD7 bytes-per-parameter.\"}),\" At 4-bit quantization that is roughly 0.55 bytes per parameter, so a 70B model needs about 38GB before context.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"macOS does not let the GPU use all your RAM.\"}),\" There is an undocumented ceiling controlled by \",(0,t.jsx)(e.code,{children:\"iogpu.wired_limit_mb\"}),\". On a 32GB Mac you do not get 32GB for the model.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Apple's headline AI numbers describe prompt processing, not text generation.\"}),\" Those have different bottlenecks. Read the wording carefully.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"The M5 Ultra's capacity is not new.\"}),\" The M3 Ultra already offered 512GB in 2025. What changed is bandwidth: 819GB/s to 1.2TB/s.\"]}),`\n`]}),`\n`,(0,t.jsx)(e.h2,{id:\"the-two-numbers\",children:\"The Two Numbers\"}),`\n`,(0,t.jsx)(e.p,{children:\"Almost every bad Mac-for-AI buying decision comes from collapsing these into one.\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Capacity is a wall.\"}),\" If the model plus its context does not fit in memory, it does not run. There is no graceful degradation on Apple silicon the way there is with a discrete GPU that can spill to system RAM \\u2014 on a Mac, unified memory \",(0,t.jsx)(e.em,{children:\"is\"}),\" the system RAM, and once you exceed what macOS will hand to the GPU you either fail to load or fall back to something much slower.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Bandwidth is a speed limit.\"}),\" To generate one token, a dense model reads its entire weight set from memory. Every token. That means the theoretical ceiling on generation speed is memory bandwidth divided by model size, and no amount of GPU cores changes it.\"]}),`\n`,(0,t.jsx)(e.p,{children:\"This is why a machine can hold a 200GB model and still feel unusable, and why a small model on a modest machine can feel instant.\"}),`\n`,(0,t.jsx)(e.h2,{id:\"capacity-what-fits\",children:\"Capacity: What Fits\"}),`\n`,(0,t.jsx)(e.p,{children:\"Weights are the dominant cost, and they are predictable:\"}),`\n`,(0,t.jsx)(t.Fragment,{children:(0,t.jsx)(e.pre,{className:\"shiki shiki-themes github-light github-dark\",style:{\"--shiki-light\":\"#24292e\",\"--shiki-dark\":\"#e1e4e8\",\"--shiki-light-bg\":\"#fff\",\"--shiki-dark-bg\":\"#24292e\"},tabIndex:\"0\",icon:'<svg viewBox=\"0 0 24 24\"><path d=\"M 6,1 C 4.354992,1 3,2.354992 3,4 v 16 c 0,1.645008 1.354992,3 3,3 h 12 c 1.645008,0 3,-1.354992 3,-3 V 8 7 A 1.0001,1.0001 0 0 0 20.707031,6.2929687 l -5,-5 A 1.0001,1.0001 0 0 0 15,1 h -1 z m 0,2 h 7 v 3 c 0,1.645008 1.354992,3 3,3 h 3 v 11 c 0,0.564129 -0.435871,1 -1,1 H 6 C 5.4358712,21 5,20.564129 5,20 V 4 C 5,3.4358712 5.4358712,3 6,3 Z M 15,3.4140625 18.585937,7 H 16 C 15.435871,7 15,6.5641288 15,6 Z\" fill=\"currentColor\" /></svg>',children:(0,t.jsx)(e.code,{children:(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{children:\"memory for weights \\u2248 parameters \\xD7 bytes per parameter\"})})})})}),`\n`,(0,t.jsx)(e.p,{children:\"Bytes per parameter depends on quantization:\"}),`\n`,(0,t.jsxs)(e.table,{children:[(0,t.jsx)(e.thead,{children:(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.th,{children:\"Precision\"}),(0,t.jsx)(e.th,{children:\"Bytes per parameter\"}),(0,t.jsx)(e.th,{children:\"Typical use\"})]})}),(0,t.jsxs)(e.tbody,{children:[(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"FP16 / BF16\"}),(0,t.jsx)(e.td,{children:\"2.0\"}),(0,t.jsx)(e.td,{children:\"Training, maximum fidelity\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"8-bit (Q8)\"}),(0,t.jsx)(e.td,{children:\"~1.0\"}),(0,t.jsx)(e.td,{children:\"Near-lossless inference\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"4-bit (Q4_K_M)\"}),(0,t.jsx)(e.td,{children:\"~0.55\"}),(0,t.jsx)(e.td,{children:\"The practical default\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"3-bit\"}),(0,t.jsx)(e.td,{children:\"~0.42\"}),(0,t.jsx)(e.td,{children:\"Noticeable quality loss\"})]})]})]}),`\n`,(0,t.jsx)(e.p,{children:\"Four-bit K-quants average a little above 4 bits once you include the scaling metadata, which is why 0.55 rather than 0.50 is the honest figure. Applying it:\"}),`\n`,(0,t.jsxs)(e.table,{children:[(0,t.jsx)(e.thead,{children:(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.th,{children:\"Model size\"}),(0,t.jsx)(e.th,{children:\"4-bit\"}),(0,t.jsx)(e.th,{children:\"8-bit\"}),(0,t.jsx)(e.th,{children:\"FP16\"})]})}),(0,t.jsxs)(e.tbody,{children:[(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"8B\"}),(0,t.jsx)(e.td,{children:\"4.4 GB\"}),(0,t.jsx)(e.td,{children:\"8 GB\"}),(0,t.jsx)(e.td,{children:\"16 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"14B\"}),(0,t.jsx)(e.td,{children:\"7.7 GB\"}),(0,t.jsx)(e.td,{children:\"14 GB\"}),(0,t.jsx)(e.td,{children:\"28 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"32B\"}),(0,t.jsx)(e.td,{children:\"17.6 GB\"}),(0,t.jsx)(e.td,{children:\"32 GB\"}),(0,t.jsx)(e.td,{children:\"64 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"70B\"}),(0,t.jsx)(e.td,{children:\"38.5 GB\"}),(0,t.jsx)(e.td,{children:\"70 GB\"}),(0,t.jsx)(e.td,{children:\"140 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"120B\"}),(0,t.jsx)(e.td,{children:\"66 GB\"}),(0,t.jsx)(e.td,{children:\"120 GB\"}),(0,t.jsx)(e.td,{children:\"240 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"235B\"}),(0,t.jsx)(e.td,{children:\"129 GB\"}),(0,t.jsx)(e.td,{children:\"235 GB\"}),(0,t.jsx)(e.td,{children:\"\\u2014\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"400B\"}),(0,t.jsx)(e.td,{children:\"220 GB\"}),(0,t.jsx)(e.td,{children:\"400 GB\"}),(0,t.jsx)(e.td,{children:\"\\u2014\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"671B\"}),(0,t.jsx)(e.td,{children:\"369 GB\"}),(0,t.jsx)(e.td,{children:\"671 GB\"}),(0,t.jsx)(e.td,{children:\"\\u2014\"})]})]})]}),`\n`,(0,t.jsx)(e.h3,{id:\"then-add-context\",children:\"Then add context\"}),`\n`,(0,t.jsxs)(e.p,{children:[\"The KV cache holds the attention keys and values for every token in the conversation, and it grows linearly with context length. Its size per token varies a lot by architecture \\u2014 models using grouped-query attention are far cheaper here than older designs \\u2014 but a workable range is \",(0,t.jsx)(e.strong,{children:\"0.1 MB to 0.5 MB per token\"}),\".\"]}),`\n`,(0,t.jsxs)(e.p,{children:[\"At 32,000 tokens of context that is roughly \",(0,t.jsx)(e.strong,{children:\"3GB to 16GB on top of the weights\"}),\". At 128,000 tokens it can exceed the model itself.\"]}),`\n`,(0,t.jsx)(e.p,{children:'This is the single most common reason a model that \"should fit\" does not. If you plan to feed it long documents or hold long conversations, budget for it explicitly rather than sizing to the weights and hoping.'}),`\n`,(0,t.jsx)(e.h3,{id:\"the-practical-rule\",children:\"The practical rule\"}),`\n`,(0,t.jsx)(t.Fragment,{children:(0,t.jsx)(e.pre,{className:\"shiki shiki-themes github-light github-dark\",style:{\"--shiki-light\":\"#24292e\",\"--shiki-dark\":\"#e1e4e8\",\"--shiki-light-bg\":\"#fff\",\"--shiki-dark-bg\":\"#24292e\"},tabIndex:\"0\",icon:'<svg viewBox=\"0 0 24 24\"><path d=\"M 6,1 C 4.354992,1 3,2.354992 3,4 v 16 c 0,1.645008 1.354992,3 3,3 h 12 c 1.645008,0 3,-1.354992 3,-3 V 8 7 A 1.0001,1.0001 0 0 0 20.707031,6.2929687 l -5,-5 A 1.0001,1.0001 0 0 0 15,1 h -1 z m 0,2 h 7 v 3 c 0,1.645008 1.354992,3 3,3 h 3 v 11 c 0,0.564129 -0.435871,1 -1,1 H 6 C 5.4358712,21 5,20.564129 5,20 V 4 C 5,3.4358712 5.4358712,3 6,3 Z M 15,3.4140625 18.585937,7 H 16 C 15.435871,7 15,6.5641288 15,6 Z\" fill=\"currentColor\" /></svg>',children:(0,t.jsx)(e.code,{children:(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{children:\"memory you need \\u2248 (weights) + (KV cache for your context) + 2\\u20133 GB overhead\"})})})})}),`\n`,(0,t.jsx)(e.p,{children:\"And then check that against what macOS will actually give you, which is not the number on the box.\"}),`\n`,(0,t.jsx)(e.h2,{id:\"the-ceiling-nobody-mentions\",children:\"The Ceiling Nobody Mentions\"}),`\n`,(0,t.jsx)(e.p,{children:\"macOS reserves unified memory for the system and caps how much the GPU may wire down. The knob is a sysctl:\"}),`\n`,(0,t.jsx)(t.Fragment,{children:(0,t.jsx)(e.pre,{className:\"shiki shiki-themes github-light github-dark\",style:{\"--shiki-light\":\"#24292e\",\"--shiki-dark\":\"#e1e4e8\",\"--shiki-light-bg\":\"#fff\",\"--shiki-dark-bg\":\"#24292e\"},tabIndex:\"0\",icon:'<svg viewBox=\"0 0 24 24\"><path d=\"m 4,4 a 1,1 0 0 0 -0.7070312,0.2929687 1,1 0 0 0 0,1.4140625 L 8.5859375,11 3.2929688,16.292969 a 1,1 0 0 0 0,1.414062 1,1 0 0 0 1.4140624,0 l 5.9999998,-6 a 1.0001,1.0001 0 0 0 0,-1.414062 L 4.7070312,4.2929687 A 1,1 0 0 0 4,4 Z m 8,14 a 1,1 0 0 0 -1,1 1,1 0 0 0 1,1 h 8 a 1,1 0 0 0 1,-1 1,1 0 0 0 -1,-1 z\" fill=\"currentColor\" /></svg>',children:(0,t.jsxs)(e.code,{children:[(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"$\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" sysctl\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" iogpu.wired_limit_mb\"})]}),`\n`,(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"iogpu.wired_limit_mb:\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#005CC5\",\"--shiki-dark\":\"#79B8FF\"},children:\" 0\"})]})]})})}),`\n`,(0,t.jsx)(e.p,{children:\"Zero means automatic. Apple does not document what automatic resolves to, and I am not going to invent a formula \\u2014 but community measurement consistently puts it somewhere around 65\\u201375% of total memory, with larger machines allowed a larger share.\"}),`\n`,(0,t.jsx)(e.p,{children:\"The practical consequence, on the machines you might buy:\"}),`\n`,(0,t.jsxs)(e.table,{children:[(0,t.jsx)(e.thead,{children:(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.th,{children:\"Installed memory\"}),(0,t.jsx)(e.th,{children:\"Roughly available to the GPU\"})]})}),(0,t.jsxs)(e.tbody,{children:[(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"16 GB\"}),(0,t.jsx)(e.td,{children:\"~10\\u201312 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"32 GB\"}),(0,t.jsx)(e.td,{children:\"~21\\u201324 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"64 GB\"}),(0,t.jsx)(e.td,{children:\"~42\\u201348 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"128 GB\"}),(0,t.jsx)(e.td,{children:\"~85\\u201396 GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"512 GB\"}),(0,t.jsx)(e.td,{children:\"~340\\u2013384 GB\"})]})]})]}),`\n`,(0,t.jsxs)(e.p,{children:[\"You can raise it. It takes effect immediately and does \",(0,t.jsx)(e.strong,{children:\"not\"}),\" survive a reboot:\"]}),`\n`,(0,t.jsx)(t.Fragment,{children:(0,t.jsx)(e.pre,{className:\"shiki shiki-themes github-light github-dark\",style:{\"--shiki-light\":\"#24292e\",\"--shiki-dark\":\"#e1e4e8\",\"--shiki-light-bg\":\"#fff\",\"--shiki-dark-bg\":\"#24292e\"},tabIndex:\"0\",icon:'<svg viewBox=\"0 0 24 24\"><path d=\"m 4,4 a 1,1 0 0 0 -0.7070312,0.2929687 1,1 0 0 0 0,1.4140625 L 8.5859375,11 3.2929688,16.292969 a 1,1 0 0 0 0,1.414062 1,1 0 0 0 1.4140624,0 l 5.9999998,-6 a 1.0001,1.0001 0 0 0 0,-1.414062 L 4.7070312,4.2929687 A 1,1 0 0 0 4,4 Z m 8,14 a 1,1 0 0 0 -1,1 1,1 0 0 0 1,1 h 8 a 1,1 0 0 0 1,-1 1,1 0 0 0 -1,-1 z\" fill=\"currentColor\" /></svg>',children:(0,t.jsxs)(e.code,{children:[(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6A737D\",\"--shiki-dark\":\"#6A737D\"},children:\"# Allow the GPU to wire 28GB on a 32GB Mac\"})}),`\n`,(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"sudo\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" sysctl\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" iogpu.wired_limit_mb=\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#005CC5\",\"--shiki-dark\":\"#79B8FF\"},children:\"28672\"})]}),`\n`,(0,t.jsx)(e.span,{className:\"line\"}),`\n`,(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6A737D\",\"--shiki-dark\":\"#6A737D\"},children:\"# Back to automatic\"})}),`\n`,(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"sudo\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" sysctl\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" iogpu.wired_limit_mb=\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#005CC5\",\"--shiki-dark\":\"#79B8FF\"},children:\"0\"})]})]})})}),`\n`,(0,t.jsxs)(e.p,{children:[\"Be careful with this. Everything you take, you take from the operating system and every other application. Push it too close to your total and you will hit swapping, beachballs, or an out-of-memory kill \\u2014 and on a Mac with a fast SSD, swapping under memory pressure is also a sustained write load. If you have seen the \",(0,t.jsx)(e.code,{children:\"System has run out of application memory\"}),\" dialog, this is one way to get there.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[\"A reasonable ceiling is total memory minus 6\\u20138GB for macOS and your other apps. Set it, run your model, and watch memory pressure in Activity Monitor. If you are seeing pressure spikes generally, our \",(0,t.jsx)(e.a,{href:\"/blog/macos-tahoe-ram-management-memory-optimization-complete-guide-2025\",children:\"Mac memory management guide\"}),\" covers the diagnostics.\"]}),`\n`,(0,t.jsx)(e.p,{children:(0,t.jsx)(e.img,{alt:\"Memory capacity versus bandwidth\",src:\"/images/blog/mac-unified-memory-local-llm-how-much-ram-m5-ultra-512gb-guide-2026/content-1.webp\",width:\"1024\",height:\"572\"})}),`\n`,(0,t.jsx)(e.h2,{id:\"bandwidth-how-fast-it-answers\",children:\"Bandwidth: How Fast It Answers\"}),`\n`,(0,t.jsx)(e.p,{children:\"Here is the estimate that matters, and here is exactly what it is:\"}),`\n`,(0,t.jsx)(t.Fragment,{children:(0,t.jsx)(e.pre,{className:\"shiki shiki-themes github-light github-dark\",style:{\"--shiki-light\":\"#24292e\",\"--shiki-dark\":\"#e1e4e8\",\"--shiki-light-bg\":\"#fff\",\"--shiki-dark-bg\":\"#24292e\"},tabIndex:\"0\",icon:'<svg viewBox=\"0 0 24 24\"><path d=\"M 6,1 C 4.354992,1 3,2.354992 3,4 v 16 c 0,1.645008 1.354992,3 3,3 h 12 c 1.645008,0 3,-1.354992 3,-3 V 8 7 A 1.0001,1.0001 0 0 0 20.707031,6.2929687 l -5,-5 A 1.0001,1.0001 0 0 0 15,1 h -1 z m 0,2 h 7 v 3 c 0,1.645008 1.354992,3 3,3 h 3 v 11 c 0,0.564129 -0.435871,1 -1,1 H 6 C 5.4358712,21 5,20.564129 5,20 V 4 C 5,3.4358712 5.4358712,3 6,3 Z M 15,3.4140625 18.585937,7 H 16 C 15.435871,7 15,6.5641288 15,6 Z\" fill=\"currentColor\" /></svg>',children:(0,t.jsx)(e.code,{children:(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{children:\"tokens per second \\u2248 (memory bandwidth \\xD7 efficiency) \\xF7 weight size in bytes\"})})})})}),`\n`,(0,t.jsx)(e.p,{children:\"Efficiency accounts for the gap between peak theoretical bandwidth and what an inference engine achieves \\u2014 on Apple silicon with a well-optimized runtime, roughly 0.6 to 0.8. I use 0.7 below.\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"These are computed estimates from published bandwidth figures, not benchmarks I ran.\"}),\" Treat them as the right order of magnitude and the correct relative ordering, not as measurements. Real numbers vary with the runtime, the quantization, the context length, and thermal behavior.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[\"Generation speed for a \",(0,t.jsx)(e.strong,{children:\"dense\"}),\" model at 4-bit:\"]}),`\n`,(0,t.jsxs)(e.table,{children:[(0,t.jsx)(e.thead,{children:(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.th,{}),(0,t.jsx)(e.th,{children:\"M6 (170GB/s)\"}),(0,t.jsx)(e.th,{children:\"M5 Pro (307GB/s)\"}),(0,t.jsx)(e.th,{children:\"M5 Max (614GB/s)\"}),(0,t.jsx)(e.th,{children:\"M5 Ultra (1.2TB/s)\"})]})}),(0,t.jsxs)(e.tbody,{children:[(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"8B (4.4GB)\"}),(0,t.jsx)(e.td,{children:\"~27 tok/s\"}),(0,t.jsx)(e.td,{children:\"~49 tok/s\"}),(0,t.jsx)(e.td,{children:\"~98 tok/s\"}),(0,t.jsx)(e.td,{children:\"~190 tok/s\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"32B (17.6GB)\"}),(0,t.jsx)(e.td,{children:\"~7 tok/s\"}),(0,t.jsx)(e.td,{children:\"~12 tok/s\"}),(0,t.jsx)(e.td,{children:\"~24 tok/s\"}),(0,t.jsx)(e.td,{children:\"~48 tok/s\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"70B (38.5GB)\"}),(0,t.jsx)(e.td,{children:\"does not fit\"}),(0,t.jsx)(e.td,{children:\"~6 tok/s\"}),(0,t.jsx)(e.td,{children:\"~11 tok/s\"}),(0,t.jsx)(e.td,{children:\"~22 tok/s\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"120B (66GB)\"}),(0,t.jsx)(e.td,{children:\"does not fit\"}),(0,t.jsx)(e.td,{children:\"does not fit\"}),(0,t.jsx)(e.td,{children:\"~7 tok/s\"}),(0,t.jsx)(e.td,{children:\"~13 tok/s\"})]})]})]}),`\n`,(0,t.jsx)(e.p,{children:\"For reference: comfortable reading speed is around 10 tokens per second. Below about 5, most people stop using the thing.\"}),`\n`,(0,t.jsx)(e.p,{children:\"Look at the 32B row. That is one model, and it goes from barely tolerable to genuinely quick across the range \\u2014 with no change in whether it fits. That is bandwidth, not capacity, and it is the axis people forget when they buy the cheapest machine that technically holds the model.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"why-moe-models-change-the-answer\",children:\"Why MoE models change the answer\"}),`\n`,(0,t.jsx)(e.p,{children:\"Mixture-of-experts models break the rule above, and this is the whole reason a 512GB machine is interesting.\"}),`\n`,(0,t.jsxs)(e.p,{children:[\"An MoE model stores many expert subnetworks but activates only a few per token. A 235B model with 22B active parameters must \",(0,t.jsx)(e.em,{children:\"hold\"}),\" 235B worth of weights, but only \",(0,t.jsx)(e.em,{children:\"reads\"}),\" about 22B worth to produce each token.\"]}),`\n`,(0,t.jsx)(e.p,{children:\"So:\"}),`\n`,(0,t.jsxs)(e.ul,{children:[`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Capacity\"}),\" is governed by total parameters \\u2014 129GB at 4-bit.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Speed\"}),\" is governed by active parameters \\u2014 like a 22B dense model, so roughly 40 tok/s on an M5 Ultra rather than the ~9 tok/s a 235B dense model would manage.\"]}),`\n`]}),`\n`,(0,t.jsx)(e.p,{children:\"This is the asymmetry that makes large MoE models practical on Apple silicon in a way they are not on a consumer GPU: you need the capacity, which Apple sells in a way nobody else does at this price, and you get speed proportional to a much smaller model.\"}),`\n`,(0,t.jsxs)(e.p,{children:['It is also why \"512GB runs frontier models\" is true but incomplete. It runs ',(0,t.jsx)(e.em,{children:\"sparse\"}),\" frontier models at usable speed. A hypothetical 400B dense model at 4-bit would occupy 220GB and generate around 4 tokens per second on an M5 Ultra. It fits. You would not enjoy it.\"]}),`\n`,(0,t.jsx)(e.h2,{id:\"prompt-processing-is-a-different-problem\",children:\"Prompt Processing Is a Different Problem\"}),`\n`,(0,t.jsx)(e.p,{children:\"Read Apple's claims for the new Mac mini precisely:\"}),`\n`,(0,t.jsxs)(e.blockquote,{children:[`\n`,(0,t.jsxs)(e.p,{children:[\"Up to 4.8x faster \",(0,t.jsx)(e.strong,{children:\"LLM prompt processing\"}),` versus M4\nUp to 13.5x faster `,(0,t.jsx)(e.strong,{children:\"LLM prompt processing\"}),\" versus M1\"]}),`\n`]}),`\n`,(0,t.jsxs)(e.p,{children:[\"Prompt processing \\u2014 prefill \\u2014 is the phase where the model ingests what you gave it. Every token of your input is processed in parallel, which makes it \",(0,t.jsx)(e.strong,{children:\"compute-bound\"}),\": it scales with GPU throughput, with the Neural Accelerators now built into each GPU core, and on the M6 with the first dual Neural Engine Apple has shipped.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[\"Token generation \\u2014 decode \\u2014 happens one token at a time, each requiring a full pass over the weights. It is \",(0,t.jsx)(e.strong,{children:\"bandwidth-bound\"}),\", and no amount of compute fixes it.\"]}),`\n`,(0,t.jsx)(e.p,{children:\"Apple chose the compute-bound half to advertise. That is not misleading; it is genuinely the half that improved most this generation. But it means the marketing numbers describe how quickly the machine reads your 50-page document, not how quickly it writes the summary.\"}),`\n`,(0,t.jsx)(e.p,{children:\"Which one you care about depends on your work:\"}),`\n`,(0,t.jsxs)(e.ul,{children:[`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Long inputs, short outputs\"}),\" \\u2014 summarizing documents, classifying, extracting from a codebase \\u2014 is prefill-dominated. Apple's numbers apply.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Short inputs, long outputs\"}),\" \\u2014 drafting, chat, code generation \\u2014 is decode-dominated. Bandwidth is your number.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Long inputs and long outputs\"}),\" \\u2014 agentic workflows over a large context \\u2014 needs both, and needs capacity for the KV cache on top.\"]}),`\n`]}),`\n`,(0,t.jsx)(e.h2,{id:\"what-actually-changed-with-the-m5-ultra\",children:\"What Actually Changed with the M5 Ultra\"}),`\n`,(0,t.jsx)(e.p,{children:\"Some of the launch coverage implied that holding a frontier model on a desktop is new. It is not, and getting this right changes the upgrade decision for anyone who already owns a Mac Studio.\"}),`\n`,(0,t.jsxs)(e.table,{children:[(0,t.jsx)(e.thead,{children:(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.th,{}),(0,t.jsx)(e.th,{children:\"M3 Ultra (2025)\"}),(0,t.jsx)(e.th,{children:\"M5 Ultra (2026)\"})]})}),(0,t.jsxs)(e.tbody,{children:[(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"Max unified memory\"}),(0,t.jsx)(e.td,{children:\"512GB\"}),(0,t.jsx)(e.td,{children:\"512GB\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"Memory bandwidth\"}),(0,t.jsx)(e.td,{children:\"819GB/s\"}),(0,t.jsx)(e.td,{children:\"1.2TB/s\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"GPU cores\"}),(0,t.jsx)(e.td,{children:\"up to 80\"}),(0,t.jsx)(e.td,{children:\"up to 80\"})]}),(0,t.jsxs)(e.tr,{children:[(0,t.jsx)(e.td,{children:\"CPU cores\"}),(0,t.jsx)(e.td,{children:\"32\"}),(0,t.jsx)(e.td,{children:\"up to 36\"})]})]})]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Capacity did not change. GPU core count did not change.\"}),\" Bandwidth went up about 1.47x, and the CPU gained four cores.\"]}),`\n`,(0,t.jsx)(e.p,{children:`This lines up exactly with Apple's own comparisons against the M3 Ultra \\u2014 \"up to 1.3x higher multithreaded performance\" and \"up to 1.8x faster graphics\" are modest numbers, and Apple's big claim, \"up to 4.3x the peak AI compute performance,\" comes from the Neural Accelerators now built into each GPU core rather than from more or bigger cores.`}),`\n`,(0,t.jsx)(e.p,{children:\"So for local inference specifically:\"}),`\n`,(0,t.jsxs)(e.ul,{children:[`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Generation speed\"}),\" improves roughly in line with bandwidth \\u2014 call it 1.4x or so on the same model.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Prompt processing\"}),\" improves much more, in line with that 4.3x AI compute figure.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"What you can load\"}),\" is unchanged.\"]}),`\n`]}),`\n`,(0,t.jsxs)(e.p,{children:[\"If you own an M3 Ultra Mac Studio with 512GB and your constraint is which models fit, this generation does not solve a problem you have. If your constraint is waiting for long prompts to process, it addresses that directly. Our \",(0,t.jsx)(e.a,{href:\"/blog/mac-studio-2025-m4-max-m3-ultra-performance-guide\",children:\"Mac Studio M4 Max and M3 Ultra performance guide\"}),\" covers the previous generation in detail.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[\"One more scheduling note worth knowing before you order: \",(0,t.jsx)(e.strong,{children:\"the 512GB configuration does not ship on 22 September with everything else.\"}),\" Apple has it arriving in late October.\"]}),`\n`,(0,t.jsx)(e.p,{children:(0,t.jsx)(e.img,{alt:\"Choosing a Mac configuration for local AI\",src:\"/images/blog/mac-unified-memory-local-llm-how-much-ram-m5-ultra-512gb-guide-2026/content-2.webp\",width:\"1024\",height:\"572\"})}),`\n`,(0,t.jsx)(e.h2,{id:\"which-machine-by-what-you-actually-do\",children:\"Which Machine, By What You Actually Do\"}),`\n`,(0,t.jsx)(e.h3,{id:\"16gb--not-for-this\",children:\"16GB \\u2014 not for this\"}),`\n`,(0,t.jsx)(e.p,{children:\"You have roughly 10\\u201312GB available to the GPU. That runs 8B models at 4-bit with modest context and nothing larger. It is fine for autocomplete-class models and small local assistants, and it will frustrate you at anything else. In 2026, 16GB is a machine that does other things well and runs LLMs incidentally.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"32gb-m6-mac-mini-maxed--1299--the-entry-point\",children:\"32GB (M6 Mac mini, maxed \\u2014 $1,299) \\u2014 the entry point\"}),`\n`,(0,t.jsx)(e.p,{children:\"About 21\\u201324GB to work with. This runs 8B and 14B models comfortably and a 32B model at 4-bit with short context, at around 7 tokens per second. That last figure is the honest limit: it fits, but at the edge of pleasant.\"}),`\n`,(0,t.jsx)(e.p,{children:\"Good for: local coding assistants, private document work, learning the tooling. The $400 upgrade from 16GB is not optional if local models are a reason you are buying the machine.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"64gb-m5-pro-mac-mini--1699--the-sweet-spot-for-most-people\",children:\"64GB (M5 Pro Mac mini \\u2014 $1,699+) \\u2014 the sweet spot for most people\"}),`\n`,(0,t.jsx)(e.p,{children:\"About 42\\u201348GB available, and 307GB/s. A 70B model at 4-bit fits with room for real context, generating around 6 tokens per second \\u2014 usable for batch work, slow for interactive chat. A 32B model runs at roughly 12 tokens per second, which is comfortable.\"}),`\n`,(0,t.jsx)(e.p,{children:\"This is where most serious local-model use lands, and it is the configuration I would point most people at. It also brings Thunderbolt 5, which the M6 mini does not have.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"128gb-m5-max-mac-studio--2499--the-professional-tier\",children:\"128GB (M5 Max Mac Studio \\u2014 $2,499+) \\u2014 the professional tier\"}),`\n`,(0,t.jsx)(e.p,{children:\"About 85\\u201396GB available at 614GB/s. 70B at 4-bit runs at roughly 11 tokens per second, comfortably interactive. 120B-class models fit. Mid-size MoE models become practical.\"}),`\n`,(0,t.jsx)(e.p,{children:\"This is the configuration for someone whose work depends on local inference daily \\u2014 the bandwidth roughly doubles the M5 Pro's generation speed on the same model.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"512gb-m5-ultra-mac-studio--5499-over-18000-loaded--capacity-you-cannot-buy-elsewhere\",children:\"512GB (M5 Ultra Mac Studio \\u2014 $5,499+, over $18,000 loaded) \\u2014 capacity you cannot buy elsewhere\"}),`\n`,(0,t.jsx)(e.p,{children:\"About 340\\u2013384GB available at 1.2TB/s. This holds 400B-class MoE models at 4-bit and generates at the speed of their active parameter count, which is the specific thing no other desktop does at this price.\"}),`\n`,(0,t.jsx)(e.p,{children:\"Buy this if you are capacity-bound in a way nothing else solves. If your models fit in 128GB, the Ultra buys you speed you can get more cheaply.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"if-you-are-not-buying-a-new-mac\",children:\"If you are not buying a new Mac\"}),`\n`,(0,t.jsx)(e.p,{children:\"An existing M1/M2/M3/M4 Mac with adequate memory runs local models fine \\u2014 bandwidth on the Pro and Max tiers has been good for several generations. Check your own numbers:\"}),`\n`,(0,t.jsx)(t.Fragment,{children:(0,t.jsx)(e.pre,{className:\"shiki shiki-themes github-light github-dark\",style:{\"--shiki-light\":\"#24292e\",\"--shiki-dark\":\"#e1e4e8\",\"--shiki-light-bg\":\"#fff\",\"--shiki-dark-bg\":\"#24292e\"},tabIndex:\"0\",icon:'<svg viewBox=\"0 0 24 24\"><path d=\"m 4,4 a 1,1 0 0 0 -0.7070312,0.2929687 1,1 0 0 0 0,1.4140625 L 8.5859375,11 3.2929688,16.292969 a 1,1 0 0 0 0,1.414062 1,1 0 0 0 1.4140624,0 l 5.9999998,-6 a 1.0001,1.0001 0 0 0 0,-1.414062 L 4.7070312,4.2929687 A 1,1 0 0 0 4,4 Z m 8,14 a 1,1 0 0 0 -1,1 1,1 0 0 0 1,1 h 8 a 1,1 0 0 0 1,-1 1,1 0 0 0 -1,-1 z\" fill=\"currentColor\" /></svg>',children:(0,t.jsxs)(e.code,{children:[(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6A737D\",\"--shiki-dark\":\"#6A737D\"},children:\"# Total memory in GB\"})}),`\n`,(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#005CC5\",\"--shiki-dark\":\"#79B8FF\"},children:\"echo\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:' \"$(( '}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"$(sysctl\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#005CC5\",\"--shiki-dark\":\"#79B8FF\"},children:\" -n\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:' hw.memsize) / 1073741824 )) GB\"'})]}),`\n`,(0,t.jsx)(e.span,{className:\"line\"}),`\n`,(0,t.jsx)(e.span,{className:\"line\",children:(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6A737D\",\"--shiki-dark\":\"#6A737D\"},children:\"# Chip and current GPU wired limit\"})}),`\n`,(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"sysctl\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#005CC5\",\"--shiki-dark\":\"#79B8FF\"},children:\" -n\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" machdep.cpu.brand_string\"})]}),`\n`,(0,t.jsxs)(e.span,{className:\"line\",children:[(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#6F42C1\",\"--shiki-dark\":\"#B392F0\"},children:\"sysctl\"}),(0,t.jsx)(e.span,{style:{\"--shiki-light\":\"#032F62\",\"--shiki-dark\":\"#9ECBFF\"},children:\" iogpu.wired_limit_mb\"})]})]})})}),`\n`,(0,t.jsx)(e.p,{children:\"Then apply the tables above. A 64GB M1 Max at 400GB/s is a perfectly good local-inference machine and always was.\"}),`\n`,(0,t.jsx)(e.h2,{id:\"a-worked-example\",children:\"A Worked Example\"}),`\n`,(0,t.jsx)(e.p,{children:\"Abstract tables are easy to nod along to and hard to act on. Here is the reasoning applied to one concrete case.\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"The requirement:\"}),\" a private coding assistant that reads a moderately large codebase, holds a 32,000-token context, and answers interactively. You want a 32B-class model because smaller ones are noticeably worse at code.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Step 1 \\u2014 weights.\"}),\" 32B at 4-bit: 32 \\xD7 0.55 = \",(0,t.jsx)(e.strong,{children:\"17.6GB\"}),\".\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Step 2 \\u2014 context.\"}),\" 32,000 tokens at roughly 0.2 MB per token for a modern grouped-query-attention model: \",(0,t.jsx)(e.strong,{children:\"about 6.4GB\"}),\". This is the term people forget, and here it is over a third of the weight size.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Step 3 \\u2014 overhead.\"}),\" Runtime, Metal buffers, tokenizer: \",(0,t.jsx)(e.strong,{children:\"2\\u20133GB\"}),\".\"]}),`\n`,(0,t.jsx)(e.p,{children:(0,t.jsx)(e.strong,{children:\"Total: roughly 26\\u201327GB that must be wired for the GPU.\"})}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Step 4 \\u2014 translate to installed memory.\"}),\" At a ~70% automatic ceiling, 27GB wired needs about 38GB installed. A 32GB Mac gets you 21\\u201324GB automatically \\u2014 not enough. You could raise \",(0,t.jsx)(e.code,{children:\"iogpu.wired_limit_mb\"}),\" to 27GB on a 32GB machine, which leaves 5GB for macOS and everything else. That will work while nothing else is running and fall apart the moment you open a browser.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Conclusion: 64GB installed.\"}),\" Which lands on the M5 Pro Mac mini, not the M6.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Step 5 \\u2014 check the speed.\"}),\" 307GB/s \\xD7 0.7 \\xF7 17.6GB \\u2248 \",(0,t.jsx)(e.strong,{children:\"12 tokens per second\"}),\". Comfortable for interactive use.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[\"Now change one variable. Keep everything else and ask for a 128,000-token context: the KV cache goes to roughly 25GB, total demand to about 45GB, and required installed memory to 64GB \",(0,t.jsx)(e.em,{children:\"minimum\"}),\" with the wired limit raised \\u2014 or 128GB to be comfortable. One setting moved the answer up an entire product tier.\"]}),`\n`,(0,t.jsx)(e.p,{children:\"That is the whole method. Weights, context, overhead, divide by the ceiling, then check bandwidth against the weight size.\"}),`\n`,(0,t.jsx)(e.h2,{id:\"runtimes-do-not-all-behave-the-same\",children:\"Runtimes Do Not All Behave the Same\"}),`\n`,(0,t.jsx)(e.p,{children:\"The arithmetic above describes the hardware. What you actually observe depends on the software, and the differences are large enough to change a buying decision.\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"MLX\"}),` is Apple's own array framework, built for unified memory. It does not copy between \"CPU memory\" and \"GPU memory\" because on Apple silicon there is no such distinction, and it is generally the most memory-efficient option on a Mac. If you are buying hardware specifically for local inference, MLX-based tooling is what makes the hardware look good.`]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"llama.cpp\"}),\", and the tools built on it, is the most portable option and has excellent Metal support. Its quantization formats \\u2014 the \",(0,t.jsx)(e.code,{children:\"Q4_K_M\"}),\" family the tables above assume \\u2014 are the de facto standard, and its memory behavior is predictable and well documented.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Ollama\"}),' wraps llama.cpp with model management. Convenient, and the thing to know is that it keeps models resident in memory after use for a configurable period. If you are near your ceiling, a model you finished with an hour ago may still be occupying memory. That is a frequent cause of \"it worked yesterday.\"']}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"LM Studio\"}),\" is the graphical option and is honest about memory, showing you what will and will not fit before you load it. It is a good way to develop intuition for these numbers without doing arithmetic.\"]}),`\n`,(0,t.jsx)(e.p,{children:\"Two behaviors matter regardless of which you use:\"}),`\n`,(0,t.jsxs)(e.ol,{children:[`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Whether the KV cache is preallocated.\"}),\" Some runtimes reserve the full context window at load time; others grow it as the conversation does. Preallocation means a model that would run in practice fails to load. If a model refuses to load at 128K but loads at 8K, this is why \\u2014 and lowering the configured context is the fix, not a smaller model.\"]}),`\n`,(0,t.jsxs)(e.li,{children:[(0,t.jsx)(e.strong,{children:\"Whether unused models are unloaded.\"}),\" Check your runtime's keep-alive setting before concluding your machine is too small.\"]}),`\n`]}),`\n`,(0,t.jsx)(e.h2,{id:\"buying-headroom-without-buying-memory\",children:\"Buying Headroom Without Buying Memory\"}),`\n`,(0,t.jsx)(e.p,{children:\"Before spending $400 on a memory tier, several techniques change the arithmetic.\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Quantize harder.\"}),\" Dropping from 8-bit to 4-bit halves your memory requirement for modest quality loss. This is almost always a better trade than moving to a smaller model at higher precision \\u2014 a 32B model at 4-bit generally beats a 14B model at 8-bit, at similar memory.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Quantize the KV cache.\"}),\" Most runtimes can store the KV cache at 8-bit or lower. On long contexts, where the cache rivals the weights, this is the single largest saving available and the quality cost is small.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Right-size the context.\"}),\" A 128K context window that you use 4,000 tokens of still costs full price in preallocating runtimes. Configure the context you actually use.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Prefer MoE models.\"}),\" As covered above, a mixture-of-experts model gives you the quality of a large model at the generation speed of a small one. The cost is capacity, which is exactly what a Mac has more of than the alternatives.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Use speculative decoding.\"}),\" A small draft model proposes several tokens and the large model verifies them in one pass. Because verification is parallel, this attacks the bandwidth bottleneck directly and can meaningfully raise tokens per second \\u2014 at the cost of holding a second, small model in memory. On a bandwidth-limited machine with capacity to spare, it is a good trade.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Close things.\"}),\" Obvious, and still the answer surprisingly often. Browsers with many tabs, virtual machines, and video editors are all competing for the same pool. There is no separate VRAM to fall back on.\"]}),`\n`,(0,t.jsx)(e.h2,{id:\"troubleshooting-common-issues\",children:\"Troubleshooting Common Issues\"}),`\n`,(0,t.jsx)(e.h3,{id:\"the-model-loads-but-generation-is-extremely-slow\",children:\"The model loads but generation is extremely slow\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Problem\"}),\": You have exceeded what macOS will give the GPU, and work is spilling to swap. Symptoms are a first token that takes many seconds and generation that stutters rather than streaming evenly.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Solution\"}),\": Check memory pressure in Activity Monitor while generating. If it is yellow or red, either use a smaller quantization, reduce your context window, or raise \",(0,t.jsx)(e.code,{children:\"iogpu.wired_limit_mb\"}),\" \\u2014 leaving at least 6\\u20138GB for the system. If pressure is green and it is still slow, you are simply bandwidth-limited and the fix is a smaller model.\"]}),`\n`,(0,t.jsx)(e.h3,{id:\"system-has-run-out-of-application-memory\",children:'\"System has run out of application memory\"'}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Problem\"}),\": Usually a wired limit set too aggressively, or a large context that grew during a long session.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Solution\"}),\": Reset with \",(0,t.jsx)(e.code,{children:\"sudo sysctl iogpu.wired_limit_mb=0\"}),\" and restart the inference process. Reboot if the system stays unhappy \\u2014 the setting is not persistent, so a restart clears it regardless. Related causes are covered in our \",(0,t.jsx)(e.a,{href:\"/blog/macos-tahoe-memory-leak-fix-complete-guide-2026\",children:\"memory leak troubleshooting guide\"}),\".\"]}),`\n`,(0,t.jsx)(e.h3,{id:\"it-fits-on-paper-but-fails-to-load\",children:\"It fits on paper but fails to load\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Problem\"}),\": You sized for the weights and forgot the KV cache, or the runtime is allocating the full context up front rather than growing it.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Solution\"}),\": Lower the context length in your runtime's settings and try again. If a 32K context loads and 128K does not, that confirms the diagnosis, and the arithmetic in the capacity section tells you how much you need for the context you want.\"]}),`\n`,(0,t.jsx)(e.h3,{id:\"the-mac-gets-hot-and-slows-down-over-a-long-session\",children:\"The Mac gets hot and slows down over a long session\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Problem\"}),\": Sustained inference is a sustained load on both GPU and memory. Laptops throttle. Desktops mostly do not.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Solution\"}),\": This is one of the real arguments for a Mac mini or Mac Studio over a MacBook for this workload \\u2014 sustained thermal headroom. On a laptop, see our \",(0,t.jsx)(e.a,{href:\"/blog/macbook-pro-m5-overheating-thermal-throttling-fix-guide-2026\",children:\"MacBook Pro M5 thermal throttling guide\"}),\".\"]}),`\n`,(0,t.jsx)(e.h3,{id:\"fast-prompt-processing-slow-generation\",children:\"Fast prompt processing, slow generation\"}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Problem\"}),\": Nothing is wrong. This is the expected shape of the hardware.\"]}),`\n`,(0,t.jsxs)(e.p,{children:[(0,t.jsx)(e.strong,{children:\"Solution\"}),\": See the prefill-versus-decode section above. If generation speed is what you need, the fix is a smaller model, a more aggressive quantization, an MoE model with a small active parameter count, or more bandwidth \\u2014 in that order of cost.\"]}),`\n`,(0,t.jsx)(e.h2,{id:\"faq\",children:\"FAQ\"}),`\n`,(0,t.jsx)(e.h3,{id:\"how-much-ram-do-i-need-to-run-a-70b-model\",children:\"How much RAM do I need to run a 70B model?\"}),`\n`,(0,t.jsxs)(e.p,{children:[\"At 4-bit quantization the weights are about 38.5GB. Add context and overhead and you want \",(0,t.jsx)(e.strong,{children:\"64GB installed\"}),\" as a realistic minimum, because macOS will only hand roughly 42\\u201348GB of that to the GPU. 48GB installed is too tight once you include any meaningful context.\"]}),`\n`,(0,t.jsx)(e.h3,{id:\"is-16gb-enough-for-local-ai-in-2026\",children:\"Is 16GB enough for local AI in 2026?\"}),`\n`,(0,t.jsx)(e.p,{children:\"For 8B models at 4-bit with short context, yes. For anything larger, no. If local models are a reason you are buying the machine, treat 32GB as the floor and 64GB as the target.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"does-the-neural-engine-speed-up-local-llms\",children:\"Does the Neural Engine speed up local LLMs?\"}),`\n`,(0,t.jsx)(e.p,{children:`It contributes to prompt processing, which is where Apple's \"4.8x faster LLM prompt processing\" figure comes from, and the M6's dual 16-core Neural Engine is a real step. It does not change the bandwidth limit on token generation. Most popular local-inference runtimes on the Mac lean primarily on the GPU via Metal.`}),`\n`,(0,t.jsx)(e.h3,{id:\"should-i-buy-the-512gb-mac-studio\",children:\"Should I buy the 512GB Mac Studio?\"}),`\n`,(0,t.jsx)(e.p,{children:\"Only if you are capacity-bound. If the models you actually run fit in 128GB, an M5 Max Mac Studio does the same work for a third of the price. Also note the 512GB option does not ship until late October, and the fully configured machine reaches $18,299.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"can-i-add-memory-later\",children:\"Can I add memory later?\"}),`\n`,(0,t.jsx)(e.p,{children:\"No. Unified memory is part of the chip package on every Apple silicon Mac. This is the specification most worth overbuying, because it is the only one you can never revisit.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"is-a-mac-better-than-a-pc-with-a-discrete-gpu-for-this\",children:\"Is a Mac better than a PC with a discrete GPU for this?\"}),`\n`,(0,t.jsx)(e.p,{children:\"Different trade-offs. A discrete GPU has far higher memory bandwidth but much less memory \\u2014 consumer cards top out well below what a Mac Studio offers. A Mac wins decisively on capacity per dollar and on power draw, and loses on raw throughput for models small enough to fit in VRAM. If your models are large, the Mac is often the only single-box option. If they are small, a GPU is faster.\"}),`\n`,(0,t.jsx)(e.h3,{id:\"does-quantization-hurt-quality\",children:\"Does quantization hurt quality?\"}),`\n`,(0,t.jsx)(e.p,{children:\"8-bit is close to lossless for most purposes. 4-bit K-quants are the practical default and the degradation is modest for most tasks. Below 4-bit, quality falls off noticeably. Given the memory tables above, dropping from 8-bit to 4-bit roughly halves your memory requirement \\u2014 usually a better trade than moving to a smaller model.\"}),`\n`,(0,t.jsx)(e.h2,{id:\"conclusion\",children:\"Conclusion\"}),`\n`,(0,t.jsx)(e.p,{children:\"The 512GB headline is real, and for a small number of people it is decisive. For everyone else the useful version of this article is three lines of arithmetic: weights are parameters times bytes-per-parameter, context adds 0.1\\u20130.5 MB per token on top, and macOS gives the GPU only about 65\\u201375% of what you bought.\"}),`\n`,(0,t.jsx)(e.p,{children:\"Run those numbers against the models you actually intend to use and the answer usually lands on 64GB. Then check the second number \\u2014 bandwidth \\u2014 because it decides whether the model that fits is a model you will keep using. The gap between a 32B model at 7 tokens per second and the same model at 24 tokens per second is the gap between a demo and a tool, and no amount of capacity closes it.\"}),`\n`,(0,t.jsx)(e.p,{children:\"And whichever you choose, choose carefully. It is soldered.\"}),`\n`,(0,t.jsx)(e.p,{children:(0,t.jsxs)(e.em,{children:[\"Related reading: \",(0,t.jsx)(e.a,{href:\"/blog/mac-mini-m6-mac-studio-m5-ultra-2026-specs-price-buying-guide\",children:\"Mac mini M6 and Mac Studio M5 Ultra: what actually changed\"}),\", \",(0,t.jsx)(e.a,{href:\"/blog/local-ai-apps-mac-privacy-guide-2026\",children:\"local AI apps on Mac and what they do with your data\"}),\", and \",(0,t.jsx)(e.a,{href:\"/blog/mac-ram-prices-rising-2026-dram-shortage-memory-buying-guide\",children:\"Mac RAM prices are rising in 2026\"}),\".\"]})})]})}function c(i={}){let{wrapper:e}=i.components||{};return e?(0,t.jsx)(e,{...i,children:(0,t.jsx)(d,{...i})}):d(i)}return k(B);})();\n;return Component;","toc":[{"title":"Key Takeaways","url":"#key-takeaways","depth":2},{"title":"The Two Numbers","url":"#the-two-numbers","depth":2},{"title":"Capacity: What Fits","url":"#capacity-what-fits","depth":2},{"title":"Then add context","url":"#then-add-context","depth":3},{"title":"The practical rule","url":"#the-practical-rule","depth":3},{"title":"The Ceiling Nobody Mentions","url":"#the-ceiling-nobody-mentions","depth":2},{"title":"Bandwidth: How Fast It Answers","url":"#bandwidth-how-fast-it-answers","depth":2},{"title":"Why MoE models change the answer","url":"#why-moe-models-change-the-answer","depth":3},{"title":"Prompt Processing Is a Different Problem","url":"#prompt-processing-is-a-different-problem","depth":2},{"title":"What Actually Changed with the M5 Ultra","url":"#what-actually-changed-with-the-m5-ultra","depth":2},{"title":"Which Machine, By What You Actually Do","url":"#which-machine-by-what-you-actually-do","depth":2},{"title":"16GB — not for this","url":"#16gb--not-for-this","depth":3},{"title":"32GB (M6 Mac mini, maxed — $1,299) — the entry point","url":"#32gb-m6-mac-mini-maxed--1299--the-entry-point","depth":3},{"title":"64GB (M5 Pro Mac mini — $1,699+) — the sweet spot for most people","url":"#64gb-m5-pro-mac-mini--1699--the-sweet-spot-for-most-people","depth":3},{"title":"128GB (M5 Max Mac Studio — $2,499+) — the professional tier","url":"#128gb-m5-max-mac-studio--2499--the-professional-tier","depth":3},{"title":"512GB (M5 Ultra Mac Studio — $5,499+, over $18,000 loaded) — capacity you cannot buy elsewhere","url":"#512gb-m5-ultra-mac-studio--5499-over-18000-loaded--capacity-you-cannot-buy-elsewhere","depth":3},{"title":"If you are not buying a new Mac","url":"#if-you-are-not-buying-a-new-mac","depth":3},{"title":"A Worked Example","url":"#a-worked-example","depth":2},{"title":"Runtimes Do Not All Behave the Same","url":"#runtimes-do-not-all-behave-the-same","depth":2},{"title":"Buying Headroom Without Buying Memory","url":"#buying-headroom-without-buying-memory","depth":2},{"title":"Troubleshooting Common Issues","url":"#troubleshooting-common-issues","depth":2},{"title":"The model loads but generation is extremely slow","url":"#the-model-loads-but-generation-is-extremely-slow","depth":3},{"title":"\"System has run out of application memory\"","url":"#system-has-run-out-of-application-memory","depth":3},{"title":"It fits on paper but fails to load","url":"#it-fits-on-paper-but-fails-to-load","depth":3},{"title":"The Mac gets hot and slows down over a long session","url":"#the-mac-gets-hot-and-slows-down-over-a-long-session","depth":3},{"title":"Fast prompt processing, slow generation","url":"#fast-prompt-processing-slow-generation","depth":3},{"title":"FAQ","url":"#faq","depth":2},{"title":"How much RAM do I need to run a 70B model?","url":"#how-much-ram-do-i-need-to-run-a-70b-model","depth":3},{"title":"Is 16GB enough for local AI in 2026?","url":"#is-16gb-enough-for-local-ai-in-2026","depth":3},{"title":"Does the Neural Engine speed up local LLMs?","url":"#does-the-neural-engine-speed-up-local-llms","depth":3},{"title":"Should I buy the 512GB Mac Studio?","url":"#should-i-buy-the-512gb-mac-studio","depth":3},{"title":"Can I add memory later?","url":"#can-i-add-memory-later","depth":3},{"title":"Is a Mac better than a PC with a discrete GPU for this?","url":"#is-a-mac-better-than-a-pc-with-a-discrete-gpu-for-this","depth":3},{"title":"Does quantization hurt quality?","url":"#does-quantization-hurt-quality","depth":3},{"title":"Conclusion","url":"#conclusion","depth":2}],"estimatedTime":22}