Public evidence synthesis
Local Model Fit for Mac mini and Mac Studio
A planning tool for the latest Qwen3.8, DeepSeek V4, GLM-5.3, and Gemma 4 open-weight representatives across Mac mini and Mac Studio configurations. Start with public file facts and capacity arithmetic, then separate runtime and workflow evidence before you decide whether an upgrade is justified.
Short answer
Use this page to compare the newest high-interest representatives by family. Legacy releases remain available only for compatibility links; they are intentionally excluded from the primary matrix. A model may fit the memory budget while tool calling, JSON, vision, concurrency, or sustained behavior remains unknown.
Interactive tool
Check whether your model fits your Mac.
Choose a model or a Mac configuration to see the weight budget, memory headroom, runtime evidence, and the next step into the free Advisor.
This planning tool combines official model cards, public GGUF/QAT file metadata, runtime documentation, published evaluations, and clearly labeled community reports. It does not claim that Keep or Upgrade ran every model, and Apple or model vendors do not endorse this page.
Download CSVDownload JSON ยท Same source data as this page
The downloads are citation-friendly snapshots with source URLs, evidence levels, and the data date. Link to this Hub as the canonical explanation; the download endpoints stay out of the sitemap.
Mac mini M6 ยท 16GB153GB/s ยท US$899 ยท starting price
Model evidence matrix
Latest and high-interest open-weight models
The primary matrix prioritizes current flagship releases and useful Mac planning representatives. Capacity uses the selected public artifact size when available; parameter-based numbers are planning estimates only. Expand Sources and limits on any row to inspect attributable community signals and their limits.
| Family / tier | Model and version | Quantization / artifact | Capacity on selected Mac | Runtime evidence | Confidence | Details |
|---|---|---|---|---|---|---|
| QwenSmall | Qwen3.8 27BQwen3.8 ยท 27B dense | Q4_0 ยท GGUF16.1 GB ยท Public file metadata | No fit-8.06GB remaining after reserve | Official supportmlx2 social signals ยท see sources | highChecked 2026-08-28 | Sources and limits
|
| QwenMainstream | Qwen3.8-Flash-Next 180B (6B active)Qwen3.8-Flash-Next ยท 180B total / 6B active | BF16 ยท Safetensors360.0 GB ยท Public file metadata | No fit-352GB remaining after reserve | Community reportedmlx1 social signal ยท see sources | mediumChecked 2026-08-28 | Sources and limits
|
| QwenBoundary | Qwen3.8 2.4T-A95BQwen3.8 ยท 2400B total / 95B active | Q4_K_M ยท GGUF1440.0 GB ยท Public model source; exact file not registered | No fit-1432GB remaining after reserve | Official supportmlx | mediumChecked 2026-08-28 | Sources and limits
|
| DeepSeekMainstream | DeepSeek-V4-Flash 284B-A13BDeepSeek V4 ยท 284B total / 13B active | MLX-4bit ยท MLX170.4 GB ยท Public model source; exact file not registered | No fit-162.4GB remaining after reserve | Community reportedmlx7 social signals ยท see sources | highChecked 2026-08-28 | Sources and limits
|
| DeepSeekBoundary | DeepSeek-V4-Pro 1.6T-A49BDeepSeek V4 ยท 1600B total / 49B active | MLX-4bit ยท MLX960.0 GB ยท Public model source; exact file not registered | No fit-952GB remaining after reserve | Unknownllama.cpp | highChecked 2026-08-28 | Sources and limits
|
| GLMMainstream | GLM-5.3-Flash 320B-A18BGLM-5.3-Flash ยท 320B total / 18B active | MLX-4bit ยท MLX204.0 GB ยท Public file metadata | No fit-195.99GB remaining after reserve | Community reportedmlx2 social signals ยท see sources | highChecked 2026-08-28 | Sources and limits
|
| GLMBoundary | GLM-5.3 744B-A40BGLM-5.3 ยท 744B total / 40B active | MLX-4bit ยท MLX446.4 GB ยท Public model source; exact file not registered | No fit-438.4GB remaining after reserve | Unknownllama.cpp | highChecked 2026-08-28 | Sources and limits
|
| Google GemmaSmall | Gemma 4 E2BGemma 4 ยท 2.3B effective / 5.1B total ยท PLE | Q4_0 ยท QAT3.35 GB ยท Public file metadata | Headroom4.65GB remaining after reserve | Community reportedllama.cpp1 social signal ยท see sources | highChecked 2026-08-28 | Sources and limits
|
| Google GemmaMainstream | Gemma 4 12B UnifiedGemma 4 ยท 11.95B dense | Q4_0 ยท QAT6.98 GB ยท Public file metadata | Tight1.02GB remaining after reserve | Unknownllama.cpp | highChecked 2026-08-28 | Sources and limits
|
| Google GemmaBoundary | Gemma 4 26B-A4BGemma 4 ยท 25.2B total / 3.8B active | Q4_0 ยท QAT14.4 GB ยท Public file metadata | No fit-6.44GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-28 | Sources and limits
|
Qwen3.8 27B
Qwen3.8 ยท 27B dense
- Quantization / artifact
- Q4_0 ยท GGUF16.1 GB ยท Public file metadata
- Capacity on selected Mac
- -8.06GB remaining after reserve
- Runtime evidence
- Official supportmlx
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public file metadata16.1 GB ยท exact public file metadata
- Runtime sourcemlx ยท Official support
- Community evidence (2)Public reports are signals only; they never upgrade this page to a project test.
- Reddit ยท Qwen3.8-27B on a 24GB M4 Pro Mac miniMac mini ยท M4 Pro ยท 24GB ยท llama.cpp b10488 ยท GGUF Q4_K_M / IQ4_XSQ4_K_M 17.77GB; about 11.4 tok/s decode; about 16.6GB resident with q8 KV at 32K context ยท medium confidenceA first-person report says the 27B model is usable on 24GB, but the memory ceiling is visible once context and KV cache are included. Limitation: One user, one build, and one workload; not a Keep or Upgrade test.
- Tech media / blog ยท Qwen3.8-27B 4-bit on a 32GB M1 ProMacBook Pro ยท M1 Pro ยท 32GB ยท MLX-VLM ยท MLX 4-bit8โ8.7 tok/s; weights about 16.05GB; peak memory about 18.5โ21.7GB ยท medium confidenceA detailed local deployment log reports a stable 32GB starting point and explicitly discourages 16GB for this 27B representative. Limitation: Community log with different thermals, context, and prompt mix from the product matrix.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
Qwen3.8-Flash-Next 180B (6B active)
Qwen3.8-Flash-Next ยท 180B total / 6B active
- Quantization / artifact
- BF16 ยท Safetensors360.0 GB ยท Public file metadata
- Capacity on selected Mac
- -352GB remaining after reserve
- Runtime evidence
- Community reportedmlx
- Confidence
- mediumChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public file metadata360.0 GB ยท exact public file metadata
- Runtime sourcemlx ยท Community reported
- Community evidence (1)Public reports are signals only; they never upgrade this page to a project test.
- Reddit ยท Latest-model comparison thread (non-Apple hardware)4ร DGX Spark / other non-Apple systemsUsers discuss Qwen3.8 Flash as an exploration/subagent model; no Mac tuple or reproducible throughput is provided. ยท low confidenceThis is a freshness and interest signal only, not evidence that Flash Next fits a Mac configuration. Limitation: Cross-hardware discussion; excluded from Mac capacity conclusions.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
Qwen3.8 2.4T-A95B
Qwen3.8 ยท 2400B total / 95B active
- Quantization / artifact
- Q4_K_M ยท GGUF1440.0 GB ยท Public model source; exact file not registered
- Capacity on selected Mac
- -1432GB remaining after reserve
- Runtime evidence
- Official supportmlx
- Confidence
- mediumChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public model source; exact file not registered1440.0 GB ยท parameter planning estimate
- Runtime sourcemlx ยท Official support
- No trustworthy Mac first-person report captured in this refresh.Capacity remains a public-metadata planning estimate; runtime is not upgraded by absence of evidence.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
DeepSeek-V4-Flash 284B-A13B
DeepSeek V4 ยท 284B total / 13B active
- Quantization / artifact
- MLX-4bit ยท MLX170.4 GB ยท Public model source; exact file not registered
- Capacity on selected Mac
- -162.4GB remaining after reserve
- Runtime evidence
- Community reportedmlx
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public model source; exact file not registered170.4 GB ยท parameter planning estimate
- Runtime sourcemlx ยท Community reported
- Community evidence (7)Public reports are signals only; they never upgrade this page to a project test.
- Reddit ยท DeepSeek V4 Flash on an M5 Max 128GBMac ยท Apple M5 Max ยท 128GB ยท ds4 / SSD-streaming path ยท DwarfStar IQ2XXSAbout 81GB model working set and about 31.06 tok/s in the reported run ยท medium confidenceA direct Apple Silicon report shows a practical 128GB path for the Flash representative. Limitation: Community benchmark; exact context, thermal state, and build can change the result.
- Tech media / blog ยท DeepSeek V4 Flash on an M3 Max 128GBMacBook Pro ยท M3 Max ยท 128GB ยท ds4 experimental forkAbout 21 tok/s and about 81GB resident in the reported run ยท medium confidenceAn independent technical write-up corroborates that 128GB-class Apple Silicon can run a local Flash path. Limitation: The article is tied to an experimental fork and does not validate mainline Ollama or llama.cpp.
- Reddit ยท DeepSeek V4 Flash on 64GB Macs with SSD streaming64GB Apple Silicon Macs ยท 64GB ยท ds4 SSD streaming ยท IQ2XXS / 2-bit classReports cluster around roughly 10โ15 tok/s; technically possible, but not a comfortable in-memory workflow ยท medium confidence64GB appears technically viable only with aggressive quantization and storage streaming, so it is kept as a constrained edge case. Limitation: Mixed community reports and different Mac generations; do not read this as a 64GB recommendation.
- GitHub ยท MLX DeepSeek V4 residency growth and resource-limit crashM4 Max 128GB and other Apple Silicon reports ยท 128GB ยท MLX-LMReproduction reports deterministic resource-limit failure around 11,300 generated tokens ยท high confidenceThe issue is important negative evidence: a model can fit and still fail on long generations because of runtime memory behavior. Limitation: Issue state and fixes can change; check the runtime version and referenced patch before relying on MLX.
- GitHub ยท MLX cache/residency failure on an M3 Ultra 512GBMac Studio ยท M3 Ultra 512GB ยท 512GB ยท MLX-LMReports failure around 11,456โ11,488 generated tokens in a production-style run ยท high confidenceEven very large unified-memory systems are not automatically stable for long-context MLX runs. Limitation: This is a runtime bug report, not a capacity limit; a patched release may change the result.
- GitHub ยท DeepSeek V4 Flash Q8 garbled output on Apple NEONApple Silicon CPU / NEON ยท llama.cpp ยท Q8A specific repack path produced garbled output ยท high confidenceThis narrows the risk to a quantization/repack/runtime combination instead of treating every llama.cpp path as equivalent. Limitation: Do not generalize one broken Q8 repack to all DeepSeek V4 Flash formats.
- GitHub ยท ds4 Apple Silicon throughput tableMac Studio ยท M3 Ultra 512GB ยท 512GB ยท ds4 Metal/CUDA engine ยท Q4 classRepository table reports short/long generation figures around 78.95 / 35.50 tok/s ยท medium confidenceA purpose-built engine provides a useful upper-bound signal for a large-memory Mac Studio, but it is not a mainstream runtime guarantee. Limitation: Single graph worker, no batching, and engine-specific optimizations; compare only directionally.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
DeepSeek-V4-Pro 1.6T-A49B
DeepSeek V4 ยท 1600B total / 49B active
- Quantization / artifact
- MLX-4bit ยท MLX960.0 GB ยท Public model source; exact file not registered
- Capacity on selected Mac
- -952GB remaining after reserve
- Runtime evidence
- Unknownllama.cpp
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public model source; exact file not registered960.0 GB ยท parameter planning estimate
- Runtime sourcellama.cpp ยท Unknown
- No trustworthy Mac first-person report captured in this refresh.Capacity remains a public-metadata planning estimate; runtime is not upgraded by absence of evidence.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
GLM-5.3-Flash 320B-A18B
GLM-5.3-Flash ยท 320B total / 18B active
- Quantization / artifact
- MLX-4bit ยท MLX204.0 GB ยท Public file metadata
- Capacity on selected Mac
- -195.99GB remaining after reserve
- Runtime evidence
- Community reportedmlx
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public file metadata204.0 GB ยท exact public file metadata
- Runtime sourcemlx ยท Community reported
- Community evidence (2)Public reports are signals only; they never upgrade this page to a project test.
- Hugging Face ยท Community GLM-5.3-Flash MLX quantization ladderMLX ยท 2bit-lite / 2 / 3 / 4 / 6-bit MLXPublic metadata lists about 102.43GB to 295.60GB per variant; the 4-bit root mirror is about 203.99GB ยท medium confidenceThe newer repository makes the precision trade-off explicit and includes a 2bit-lite path, but it does not prove loadability or speed on a particular Mac. Limitation: Community conversion; full-repository size includes duplicate root and variant folders, and no trustworthy Apple first-person benchmark is attached.
- Reddit ยท GLM-5.3 Flash comparison thread (non-Apple hardware)Dual DGX Spark / other non-Apple systemsOne report describes roughly 22 tok/s on dual Spark and a slower, more verbose interaction style ยท low confidenceThis indicates active community comparison but cannot be translated into a Mac recommendation. Limitation: Cross-hardware, anecdotal, and workload-specific.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
GLM-5.3 744B-A40B
GLM-5.3 ยท 744B total / 40B active
- Quantization / artifact
- MLX-4bit ยท MLX446.4 GB ยท Public model source; exact file not registered
- Capacity on selected Mac
- -438.4GB remaining after reserve
- Runtime evidence
- Unknownllama.cpp
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public model source; exact file not registered446.4 GB ยท parameter planning estimate
- Runtime sourcellama.cpp ยท Unknown
- No trustworthy Mac first-person report captured in this refresh.Capacity remains a public-metadata planning estimate; runtime is not upgraded by absence of evidence.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
Gemma 4 E2B
Gemma 4 ยท 2.3B effective / 5.1B total ยท PLE
- Quantization / artifact
- Q4_0 ยท QAT3.35 GB ยท Public file metadata
- Capacity on selected Mac
- 4.65GB remaining after reserve
- Runtime evidence
- Community reportedllama.cpp
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public file metadata3.35 GB ยท exact public file metadata
- Runtime sourcellama.cpp ยท Community reported
- Community evidence (1)Public reports are signals only; they never upgrade this page to a project test.
- Tech media / blog ยท Apple Silicon local-AI comparison including Gemma 4Mac Studio / M4 Max 128GB test systems ยท 128GB ยท llama.cppGemma 4 appears in the comparison; the article warns that bandwidth is not the sole predictor and does not publish this exact matrix tuple ยท low confidenceUseful context for Apple Silicon behavior, but not an exact Gemma 4 E2B fit or performance claim. Limitation: Different model/quantization details and media test methodology from this product matrix.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
Gemma 4 12B Unified
Gemma 4 ยท 11.95B dense
- Quantization / artifact
- Q4_0 ยท QAT6.98 GB ยท Public file metadata
- Capacity on selected Mac
- 1.02GB remaining after reserve
- Runtime evidence
- Unknownllama.cpp
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public file metadata6.98 GB ยท exact public file metadata
- Runtime sourcellama.cpp ยท Unknown
- No trustworthy Mac first-person report captured in this refresh.Capacity remains a public-metadata planning estimate; runtime is not upgraded by absence of evidence.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
Gemma 4 26B-A4B
Gemma 4 ยท 25.2B total / 3.8B active
- Quantization / artifact
- Q4_0 ยท QAT14.4 GB ยท Public file metadata
- Capacity on selected Mac
- -6.44GB remaining after reserve
- Runtime evidence
- Community reportedllama.cpp
- Confidence
- highChecked 2026-08-28
Sources and limits
- Official model sourceModel family, parameters, and release context ยท Checked 2026-08-28
- Public file metadata14.4 GB ยท exact public file metadata
- Runtime sourcellama.cpp ยท Community reported
- No trustworthy Mac first-person report captured in this refresh.Capacity remains a public-metadata planning estimate; runtime is not upgraded by absence of evidence.
- Public file metadata or a parameter estimate is not a first-party benchmark.
- Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
- Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.
- Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
- The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here.
Capability evidence
Workflow capability is a separate question
Context windows are model metadata. Concurrency, tool calling, JSON reliability, vision, and sustained behavior depend on the exact model, runtime, client, memory pressure, and task. A positive cell is still not a guarantee for your workflow.
| Model | Context | Concurrency | Tool calling | JSON | Vision | Sustained |
|---|---|---|---|---|---|---|
| Qwen3.8 27BSmall | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Qwen3.8-Flash-Next 180B (6B active)Mainstream | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | Official supportThe current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking.Source | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Qwen3.8 2.4T-A95BBoundary | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| DeepSeek-V4-Flash 284B-A13BMainstream | Official supportPublic model metadata lists 1,048,576 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official DeepSeek V4 material documents tool use in the model family. The exact local template and runtime still need a workflow check.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| DeepSeek-V4-Pro 1.6T-A49BBoundary | Official supportPublic model metadata lists 1,048,576 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official DeepSeek V4 material documents tool use in the model family. The exact local template and runtime still need a workflow check.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| GLM-5.3-Flash 320B-A18BMainstream | Official supportPublic model metadata lists 1,000,000 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official GLM-5.3 material documents function calling and agent capabilities. It is not a Mac runtime integration benchmark.Source | Public benchmarkPublic GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client.Source | Official supportThe current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking.Source | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| GLM-5.3 744B-A40BBoundary | Official supportPublic model metadata lists 1,000,000 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official GLM-5.3 material documents function calling and agent capabilities. It is not a Mac runtime integration benchmark.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Gemma 4 E2BSmall | Official supportPublic model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | Official supportThe current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking.Source | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Gemma 4 12B UnifiedMainstream | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Gemma 4 26B-A4BBoundary | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
Qwen3.8 27BSmall
- Context
- Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check.Source
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Qwen3.8-Flash-Next 180B (6B active)Mainstream
- Context
- Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check.Source
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- Official supportThe current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking.Source
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Qwen3.8 2.4T-A95BBoundary
- Context
- Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check.Source
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
DeepSeek-V4-Flash 284B-A13BMainstream
- Context
- Official supportPublic model metadata lists 1,048,576 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe official DeepSeek V4 material documents tool use in the model family. The exact local template and runtime still need a workflow check.Source
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
DeepSeek-V4-Pro 1.6T-A49BBoundary
- Context
- Official supportPublic model metadata lists 1,048,576 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe official DeepSeek V4 material documents tool use in the model family. The exact local template and runtime still need a workflow check.Source
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
GLM-5.3-Flash 320B-A18BMainstream
- Context
- Official supportPublic model metadata lists 1,000,000 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe official GLM-5.3 material documents function calling and agent capabilities. It is not a Mac runtime integration benchmark.Source
- JSON
- Public benchmarkPublic GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client.Source
- Vision
- Official supportThe current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking.Source
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
GLM-5.3 744B-A40BBoundary
- Context
- Official supportPublic model metadata lists 1,000,000 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- Official supportThe official GLM-5.3 material documents function calling and agent capabilities. It is not a Mac runtime integration benchmark.Source
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Gemma 4 E2BSmall
- Context
- Official supportPublic model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- Official supportThe current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking.Source
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Gemma 4 12B UnifiedMainstream
- Context
- Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Gemma 4 26B-A4BBoundary
- Context
- Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
- Concurrency
- System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
- Tool calling
- UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
- JSON
- UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
- Vision
- UnknownVision support is not assumed from family branding or an unrelated model variant.
- Sustained
- UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
How to read the labels
- Official support Official model/runtime documentation describes the capability path; exact client behavior still needs checking.
- Public benchmark A published evaluation exists, but it is not a Mac throughput or workflow benchmark.
- Community reported Public reports or community artifacts are useful signals, not project tests or vendor guarantees.
- System inference A bounded inference from capacity/runtime architecture, not an observed result.
- Unknown No reliable public evidence captured for this exact claim.
- Project verified Only an exact model, quantization, runtime, and machine tuple in the project smoke ledger.
Bring your real workload
Turn public evidence into a personal decision.
The free Advisor adds your current Mac, context, concurrency, budget, compatibility, and offline requirements. It can say โneeds verificationโ when the public evidence is not enough.
Data access and use are subject to our Terms of use. These downloads contain evidence metadata, not model weights. Each model remains subject to its original license.
Continue the decision
- Can my Mac run a local LLM?The lowest Mac tier with real headroom for each current model.
- Best Mac for local LLMsSeparate memory, context, runtime, concurrency, and sustained work.
- Mac mini for agentic codingApply the evidence to a repository, IDE, and tool loop.
- Ollama vs MLX on MacCompare runtime paths without turning documentation into a benchmark.
- Mac mini M6 specsCheck the compact entry configuration and its local AI limits.
Questions people ask
Clear answers before you buy
Is this a benchmark run by Keep or Upgrade?
No. This is a public-evidence synthesis and planning tool. It combines model cards, public artifact metadata, runtime documentation, published evaluations, and labeled community reports from Reddit, GitHub, Hugging Face, tech media, and public videos. Only an exact project smoke tuple is marked project verified.
How is Mac capacity calculated?
The tool uses the selected public model artifact size when an exact file or shard total is available. Otherwise it uses a clearly labeled planning estimate. The model budget reserves memory for macOS, runtime overhead, and context; it is not a promise that the workflow will be comfortable.
Do active MoE parameters mean the model needs less memory?
No. Stored weights are the capacity gate. Active parameters help explain per-token compute, but they do not remove the need to keep the full stored model package available.
Can a model with an official context window run that context on every Mac?
No. The context window is model metadata. The selected Mac must also have enough memory after weights, runtime overhead, the operating system, and other applications are accounted for.
Should I buy a Mac from this table?
Use the current-model table to narrow the planning range, then run the free Advisor with your current Mac, context length, concurrency, client, budget, compatibility, and offline requirements. Public evidence is not a first-party benchmark or a purchase guarantee.
Use your own baseline
Get a free recommendation for your actual bottleneck.
The advisor asks for the hardware, workload, budget, and compatibility details that generic buying guides cannot see.
Keep or Upgrade is an independent decision-support tool. Apple product names identify the products being compared; Apple does not sponsor or endorse this site. Specifications and vendor test claims can change, so confirm local availability, pricing, and compatibility before purchasing.