Public evidence synthesis
Local Model Fit on Mac: Qwen, DeepSeek, GLM, Gemma
A planning tool for local model fit on Apple silicon. Start with public file facts and capacity arithmetic, then separate runtime and workflow evidence before you decide whether a Mac upgrade is justified.
Short answer
Use this page to compare small, mainstream, and boundary models across Qwen, DeepSeek, GLM, and Gemma. A model may fit the memory budget while tool calling, JSON, vision, concurrency, or sustained behavior remains unknown.
This planning tool combines official model cards, public GGUF/QAT file metadata, runtime documentation, published evaluations, and clearly labeled community reports. It does not claim that Keep or Upgrade ran every model, and Apple or model vendors do not endorse this page.
Mac mini M6 · 16GB153GB/s · US$899 starting price
Model evidence matrix
Small, mainstream, and boundary models
Capacity uses the selected public artifact size when available. Parameter-based numbers are planning estimates only. MoE capacity uses stored weights, not active parameters; Gemma Effective/PLE values are kept separate from MoE arithmetic.
| Family / tier | Model and version | Quantization / artifact | Capacity on selected Mac | Runtime evidence | Confidence | Details |
|---|---|---|---|---|---|---|
| QwenSmall | Qwen3.5 0.8BQwen3.5 · 0.8B dense | Q4_K_M · GGUF0.53 GB · Public file metadata | Headroom7.47GB remaining after reserve | Project verifiedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
| QwenMainstream | Qwen3.8 27BQwen3.8 · 27B dense | Q4_0 · GGUF16.1 GB · Public file metadata | No fit-8.06GB remaining after reserve | Official supportmlx | highChecked 2026-08-27 | Sources and limits
|
| QwenBoundary | Qwen3.5 397B-A17BQwen3.5 · 397B total / 17B active | MXFP4_MOE · GGUF237.3 GB · Public file metadata | No fit-229.35GB remaining after reserve | Official supportmlx | mediumChecked 2026-08-27 | Sources and limits
|
| DeepSeekSmall | DeepSeek-R1-Distill-Qwen 1.5BDeepSeek-R1 Distill · 1.5B dense | Q4_K_M · GGUF1.12 GB · Public file metadata | Headroom6.88GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
| DeepSeekMainstream | DeepSeek-R1-Distill-Qwen 32BDeepSeek-R1 Distill · 32B dense | Q4_K_M · GGUF19.9 GB · Public file metadata | No fit-11.85GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
| DeepSeekBoundary | DeepSeek-R1 671B (685B stored package)DeepSeek-R1 · 671B total / 37B active | Q4_K_M · GGUF411.0 GB · Public model source; exact file not registered | No fit-403GB remaining after reserve | Community reportedllama.cpp | lowChecked 2026-08-27 | Sources and limits
|
| GLMSmall | GLM-4.7-Flash 30B-A3BGLM-4.7-Flash · 30B total / 3B active | Q4_K_M · GGUF18.1 GB · Public file metadata | No fit-10.13GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
| GLMMainstream | GLM-4.5-Air 106B-A12BGLM-4.5-Air · 106B total / 12B active | Q4_K_M · GGUF63.6 GB · Public model source; exact file not registered | No fit-55.6GB remaining after reserve | Community reportedllama.cpp | mediumChecked 2026-08-27 | Sources and limits
|
| GLMBoundary | GLM-5 744B-A40BGLM-5 · 744B total / 40B active | Q4_K_M · GGUF446.4 GB · Public model source; exact file not registered | No fit-438.4GB remaining after reserve | Community reportedllama.cpp | lowChecked 2026-08-27 | Sources and limits
|
| Google GemmaSmall | Gemma 4 E2BGemma 4 · 2.3B effective / 5.1B total · PLE | Q4_0 · QAT3.35 GB · Public file metadata | Headroom4.65GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
| Google GemmaMainstream | Gemma 4 12B UnifiedGemma 4 · 11.95B dense | Q4_0 · QAT6.98 GB · Public file metadata | Tight1.02GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
| Google GemmaBoundary | Gemma 4 26B-A4BGemma 4 · 25.2B total / 3.8B active | Q4_0 · QAT14.4 GB · Public file metadata | No fit-6.44GB remaining after reserve | Community reportedllama.cpp | highChecked 2026-08-27 | Sources and limits
|
Capability evidence
Workflow capability is a separate question
Context windows are model metadata. Concurrency, tool calling, JSON reliability, vision, and sustained behavior depend on the exact model, runtime, client, memory pressure, and task. A positive cell is still not a guarantee for your workflow.
| Model | Context | Concurrency | Tool calling | JSON | Vision | Sustained |
|---|---|---|---|---|---|---|
| Qwen3.5 0.8BSmall | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Qwen3.8 27BMainstream | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official Qwen-Agent project documents function calling and MCP paths. The exact model template and runtime still need a workflow check.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Qwen3.5 397B-A17BBoundary | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| DeepSeek-R1-Distill-Qwen 1.5BSmall | Official supportPublic model metadata lists 32,768 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| DeepSeek-R1-Distill-Qwen 32BMainstream | Official supportPublic model metadata lists 32,768 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| DeepSeek-R1 671B (685B stored package)Boundary | Official supportPublic model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| GLM-4.7-Flash 30B-A3BSmall | Official supportPublic model metadata lists 202,752 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official GLM material describes agent/tool-use capability or publishes agentic evaluations. It is not a Mac runtime integration benchmark.Source | Public benchmarkPublic GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client.Source | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| GLM-4.5-Air 106B-A12BMainstream | UnknownNo separately verified public context value is registered for this entry.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Official supportThe official GLM material describes agent/tool-use capability or publishes agentic evaluations. It is not a Mac runtime integration benchmark.Source | Public benchmarkPublic GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client.Source | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| GLM-5 744B-A40BBoundary | Official supportPublic model metadata lists 202,752 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | Public benchmarkThe official GLM material describes agent/tool-use capability or publishes agentic evaluations. It is not a Mac runtime integration benchmark.Source | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Gemma 4 E2BSmall | Official supportPublic model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | Official supportGoogle's Gemma 4 model card describes vision-language inputs for the family; the selected quantized runtime path still needs checking.Source | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Gemma 4 12B UnifiedMainstream | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
| Gemma 4 26B-A4BBoundary | Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source | System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source | UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot. | UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here. | UnknownVision support is not assumed from family branding or an unrelated model variant. | UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration. |
How to read the labels
- Official support Official model/runtime documentation describes the capability path; exact client behavior still needs checking.
- Public benchmark A published evaluation exists, but it is not a Mac throughput or workflow benchmark.
- Community reported Public reports or community artifacts are useful signals, not project tests or vendor guarantees.
- System inference A bounded inference from capacity/runtime architecture, not an observed result.
- Unknown No reliable public evidence captured for this exact claim.
- Project verified Only an exact model, quantization, runtime, and machine tuple in the project smoke ledger.
Bring your real workload
Turn public evidence into a personal decision.
The free Advisor adds your current Mac, context, concurrency, budget, compatibility, and offline requirements. It can say “needs verification” when the public evidence is not enough.
Continue the decision
- Best Mac for local LLMsSeparate memory, context, runtime, concurrency, and sustained work.
- Mac mini for agentic codingApply the evidence to a repository, IDE, and tool loop.
- Ollama vs MLX on MacCompare runtime paths without turning documentation into a benchmark.
- Mac mini M6 specsCheck the compact entry configuration and its local AI limits.
Questions people ask
Clear answers before you buy
Is this a benchmark run by Keep or Upgrade?
No. This is a public-evidence synthesis and planning tool. It combines model cards, public artifact metadata, runtime documentation, published evaluations, and labeled community reports. Only an exact project smoke tuple is marked project verified.
How is Mac capacity calculated?
The tool uses the selected public model artifact size when an exact file or shard total is available. Otherwise it uses a clearly labeled planning estimate. The model budget reserves memory for macOS, runtime overhead, and context; it is not a promise that the workflow will be comfortable.
Do active MoE parameters mean the model needs less memory?
No. Stored weights are the capacity gate. Active parameters help explain per-token compute, but they do not remove the need to keep the full stored model package available.
Can a model with an official context window run that context on every Mac?
No. The context window is model metadata. The selected Mac must also have enough memory after weights, runtime overhead, the operating system, and other applications are accounted for.
Should I buy a Mac from this table?
Use the table to narrow the planning range, then run the free Advisor with your current Mac, context length, concurrency, client, budget, compatibility, and offline requirements. Public evidence is not a first-party benchmark or a purchase guarantee.
Use your own baseline
Get a free recommendation for your actual bottleneck.
The advisor asks for the hardware, workload, budget, and compatibility details that generic buying guides cannot see.
Keep or Upgrade is an independent decision-support tool. Apple product names identify the products being compared; Apple does not sponsor or endorse this site. Specifications and vendor test claims can change, so confirm local availability, pricing, and compatibility before purchasing.