Public evidence synthesis

Local Model Fit on Mac: Qwen, DeepSeek, GLM, Gemma

A planning tool for local model fit on Apple silicon. Start with public file facts and capacity arithmetic, then separate runtime and workflow evidence before you decide whether a Mac upgrade is justified.

Last reviewed August 27, 2026 · Independent guidance, not Apple advice

Short answer

Use this page to compare small, mainstream, and boundary models across Qwen, DeepSeek, GLM, and Gemma. A model may fit the memory budget while tool calling, JSON, vision, concurrency, or sustained behavior remains unknown.

Public evidence synthesis, not a first-party benchmark.

This planning tool combines official model cards, public GGUF/QAT file metadata, runtime documentation, published evaluations, and clearly labeled community reports. It does not claim that Keep or Upgrade ran every model, and Apple or model vendors do not endorse this page.

Mac mini M6 · 16GB153GB/s · US$899 starting price

Model evidence matrix

Small, mainstream, and boundary models

Capacity uses the selected public artifact size when available. Parameter-based numbers are planning estimates only. MoE capacity uses stored weights, not active parameters; Gemma Effective/PLE values are kept separate from MoE arithmetic.

Family / tierModel and versionQuantization / artifactCapacity on selected MacRuntime evidenceConfidenceDetails
QwenSmallQwen3.5 0.8BQwen3.5 · 0.8B denseQ4_K_M · GGUF0.53 GB · Public file metadataHeadroom7.47GB remaining after reserveProject verifiedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata0.53 GB · exact public file metadata
  • Runtime sourcellama.cpp · Project verified
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
QwenMainstreamQwen3.8 27BQwen3.8 · 27B denseQ4_0 · GGUF16.1 GB · Public file metadataNo fit-8.06GB remaining after reserveOfficial supportmlxhighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata16.1 GB · exact public file metadata
  • Runtime sourcemlx · Official support
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
QwenBoundaryQwen3.5 397B-A17BQwen3.5 · 397B total / 17B activeMXFP4_MOE · GGUF237.3 GB · Public file metadataNo fit-229.35GB remaining after reserveOfficial supportmlxmediumChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata237.3 GB · exact public file metadata
  • Runtime sourcemlx · Official support
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
DeepSeekSmallDeepSeek-R1-Distill-Qwen 1.5BDeepSeek-R1 Distill · 1.5B denseQ4_K_M · GGUF1.12 GB · Public file metadataHeadroom6.88GB remaining after reserveCommunity reportedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata1.12 GB · exact public file metadata
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
DeepSeekMainstreamDeepSeek-R1-Distill-Qwen 32BDeepSeek-R1 Distill · 32B denseQ4_K_M · GGUF19.9 GB · Public file metadataNo fit-11.85GB remaining after reserveCommunity reportedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata19.9 GB · exact public file metadata
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
DeepSeekBoundaryDeepSeek-R1 671B (685B stored package)DeepSeek-R1 · 671B total / 37B activeQ4_K_M · GGUF411.0 GB · Public model source; exact file not registeredNo fit-403GB remaining after reserveCommunity reportedllama.cpplowChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public model source; exact file not registered411.0 GB · parameter planning estimate
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
GLMSmallGLM-4.7-Flash 30B-A3BGLM-4.7-Flash · 30B total / 3B activeQ4_K_M · GGUF18.1 GB · Public file metadataNo fit-10.13GB remaining after reserveCommunity reportedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata18.1 GB · exact public file metadata
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
GLMMainstreamGLM-4.5-Air 106B-A12BGLM-4.5-Air · 106B total / 12B activeQ4_K_M · GGUF63.6 GB · Public model source; exact file not registeredNo fit-55.6GB remaining after reserveCommunity reportedllama.cppmediumChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public model source; exact file not registered63.6 GB · parameter planning estimate
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
GLMBoundaryGLM-5 744B-A40BGLM-5 · 744B total / 40B activeQ4_K_M · GGUF446.4 GB · Public model source; exact file not registeredNo fit-438.4GB remaining after reserveCommunity reportedllama.cpplowChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public model source; exact file not registered446.4 GB · parameter planning estimate
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
Google GemmaSmallGemma 4 E2BGemma 4 · 2.3B effective / 5.1B total · PLEQ4_0 · QAT3.35 GB · Public file metadataHeadroom4.65GB remaining after reserveCommunity reportedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata3.35 GB · exact public file metadata
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
Google GemmaMainstreamGemma 4 12B UnifiedGemma 4 · 11.95B denseQ4_0 · QAT6.98 GB · Public file metadataTight1.02GB remaining after reserveCommunity reportedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata6.98 GB · exact public file metadata
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor
Google GemmaBoundaryGemma 4 26B-A4BGemma 4 · 25.2B total / 3.8B activeQ4_0 · QAT14.4 GB · Public file metadataNo fit-6.44GB remaining after reserveCommunity reportedllama.cpphighChecked 2026-08-27
Sources and limits
  • Official model sourceModel family, parameters, and release context · checked 2026-08-27
  • Public file metadata14.4 GB · exact public file metadata
  • Runtime sourcellama.cpp · Community reported
  • Public file metadata or a parameter estimate is not a first-party benchmark.
  • Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.
  • Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.
Use in free advisor

Capability evidence

Workflow capability is a separate question

Context windows are model metadata. Concurrency, tool calling, JSON reliability, vision, and sustained behavior depend on the exact model, runtime, client, memory pressure, and task. A positive cell is still not a guarantee for your workflow.

ModelContextConcurrencyTool callingJSONVisionSustained
Qwen3.5 0.8BSmall
Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Qwen3.8 27BMainstream
Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
Official supportThe official Qwen-Agent project documents function calling and MCP paths. The exact model template and runtime still need a workflow check.Source
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Qwen3.5 397B-A17BBoundary
Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
DeepSeek-R1-Distill-Qwen 1.5BSmall
Official supportPublic model metadata lists 32,768 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
DeepSeek-R1-Distill-Qwen 32BMainstream
Official supportPublic model metadata lists 32,768 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
DeepSeek-R1 671B (685B stored package)Boundary
Official supportPublic model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
GLM-4.7-Flash 30B-A3BSmall
Official supportPublic model metadata lists 202,752 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
Official supportThe official GLM material describes agent/tool-use capability or publishes agentic evaluations. It is not a Mac runtime integration benchmark.Source
Public benchmarkPublic GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client.Source
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
GLM-4.5-Air 106B-A12BMainstream
UnknownNo separately verified public context value is registered for this entry.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
Official supportThe official GLM material describes agent/tool-use capability or publishes agentic evaluations. It is not a Mac runtime integration benchmark.Source
Public benchmarkPublic GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client.Source
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
GLM-5 744B-A40BBoundary
Official supportPublic model metadata lists 202,752 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
Public benchmarkThe official GLM material describes agent/tool-use capability or publishes agentic evaluations. It is not a Mac runtime integration benchmark.Source
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Gemma 4 E2BSmall
Official supportPublic model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
Official supportGoogle's Gemma 4 model card describes vision-language inputs for the family; the selected quantized runtime path still needs checking.Source
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Gemma 4 12B UnifiedMainstream
Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.
Gemma 4 26B-A4BBoundary
Official supportPublic model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights.Source
System inferenceConcurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached.Source
UnknownNo model-specific public tool-calling result is strong enough for a positive claim in this snapshot.
UnknownStructured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here.
UnknownVision support is not assumed from family branding or an unrelated model variant.
UnknownNo public, comparable sustained Apple Silicon run is registered for this model and Mac configuration.

How to read the labels

  • Official support Official model/runtime documentation describes the capability path; exact client behavior still needs checking.
  • Public benchmark A published evaluation exists, but it is not a Mac throughput or workflow benchmark.
  • Community reported Public reports or community artifacts are useful signals, not project tests or vendor guarantees.
  • System inference A bounded inference from capacity/runtime architecture, not an observed result.
  • Unknown No reliable public evidence captured for this exact claim.
  • Project verified Only an exact model, quantization, runtime, and machine tuple in the project smoke ledger.

Bring your real workload

Turn public evidence into a personal decision.

The free Advisor adds your current Mac, context, concurrency, budget, compatibility, and offline requirements. It can say “needs verification” when the public evidence is not enough.

Run the free advisor

Continue the decision

Questions people ask

Clear answers before you buy

Is this a benchmark run by Keep or Upgrade?

No. This is a public-evidence synthesis and planning tool. It combines model cards, public artifact metadata, runtime documentation, published evaluations, and labeled community reports. Only an exact project smoke tuple is marked project verified.

How is Mac capacity calculated?

The tool uses the selected public model artifact size when an exact file or shard total is available. Otherwise it uses a clearly labeled planning estimate. The model budget reserves memory for macOS, runtime overhead, and context; it is not a promise that the workflow will be comfortable.

Do active MoE parameters mean the model needs less memory?

No. Stored weights are the capacity gate. Active parameters help explain per-token compute, but they do not remove the need to keep the full stored model package available.

Can a model with an official context window run that context on every Mac?

No. The context window is model metadata. The selected Mac must also have enough memory after weights, runtime overhead, the operating system, and other applications are accounted for.

Should I buy a Mac from this table?

Use the table to narrow the planning range, then run the free Advisor with your current Mac, context length, concurrency, client, budget, compatibility, and offline requirements. Public evidence is not a first-party benchmark or a purchase guarantee.

Use your own baseline

Get a free recommendation for your actual bottleneck.

The advisor asks for the hardware, workload, budget, and compatibility details that generic buying guides cannot see.

Start the free advisor

Keep or Upgrade is an independent decision-support tool. Apple product names identify the products being compared; Apple does not sponsor or endorse this site. Specifications and vendor test claims can change, so confirm local availability, pricing, and compatibility before purchasing.