{"schemaVersion":"public-local-model-evidence-v3","dataDate":"2026-08-28","description":"Latest-model public evidence synthesis for local model and Mac memory planning, including attributable community signals; not a first-party benchmark.","records":[{"id":"qwen-small","family":"qwen","tier":"small","modelId":"qwen3.8-27b","modelVersion":"Qwen3.8","displayName":"Qwen3.8 27B","architecture":"dense","totalParamsB":27,"activeParamsB":null,"effectiveParamsB":null,"contextWindowTokens":262144,"quantization":"Q4_0","format":"GGUF","artifactUrl":"https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/main/Qwen3.8-27B-Q4_0.gguf","artifactSizeBytes":16056478688,"artifactSizeGb":16.06,"sizeBasis":"exact-public-file-metadata","modelSourceUrl":"https://github.com/QwenLM/Qwen3.8","runtime":"mlx","runtimeSourceUrl":"https://github.com/QwenLM/Qwen3.8","runtimeEvidenceLevel":"official-support","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json","checkedAt":"2026-08-28","note":"Public model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://github.com/QwenLM/Qwen-Agent","checkedAt":"2026-08-28","note":"The current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[{"platform":"reddit","sourceUrl":"https://www.reddit.com/r/ollama/comments/1vrqi3d/qwen3827b_on_a_24gb_m4_pro_mac_mini_benchmarks/","checkedAt":"2026-08-28","evidenceType":"first-person-run","title":"Qwen3.8-27B on a 24GB M4 Pro Mac mini","hardware":"Mac mini · M4 Pro","memoryGb":24,"runtime":"llama.cpp b10488","quantization":"GGUF Q4_K_M / IQ4_XS","observed":"Q4_K_M 17.77GB; about 11.4 tok/s decode; about 16.6GB resident with q8 KV at 32K context","confidence":"medium","summary":"A first-person report says the 27B model is usable on 24GB, but the memory ceiling is visible once context and KV cache are included.","limitations":"One user, one build, and one workload; not a Keep or Upgrade test."},{"platform":"media","sourceUrl":"https://blog.dinosaurliu.com/2026/08/20/M1_Pro_32GB%E6%9C%AC%E5%9C%B0%E9%83%A8%E7%BD%B2Qwen3.8_27B%E5%AE%9E%E5%BD%95/","checkedAt":"2026-08-28","evidenceType":"first-person-run","title":"Qwen3.8-27B 4-bit on a 32GB M1 Pro","hardware":"MacBook Pro · M1 Pro","memoryGb":32,"runtime":"MLX-VLM","quantization":"MLX 4-bit","observed":"8–8.7 tok/s; weights about 16.05GB; peak memory about 18.5–21.7GB","confidence":"medium","summary":"A detailed local deployment log reports a stable 32GB starting point and explicitly discourages 16GB for this 27B representative.","limitations":"Community log with different thermals, context, and prompt mix from the product matrix."}]},{"id":"qwen-mainstream","family":"qwen","tier":"mainstream","modelId":"qwen3.8-flash-next-180b","modelVersion":"Qwen3.8-Flash-Next","displayName":"Qwen3.8-Flash-Next 180B (6B active)","architecture":"moe","totalParamsB":180,"activeParamsB":6,"effectiveParamsB":null,"contextWindowTokens":262144,"quantization":"BF16","format":"Safetensors","artifactUrl":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next","artifactSizeBytes":360000192888,"artifactSizeGb":360,"sizeBasis":"exact-public-file-metadata","modelSourceUrl":"https://github.com/QwenLM/Qwen3.8","runtime":"mlx","runtimeSourceUrl":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next","runtimeEvidenceLevel":"community-reported","confidence":"medium","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json","checkedAt":"2026-08-28","note":"Public model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://github.com/QwenLM/Qwen3.8","checkedAt":"2026-08-28","note":"The current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"official-support","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json","checkedAt":"2026-08-28","note":"The current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[{"platform":"reddit","sourceUrl":"https://www.reddit.com/r/LocalLLaMA/comments/1vzqubc/deepseek_v4_0731_qwen_38_flash_glm_53_flash_and/","checkedAt":"2026-08-28","evidenceType":"comparison","title":"Latest-model comparison thread (non-Apple hardware)","hardware":"4× DGX Spark / other non-Apple systems","memoryGb":null,"runtime":null,"quantization":null,"observed":"Users discuss Qwen3.8 Flash as an exploration/subagent model; no Mac tuple or reproducible throughput is provided.","confidence":"low","summary":"This is a freshness and interest signal only, not evidence that Flash Next fits a Mac configuration.","limitations":"Cross-hardware discussion; excluded from Mac capacity conclusions."}]},{"id":"qwen-boundary","family":"qwen","tier":"boundary","modelId":"qwen3.8-2.4t-a95b","modelVersion":"Qwen3.8","displayName":"Qwen3.8 2.4T-A95B","architecture":"moe","totalParamsB":2400,"activeParamsB":95,"effectiveParamsB":null,"contextWindowTokens":262144,"quantization":"Q4_K_M","format":"GGUF","artifactUrl":"https://github.com/QwenLM/Qwen3.8","artifactSizeBytes":null,"artifactSizeGb":1440,"sizeBasis":"parameter-planning-estimate","modelSourceUrl":"https://github.com/QwenLM/Qwen3.8","runtime":"mlx","runtimeSourceUrl":"https://github.com/QwenLM/Qwen3.8","runtimeEvidenceLevel":"official-support","confidence":"medium","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/raw/main/config.json","checkedAt":"2026-08-28","note":"Public model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://github.com/QwenLM/Qwen-Agent","checkedAt":"2026-08-28","note":"The current Qwen material documents tool or agent paths. The exact model template and runtime still need a workflow check."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[]},{"id":"deepseek-mainstream","family":"deepseek","tier":"mainstream","modelId":"deepseek-v4-flash-284b-a13b","modelVersion":"DeepSeek V4","displayName":"DeepSeek-V4-Flash 284B-A13B","architecture":"moe","totalParamsB":284,"activeParamsB":13,"effectiveParamsB":null,"contextWindowTokens":1048576,"quantization":"MLX-4bit","format":"MLX","artifactUrl":"https://deepseek.com/en/news/v4-preview/","artifactSizeBytes":null,"artifactSizeGb":170.4,"sizeBasis":"parameter-planning-estimate","modelSourceUrl":"https://deepseek.com/en/news/v4-preview/","runtime":"mlx","runtimeSourceUrl":"https://huggingface.co/mlx-community/DeepSeek-V4-Flash-4bit","runtimeEvidenceLevel":"community-reported","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/raw/main/config.json","checkedAt":"2026-08-28","note":"Public model metadata lists 1,048,576 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://deepseek.com/en/news/v4-preview/","checkedAt":"2026-08-28","note":"The official DeepSeek V4 material documents tool use in the model family. The exact local template and runtime still need a workflow check."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[{"platform":"reddit","sourceUrl":"https://www.reddit.com/r/LocalLLM/comments/1veo5e3/a_few_deepseek_v4_flash_runs_on_apple_m5_max_128gb/","checkedAt":"2026-08-28","evidenceType":"first-person-run","title":"DeepSeek V4 Flash on an M5 Max 128GB","hardware":"Mac · Apple M5 Max","memoryGb":128,"runtime":"ds4 / SSD-streaming path","quantization":"DwarfStar IQ2XXS","observed":"About 81GB model working set and about 31.06 tok/s in the reported run","confidence":"medium","summary":"A direct Apple Silicon report shows a practical 128GB path for the Flash representative.","limitations":"Community benchmark; exact context, thermal state, and build can change the result."},{"platform":"media","sourceUrl":"https://ronreiter.com/posts/running-deepseek-v4-flash-locally/","checkedAt":"2026-08-28","evidenceType":"first-person-run","title":"DeepSeek V4 Flash on an M3 Max 128GB","hardware":"MacBook Pro · M3 Max","memoryGb":128,"runtime":"ds4 experimental fork","quantization":null,"observed":"About 21 tok/s and about 81GB resident in the reported run","confidence":"medium","summary":"An independent technical write-up corroborates that 128GB-class Apple Silicon can run a local Flash path.","limitations":"The article is tied to an experimental fork and does not validate mainline Ollama or llama.cpp."},{"platform":"reddit","sourceUrl":"https://www.reddit.com/r/LocalLLaMA/comments/1vge4l5/deepseek_v4_flash_0731_at_1017_ts_nothink_on/","checkedAt":"2026-08-28","evidenceType":"first-person-run","title":"DeepSeek V4 Flash on 64GB Macs with SSD streaming","hardware":"64GB Apple Silicon Macs","memoryGb":64,"runtime":"ds4 SSD streaming","quantization":"IQ2XXS / 2-bit class","observed":"Reports cluster around roughly 10–15 tok/s; technically possible, but not a comfortable in-memory workflow","confidence":"medium","summary":"64GB appears technically viable only with aggressive quantization and storage streaming, so it is kept as a constrained edge case.","limitations":"Mixed community reports and different Mac generations; do not read this as a 64GB recommendation."},{"platform":"github","sourceUrl":"https://github.com/ml-explore/mlx-lm/issues/1332","checkedAt":"2026-08-28","evidenceType":"issue-report","title":"MLX DeepSeek V4 residency growth and resource-limit crash","hardware":"M4 Max 128GB and other Apple Silicon reports","memoryGb":128,"runtime":"MLX-LM","quantization":null,"observed":"Reproduction reports deterministic resource-limit failure around 11,300 generated tokens","confidence":"high","summary":"The issue is important negative evidence: a model can fit and still fail on long generations because of runtime memory behavior.","limitations":"Issue state and fixes can change; check the runtime version and referenced patch before relying on MLX."},{"platform":"github","sourceUrl":"https://github.com/ml-explore/mlx-lm/issues/1662","checkedAt":"2026-08-28","evidenceType":"issue-report","title":"MLX cache/residency failure on an M3 Ultra 512GB","hardware":"Mac Studio · M3 Ultra 512GB","memoryGb":512,"runtime":"MLX-LM","quantization":null,"observed":"Reports failure around 11,456–11,488 generated tokens in a production-style run","confidence":"high","summary":"Even very large unified-memory systems are not automatically stable for long-context MLX runs.","limitations":"This is a runtime bug report, not a capacity limit; a patched release may change the result."},{"platform":"github","sourceUrl":"https://github.com/ggml-org/llama.cpp/issues/25837","checkedAt":"2026-08-28","evidenceType":"issue-report","title":"DeepSeek V4 Flash Q8 garbled output on Apple NEON","hardware":"Apple Silicon CPU / NEON","memoryGb":null,"runtime":"llama.cpp","quantization":"Q8","observed":"A specific repack path produced garbled output","confidence":"high","summary":"This narrows the risk to a quantization/repack/runtime combination instead of treating every llama.cpp path as equivalent.","limitations":"Do not generalize one broken Q8 repack to all DeepSeek V4 Flash formats."},{"platform":"github","sourceUrl":"https://github.com/antirez/ds4","checkedAt":"2026-08-28","evidenceType":"comparison","title":"ds4 Apple Silicon throughput table","hardware":"Mac Studio · M3 Ultra 512GB","memoryGb":512,"runtime":"ds4 Metal/CUDA engine","quantization":"Q4 class","observed":"Repository table reports short/long generation figures around 78.95 / 35.50 tok/s","confidence":"medium","summary":"A purpose-built engine provides a useful upper-bound signal for a large-memory Mac Studio, but it is not a mainstream runtime guarantee.","limitations":"Single graph worker, no batching, and engine-specific optimizations; compare only directionally."}]},{"id":"deepseek-boundary","family":"deepseek","tier":"boundary","modelId":"deepseek-v4-pro-1.6t-a49b","modelVersion":"DeepSeek V4","displayName":"DeepSeek-V4-Pro 1.6T-A49B","architecture":"moe","totalParamsB":1600,"activeParamsB":49,"effectiveParamsB":null,"contextWindowTokens":1048576,"quantization":"MLX-4bit","format":"MLX","artifactUrl":"https://deepseek.com/en/news/v4-preview/","artifactSizeBytes":null,"artifactSizeGb":960,"sizeBasis":"parameter-planning-estimate","modelSourceUrl":"https://deepseek.com/en/news/v4-preview/","runtime":"llama.cpp","runtimeSourceUrl":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro","runtimeEvidenceLevel":"unknown","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/raw/main/config.json","checkedAt":"2026-08-28","note":"Public model metadata lists 1,048,576 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://deepseek.com/en/news/v4-preview/","checkedAt":"2026-08-28","note":"The official DeepSeek V4 material documents tool use in the model family. The exact local template and runtime still need a workflow check."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[]},{"id":"glm-mainstream","family":"glm","tier":"mainstream","modelId":"glm-5.3-flash-320b-a18b","modelVersion":"GLM-5.3-Flash","displayName":"GLM-5.3-Flash 320B-A18B","architecture":"moe","totalParamsB":320,"activeParamsB":18,"effectiveParamsB":null,"contextWindowTokens":1000000,"quantization":"MLX-4bit","format":"MLX","artifactUrl":"https://huggingface.co/orcarouter/GLM-5.3-Flash-MLX","artifactSizeBytes":203992076296,"artifactSizeGb":203.99,"sizeBasis":"exact-public-file-metadata","modelSourceUrl":"https://github.com/zai-org/GLM-5","runtime":"mlx","runtimeSourceUrl":"https://huggingface.co/orcarouter/GLM-5.3-Flash-MLX","runtimeEvidenceLevel":"community-reported","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://docs.z.ai/guides/vlm/glm-5.3-flash","checkedAt":"2026-08-28","note":"Public model metadata lists 1,000,000 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://docs.z.ai/guides/vlm/glm-5.3-flash","checkedAt":"2026-08-28","note":"The official GLM-5.3 material documents function calling and agent capabilities. It is not a Mac runtime integration benchmark."},"json":{"level":"public-benchmark","sourceUrl":"https://github.com/zai-org/GLM-5","checkedAt":"2026-08-28","note":"Public GLM evaluations cover agentic or structured tasks, but they do not establish JSON validity in every client."},"vision":{"level":"official-support","sourceUrl":"https://docs.z.ai/guides/vlm/glm-5.3-flash","checkedAt":"2026-08-28","note":"The current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[{"platform":"huggingface","sourceUrl":"https://huggingface.co/orcarouter/GLM-5.3-Flash-MLX","checkedAt":"2026-08-28","evidenceType":"community-quantization","title":"Community GLM-5.3-Flash MLX quantization ladder","hardware":null,"memoryGb":null,"runtime":"MLX","quantization":"2bit-lite / 2 / 3 / 4 / 6-bit MLX","observed":"Public metadata lists about 102.43GB to 295.60GB per variant; the 4-bit root mirror is about 203.99GB","confidence":"medium","summary":"The newer repository makes the precision trade-off explicit and includes a 2bit-lite path, but it does not prove loadability or speed on a particular Mac.","limitations":"Community conversion; full-repository size includes duplicate root and variant folders, and no trustworthy Apple first-person benchmark is attached."},{"platform":"reddit","sourceUrl":"https://www.reddit.com/r/LocalLLaMA/comments/1vzqubc/deepseek_v4_0731_qwen_38_flash_glm_53_flash_and/","checkedAt":"2026-08-28","evidenceType":"comparison","title":"GLM-5.3 Flash comparison thread (non-Apple hardware)","hardware":"Dual DGX Spark / other non-Apple systems","memoryGb":null,"runtime":null,"quantization":null,"observed":"One report describes roughly 22 tok/s on dual Spark and a slower, more verbose interaction style","confidence":"low","summary":"This indicates active community comparison but cannot be translated into a Mac recommendation.","limitations":"Cross-hardware, anecdotal, and workload-specific."}]},{"id":"glm-boundary","family":"glm","tier":"boundary","modelId":"glm-5.3-744b-a40b","modelVersion":"GLM-5.3","displayName":"GLM-5.3 744B-A40B","architecture":"moe","totalParamsB":744,"activeParamsB":40,"effectiveParamsB":null,"contextWindowTokens":1000000,"quantization":"MLX-4bit","format":"MLX","artifactUrl":"https://github.com/zai-org/GLM-5","artifactSizeBytes":null,"artifactSizeGb":446.4,"sizeBasis":"parameter-planning-estimate","modelSourceUrl":"https://github.com/zai-org/GLM-5","runtime":"llama.cpp","runtimeSourceUrl":"https://github.com/zai-org/GLM-5","runtimeEvidenceLevel":"unknown","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://docs.z.ai/guides/llm/glm-5.3","checkedAt":"2026-08-28","note":"Public model metadata lists 1,000,000 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"official-support","sourceUrl":"https://docs.z.ai/guides/llm/glm-5.3","checkedAt":"2026-08-28","note":"The official GLM-5.3 material documents function calling and agent capabilities. It is not a Mac runtime integration benchmark."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[]},{"id":"gemma-small","family":"gemma","tier":"small","modelId":"gemma-4-e2b","modelVersion":"Gemma 4","displayName":"Gemma 4 E2B","architecture":"effective-ple","totalParamsB":5.1,"activeParamsB":null,"effectiveParamsB":2.3,"contextWindowTokens":131072,"quantization":"Q4_0","format":"QAT","artifactUrl":"https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-gguf/blob/main/gemma-4-E2B_q4_0-it.gguf","artifactSizeBytes":3349516256,"artifactSizeGb":3.35,"sizeBasis":"exact-public-file-metadata","modelSourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","runtime":"llama.cpp","runtimeSourceUrl":"https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-gguf/resolve/main/gemma-4-E2B_q4_0-it.gguf?download=true","runtimeEvidenceLevel":"community-reported","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","checkedAt":"2026-08-28","note":"Public model metadata lists 131,072 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No model-specific public tool-calling result is strong enough for a positive claim in this snapshot."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"official-support","sourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","checkedAt":"2026-08-28","note":"The current model documentation describes multimodal inputs for this representative; the selected Mac runtime path still needs checking."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[{"platform":"media","sourceUrl":"https://www.tomshardware.com/desktops/exploring-apple-silicons-local-ai-performance-with-the-mac-studio-and-m4-max-m4-max-beats-gb10-and-strix-halo-in-decode-throughput-but-memory-bandwidth-isnt-everything","checkedAt":"2026-08-28","evidenceType":"media-review","title":"Apple Silicon local-AI comparison including Gemma 4","hardware":"Mac Studio / M4 Max 128GB test systems","memoryGb":128,"runtime":"llama.cpp","quantization":null,"observed":"Gemma 4 appears in the comparison; the article warns that bandwidth is not the sole predictor and does not publish this exact matrix tuple","confidence":"low","summary":"Useful context for Apple Silicon behavior, but not an exact Gemma 4 E2B fit or performance claim.","limitations":"Different model/quantization details and media test methodology from this product matrix."}]},{"id":"gemma-mainstream","family":"gemma","tier":"mainstream","modelId":"gemma-4-12b-unified","modelVersion":"Gemma 4","displayName":"Gemma 4 12B Unified","architecture":"dense","totalParamsB":11.95,"activeParamsB":null,"effectiveParamsB":null,"contextWindowTokens":262144,"quantization":"Q4_0","format":"QAT","artifactUrl":"https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/blob/main/gemma-4-12b-it-qat-q4_0.gguf","artifactSizeBytes":6975879296,"artifactSizeGb":6.98,"sizeBasis":"exact-public-file-metadata","modelSourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","runtime":"llama.cpp","runtimeSourceUrl":"https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/gemma-4-12b-it-qat-q4_0.gguf?download=true","runtimeEvidenceLevel":"unknown","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","checkedAt":"2026-08-28","note":"Public model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No model-specific public tool-calling result is strong enough for a positive claim in this snapshot."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[]},{"id":"gemma-boundary","family":"gemma","tier":"boundary","modelId":"gemma-4-26b-a4b","modelVersion":"Gemma 4","displayName":"Gemma 4 26B-A4B","architecture":"moe","totalParamsB":25.2,"activeParamsB":3.8,"effectiveParamsB":null,"contextWindowTokens":262144,"quantization":"Q4_0","format":"QAT","artifactUrl":"https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf/blob/main/gemma-4-26B_q4_0-it.gguf","artifactSizeBytes":14439363584,"artifactSizeGb":14.44,"sizeBasis":"exact-public-file-metadata","modelSourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","runtime":"llama.cpp","runtimeSourceUrl":"https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf","runtimeEvidenceLevel":"community-reported","confidence":"high","checkedAt":"2026-08-28","capabilities":{"context":{"level":"official-support","sourceUrl":"https://ai.google.dev/gemma/docs/core/model_card_4","checkedAt":"2026-08-28","note":"Public model metadata lists 262,144 tokens. This is not proof that the selected Mac can hold that context with the weights."},"concurrency":{"level":"system-inference","sourceUrl":"https://github.com/ggml-org/llama.cpp","checkedAt":"2026-08-28","note":"Concurrency is inferred only from memory budget and runtime architecture. No apples-to-apples public concurrency benchmark is attached."},"tool-calling":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No model-specific public tool-calling result is strong enough for a positive claim in this snapshot."},"json":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Structured JSON behavior depends on the exact model template, runtime parser, and client; no exact public result is registered here."},"vision":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"Vision support is not assumed from family branding or an unrelated model variant."},"sustained":{"level":"unknown","sourceUrl":null,"checkedAt":"2026-08-28","note":"No public, comparable sustained Apple Silicon run is registered for this model and Mac configuration."}},"limitations":["Public file metadata or a parameter estimate is not a first-party benchmark.","Community artifacts and reports do not imply vendor endorsement or Keep or Upgrade testing.","Social evidence is directional: build, context, thermals, storage streaming, and prompt mix can change the result.","Context, concurrency, tool calling, JSON, vision, and sustained behavior remain separate capability checks.","The primary matrix prioritizes current flagship releases; legacy versions remain compatibility records but are intentionally omitted here."],"socialEvidence":[]}]}