Current v0.8.0 release
macMLX v0.8.0
The current release hardens the tiered SSD KV cache end to end: the cold tier is now byte-bounded, weight-safe, survives a restart, and no longer stalls other requests while it spills.
Compatibility and upgrade notes
- Compatibility
- The hardening preserves the cache's boundaries: reuse is still exact longest-prefix only, with no released block sharing or paged KV allocation; a missing, unreadable, or version-mismatched cold index degrades to exact re-hits rather than wrong output, and a weight swap deletes stale entries instead of restoring them.
- Upgrade
- Review the tagged v0.8.0 changelog before upgrading; on the first launch after this change an over-budget cold cache directory is trimmed to the Cold (SSD) budget, which only removes regenerable cache, never model files or settings.
Shipped
Apple Silicon macOS installation
macMLX supports Apple Silicon Macs running macOS 14 or later.
Use the current project installation and Gatekeeper guidance; do not disable system-wide security protections solely to open the app.
Verified
Swift in-process inference
The default engine loads and runs MLX models inside the Swift process.
Model loading, generation, caching, and serving use Apple MLX through MacMLXCore; the default inference path does not require a Python runtime.
Verified
Apple Silicon unified memory
MLX arrays use the Mac's shared CPU/GPU memory system.
Unified memory reduces explicit transfers between CPU orchestration and integrated-GPU compute, but model weights, activations, and KV cache still consume finite physical memory.
Verified
Shared code, process-local engines
The app and CLI both import MacMLXCore, which owns inference and the server.
The products share implementation and behavior. When the app and CLI run as separate processes, they do not share one in-memory engine instance.
Verified
No Python on the default path
The released default engine is Swift-native and needs no Python runtime.
Optional compatibility engines may use subprocesses and other runtimes. This is not a claim that Python is absent everywhere in the project.
Verified
OpenAI endpoint compatibility
Chat, legacy completions, model listing, and embeddings use compatible request and response shapes.
Compatibility is endpoint-specific. macMLX model load and unload routes under /x/models are project extensions, not OpenAI-compatible model management.
Verified
Anthropic Messages compatibility
POST /v1/messages, including streaming, is available in v0.5.3.
This is Messages API compatibility only, not compatibility with the full Anthropic API.
Verified
Selected Ollama endpoints
macMLX supports /api/version, /api/tags, /api/show, /api/chat, and /api/generate.
The compatibility layer has shipped since v0.3.7. It is a selected endpoint set, not a drop-in replacement for every Ollama API.
Verified
MCP server
The CLI can expose local inference to MCP clients.
The MCP server shipped in v0.5.0 and is separate from chat-side routing to external tools.
Verified
MCP client pool
v0.5.3 includes managed MCP client connections.
The pool manages external MCP processes and connections; integrated chat-side tool routing is a separate capability released in v0.6.0.
Verified
Integrated chat tool routing
v0.6.0 ships multi-turn tool routing for OpenAI, Anthropic, and the GUI MCP loop.
Protocol-specific validation keeps tool-call histories explicit; this routing is distinct from the MCP server and client-pool infrastructure.
Verified
Local embeddings
POST /v1/embeddings shipped in v0.5.3.
Encoder-family model detection exists, while using an unsuitable chat model can still produce vectors without semantic guarantees.
Verified
Bi-encoder rerank MVP
POST /v1/rerank scores independently embedded texts with cosine similarity.
This released MVP is not a cross-encoder reranker.
Verified
Exact-prefix RAM and SSD cache
A hot RAM tier and content-addressed SSD cold tier support promotion and demotion.
The released v0.5.0 cache reuses exact full prefixes. It does not provide released block sharing or paged KV allocation.
Verified
Trie longest-prefix reuse
v0.6.0 reuses the longest compatible cached token prefix.
Multi-turn prompts can trim to the longest common prefix and incrementally prefill only the new suffix while usage retains full-prompt accounting.
Verified
Bounded model pool
Budgets, LRU eviction, pinning, cold swap, idle TTL, and probes bound multi-model use.
The pool shipped in v0.5.0 and was hardened in v0.5.3. It is not a unified adaptive controller.
Verified
Supported LoRA adapters
The native engine can apply supported LoRA adapters.
Adapter compatibility depends on the base architecture and weights; universal LoRA compatibility is not claimed.
Verified
Fourteen detected VLM families
The model library detects 14 vision-language model_type families.
This is an evidence-backed family count, not a guarantee that every checkpoint or processor variant will load.
Verified
DeepSeek V3.2 Swift overlay
v0.5.3 includes pure-Swift component parity for the DeepSeek V3.2 architecture.
A real-checkpoint smoke test remains pending and FP8 dequantization is absent, so this is not an end-to-end or universal MoE claim.
Verified
Eligibility-gated continuous batching
v0.6.0 batches only eligible dense-text requests under real concurrency, with automatic serial fallback.
The tagged 4-client benchmark measured 2.5–3.2× aggregate throughput. VLM, speculative decoding, Ollama, Anthropic, and embeddings remain serial.
Verified
Fixed prefill admission throttle
A fixed prefillBatchSize bounds rows admitted per scheduler step.
The released throttle is fixed configuration, not the planned adaptive memory controller.
Verified
Structured output
v0.6.0 supports response_format with json_object and an explicit JSON Schema subset.
Unsupported schema keywords return 400. VLM with structured output and tools with structured output are unsupported combinations and are explicitly rejected rather than silently degraded.
Verified
Speculative decoding
v0.6.0 ships the classic draft-model path through both the API and GUI.
Acceptance telemetry reports draft efficiency. Targets with non-trimmable hybrid or linear-attention caches are detected and fall back to standard decoding.
Verified
API compatibility pack
v0.6.0 adds logit_bias, logprobs and top_logprobs, XTC, per-request LoRA adapters, and tools.
An explicit compatibility matrix governs parameter combinations, and unsupported pairs return 400 instead of silently degrading.
Verified
KV-cache quantization
v0.6.0 exposes kv_bits, kv_group_size, and quantized_kv_start for compatible requests.
These controls change KV-cache precision, not model-weight precision. Nonpositive kv_bits disables the feature, and the compatibility matrix rejects unsupported model or request combinations.
Verified
Hugging Face cache discovery
v0.6.0 can discover models in configured Hugging Face cache roots without downloading duplicate weights.
Discovery confirms a local candidate, not a universal load guarantee; architecture, tokenizer, processor, and checkpoint compatibility still apply.
Verified
Per-model chat-template overrides
v0.6.2 resolves templates in user file, built-in model_type, then checkpoint order.
Built-in overrides carry standard-path parity evidence, but template support is not universal; checkpoint-specific branches and unsupported swift-jinja syntax can still define the boundary.
Verified
Track G tested models
v0.6.2 adds four checkpoint-tested native model families.
Measured real-checkpoint generation: Seed-OSS-36B 4-bit at 18.2 tok/s; Hunyuan V1 Dense 1.8B 4-bit at 80.3 tok/s; Cohere Command R7B 7B 4-bit at 21.7 tok/s; and MiniCPM3-4B 4-bit at 18.7 tok/s. Results are checkpoint-specific, not family-wide performance guarantees.
Verified
InternLM3 theoretical support
v0.6.2 ships parity-verified InternLM3 code at the theoretical support tier.
Real generation has not been demonstrated. Public checkpoints ship tokenizer.model but no tokenizer.json, while the Swift tokenizer stack requires tokenizer.json; support remains theoretical until that load-path boundary changes.
Verified
Temperature and top-p
Temperature and nucleus top-p sampling are released controls.
These are the current exposed core sampling controls.
Verified
Silicon Activity panel
A new main-window Activity tab shows live Apple Silicon readouts, a prefill/decode throughput split, and the current inference bottleneck with advice.
The panel reads GPU occupancy, memory bandwidth, thermal and memory pressure, and per-rail power without admin rights or a helper process. It is observability, not a performance guarantee: an unavailable counter renders an em dash with its reason, an idle engine shows no active generation, and estimated values are labeled rather than presented as measured.
Verified
Inference bottleneck classifier
v0.7.0 fuses the hardware samples with the engine's live prefill or decode phase to attribute what limits generation.
Because the classifier runs in-process it knows which inference phase is active, the signal an external GPU monitor cannot see. It is a smoothed heuristic, not a profiler: it prioritizes memory over thermal over compute or bandwidth, applies hysteresis and three-frame smoothing, and self-calibrates its bandwidth ceiling, so a verdict is guidance rather than a ground-truth measurement.
Verified
Sudoless silicon sampling
A runtime IOReport bridge plus samplers read GPU occupancy, memory bandwidth, thermal and memory pressure, and ANE power.
The IOReport bridge is resolved through dlopen, so a future macOS that renames or removes the private framework degrades the readouts to unavailable rather than failing to launch. Memory bandwidth is reported as an estimate, ANE exposes a power proxy with no fabricated utilization percentage, and the Media Engine, which has no duty-cycle signal, is omitted rather than guessed.
Verified
Benchmark bottleneck attribution
A benchmark run samples the silicon while generating and attaches what limited its decode steady state to the saved result.
The attribution names the decode limiter, whether memory, thermal, bandwidth, or compute, with a confidence and the representative hardware behind it. A run too short to attribute is reported as unavailable instead of inventing a verdict, and low-confidence or estimate-based calls are flagged; it describes one measured run, not a family-wide benchmark guarantee.
Verified
OCR model recognition
Dedicated OCR checkpoints receive an OCR badge distinct from the generic Vision badge, and GLM-OCR is verified end-to-end.
GLM-OCR loads through the stock VLM path with no new model code and reads image text back. The badge is intentionally narrow: it appears only on a model that actually loads, so an OCR family that is detected but not yet ported, such as dots_ocr or deepseek-ocr, earns no badge. This is model recognition, not a universal OCR-quality claim.
Verified
Byte-bounded cold KV tier
v0.8.0 bounds the on-disk cold KV cache to an explicit byte budget, pruned oldest-first, with a toggle to opt the cold tier out.
The previously unbounded cold directory now honors the Cold (SSD) budget with the Hot (RAM) budget wired alongside; an over-budget directory is trimmed on first launch, removing only regenerable cache rather than model files or settings. It does not add block sharing or paged allocation; the released cache still reuses exact full prefixes only.
Verified
Weight-identity-guarded cold entries
v0.8.0 fingerprints each cold entry against model weight identity so a weight swap cannot serve stale KV.
The fingerprint covers the model's config plus every safetensors shard; if the same path later holds re-downloaded, re-quantized, or swapped weights, the stale entry is rejected and deleted rather than restored into silently wrong output. It is an integrity guard for the cold tier's reuse, not a correctness proof for every cache path, and it does not change the released exact-prefix reuse semantics.
Verified
Persistent cross-session cold index
v0.8.0 persists a cold index so longest-prefix reuse survives a restart across sessions.
A follow-up turn that extends an earlier session's prompt reuses the shared prefix straight off disk, while a missing, unreadable, or version-mismatched index degrades to exact re-hits, never to wrong output. Reuse remains exact-prefix, not the planned paged or block-sharing virtualization; the persisted index only extends the existing longest-prefix reuse across process restarts.
Verified
Off-actor cold-tier writer
v0.8.0 moves cold-tier writes off the cache actor onto a dedicated serial background writer.
Persisting an evicted KV snapshot, hundreds of milliseconds for a large cache, no longer blocks another request's cache lookup; the write lands atomically through a temp file and rename, and the on-disk result is unchanged. It relocates disk writes to keep the store actor responsive and does not alter what is cached or the released reuse semantics.
Verified
Current limitations
Eligibility-gated continuous batching
v0.6.0 batches only eligible dense-text requests under real concurrency, with automatic serial fallback.
The tagged 4-client benchmark measured 2.5–3.2× aggregate throughput. VLM, speculative decoding, Ollama, Anthropic, and embeddings remain serial.
Verified
Structured output
v0.6.0 supports response_format with json_object and an explicit JSON Schema subset.
Unsupported schema keywords return 400. VLM with structured output and tools with structured output are unsupported combinations and are explicitly rejected rather than silently degraded.
Verified
InternLM3 theoretical support
v0.6.2 ships parity-verified InternLM3 code at the theoretical support tier.
Real generation has not been demonstrated. Public checkpoints ship tokenizer.model but no tokenizer.json, while the Swift tokenizer stack requires tokenizer.json; support remains theoretical until that load-path boundary changes.
Verified
DeepSeek V3.2 Swift overlay
v0.5.3 includes pure-Swift component parity for the DeepSeek V3.2 architecture.
A real-checkpoint smoke test remains pending and FP8 dequantization is absent, so this is not an end-to-end or universal MoE claim.
Verified
Sudoless silicon sampling
A runtime IOReport bridge plus samplers read GPU occupancy, memory bandwidth, thermal and memory pressure, and ANE power.
The IOReport bridge is resolved through dlopen, so a future macOS that renames or removes the private framework degrades the readouts to unavailable rather than failing to launch. Memory bandwidth is reported as an estimate, ANE exposes a power proxy with no fabricated utilization percentage, and the Media Engine, which has no duty-cycle signal, is omitted rather than guessed.
Verified
Planned
Paged KV, block sharing, and CoW
Paged allocation, shared blocks, and copy-on-write branching are planned.
None of these cache-virtualization features is released in v0.8.0.
Verified
Unified adaptive memory guard
A feedback controller across cache, model pool, and concurrency is planned.
Released memory probes and pool caps are separate mechanisms and must not be described as this guard.
Verified
Expanded sampling controls
top-k, min-p, presence, frequency, and repetition penalties, plus per-request seed are planned.
DeepSeek expert-routing top-k is an internal architecture operation and is unrelated to user sampling top-k.
Verified
Official sources
- macMLX v0.8.0 · github.com
- Apple Silicon macOS installation
- KV-cache quantization
- Apple Silicon unified memory
- Selected Ollama endpoints
- macMLX v0.8.0 · github.com
- MCP client pool
- Integrated chat tool routing
- Local embeddings
- Bi-encoder rerank MVP
- Weight-identity-guarded cold entries
- Trie longest-prefix reuse
- Bounded model pool
- Track G tested models
- DeepSeek V3.2 Swift overlay
- Fixed prefill admission throttle
- Eligibility-gated continuous batching
- Structured output
- Speculative decoding
- KV-cache quantization
- Per-model chat-template overrides
- macMLX v0.8.0 · github.com
- InternLM3 theoretical support
- Temperature and top-p
- Silicon Activity panel
- Silicon Activity panel
- Inference bottleneck classifier
- Inference bottleneck classifier
- Sudoless silicon sampling
- Sudoless silicon sampling
- Benchmark bottleneck attribution
- Benchmark bottleneck attribution
- OCR model recognition
- Byte-bounded cold KV tier
- Weight-identity-guarded cold entries
- Persistent cross-session cold index
- Off-actor cold-tier writer
- Expanded sampling controls
- Expanded sampling controls