Warning
First rule: never crash the laptop. The suite finds each model's safe context ceiling by extrapolating from measurements taken well below the hardware wall — it never probes into the danger zone.
MLX allocates wired (non-swappable) Metal buffers. The real crash ceiling isn't
total RAM — it's the GPU working-set limit (max_recommended_working_set_size), which the
suite reads live on each machine and treats as an exact, measured value (rounding
toward it is how you crash). On the reference M4 Pro testbed it is 17.18 GB (67% of
25.77 GB). With almost no swap free, crossing it can hard-lock the system rather than fail
gracefully. So we predict the wall and stay under it.
- 📏 Danger metric is OS-wired high-water (
vm_stat"Pages wired down"), not MLX'sget_peak_memory(), which undercounts true footprint by ~40% (it excludes the buffer cache, which the OS still wires). - ⚡ The prefill transient spike drives the crash — it grows several× faster than the steady-state KV cache and is invisible in MLX's reported peak.
- 🧬 Only
full_attentionlayers grow KV with context; sliding/linear layers don't.
Two memory classes the suite classifies and handles differently:
| Class | Models | Cache | KV quantization |
|---|---|---|---|
| 🪟 Sliding-window | Gemma, GPT-OSS | RotatingKVCache |
❌ can't — --kv-bits 4 crashes them past 5 k tokens, so they run fp16 (window caps most layers, growth is gentle) |
| 📈 Linear + full | Qwen3.5 9B / 27B | standard | ✅ 4-bit quantizes fine and lowers high-context memory (brief spike at the 5 k threshold) |
uv sync # install deps into .venv
uv run wmx-suite system # show the machine's wall, swap, baselineThe default safety cushion is 2 GB below the machine's wired-memory wall. Set
WMX_SUITE_MARGIN_GB to change it globally for system, health, characterize,
calibrate, and run; an explicit --margin on a command takes precedence:
export WMX_SUITE_MARGIN_GB=3
uv run wmx-suite health
uv run wmx-suite run --margin 2.5 --dry-run --model <hf_id> --prompt "..."The margin must be finite and non-negative. Lower margins reduce the safety cushion; use them only when you understand the hard-lock risk.
| Command | What it does |
|---|---|
uv run wmx-suite system |
Show the machine's wall, swap, baseline |
uv run wmx-suite health |
Live snapshot: current pressure + per-model ✓/✗ go-no-go |
uv run wmx-suite characterize <hf_id> |
Safe probe → fitted context ceiling (--speed quick is ~3× faster, still conservative; standard is the default, full is finer) |
uv run wmx-suite calibrate |
Measure this machine's cold-start memory overhead so pre-flight estimates are accurate on your Apple Silicon SKU (run once per machine; characterize still adapts per model) |
uv run wmx-suite list |
Ceilings for everything characterized; warns about stale fits |
uv run wmx-suite run --model <hf_id> … |
Safely launch mlx_lm.generate |
characterize refuses to launch any probe whose pre-flight base estimate already
exceeds the safe threshold — this is how oversized models (like the 27B) are handled:
predicted, never run into the wall.
run plans a launch and then execs mlx_lm.generate. It:
- picks
--kv-bitsby cache type —4for standard caches, omitted for RotatingKVCache models (Gemma, GPT-OSS) which can't quantize and would otherwise crash; - samples the live settled baseline and caps
--max-kv-sizeat the context wherelive_base + model_base + slope·chits the safe threshold, using the model's measured curve fromsuite.db(or a conservative estimate, with a warning, if uncharacterized); - refuses to launch if the model would breach the wall just to load (e.g. the 27B).
- tokenizes ordinary prompts before launch, warns above 80% of the effective context cap, and refuses prompts above it;
- refuses models such as Qwen3.5 whose custom MLX cache does not currently enforce
--max-kv-size, unless--forceexplicitly accepts the unbounded runtime cache.
# launch safely
uv run wmx-suite run --model mlx-community/Qwen3.5-9B-OptiQ-4bit --prompt "..." --max-tokens 200
# inspect the plan only — no launch
uv run wmx-suite run --dry-run --model <hf_id> --prompt "..."
--forceoverrides a refusal at your own risk;--dry-runprints the plan without launching. Stdin prompts and prompt-cache files require--forcebecause their complete effective prompt cannot be verified by the tokenizer preflight.
Every successful run records its prompt/generation tokens-per-second to the database
(output still streams live — it runs under a PTY so the experience is unchanged). list
then shows the median gen speed per model. Pass --no-log for a bare passthrough.
If artifacts in a model's cached Hugging Face snapshots are newer than its latest
characterization, list and run warn that the fit may be stale. Unused blobs,
mutable refs, and negative-lookup metadata are ignored. The suite does not
automatically re-characterize; review the cache change and run characterize again
before relying on the old ceiling.
Beyond context ceilings, the suite is a memory/perf benchmark lab for two model families.
Each command runs under the same RULE #1 safety gating and records to suite.db:
| Command | What it measures |
|---|---|
uv run wmx-suite benchmark-kokoro |
Kokoro TTS throughput (RTF / chars-per-sec) vs length |
uv run wmx-suite benchmark-kokoro-ttfa |
streaming time-to-first-audio latency |
uv run wmx-suite benchmark-kokoro-batch |
batch concurrency vs throughput |
uv run wmx-suite benchmark-kokoro-voice |
voice-switching latency |
uv run wmx-suite benchmark-kokoro-cache |
voice-cache memory overhead |
uv run wmx-suite benchmark-kokoro-baseline |
static active-synthesis RAM floor |
uv run wmx-suite benchmark-embeddings |
encoder embeddings memory surface (batch × seq_len) |
Run uv run wmx-suite benchmark-<name> --help for options.
wmx_suite/
config.py # validated runtime defaults (e.g. WMX_SUITE_MARGIN_GB)
system.py # device wall, swap, current wired memory
models.py # HF-cache config reader + memory-class classifier
profiles.py # per-machine cold-start constants (calibration)
probe.py # safe characterize/calibrate: ramp + linear fit + ceiling solve
probe_worker.py # ONE isolated (model, context) measurement -> JSON
launcher.py # safe `run` planning + exec of mlx_lm.generate
db.py # SQLite store (context fits, calibration, benchmarks)
ui.py # shared console rendering schema
views/ # per-command output rendering
cli.py # core command entry point
cli_benchmarks.py # benchmark subcommands (Kokoro TTS + embeddings)
embeddings_probe.py # embeddings memory-surface benchmark
kokoro_safety.py # RULE #1 safety gating for the Kokoro workers
probe_worker_kokoro_*.py / probe_worker_embeddings.py # isolated benchmark workers
data/suite.db # results (gitignored)
The SQLite store holds three families of tables: context-ceiling measurement
(models, probe_runs, measurements, fits, generation_log), per-machine
calibration (system_profiles, embedding_profiles), and benchmark results
(kokoro_*, embeddings_*).
v0 scaffold. Validated methodology: predicted Gemma's ceiling to within 0.5% from safe probes. Calibration of the pre-flight base estimate refines as more models are run.
Headless engine. wmx-suite is CLI- and JSON-only — no UI. It is the Apple-Silicon measurement engine behind Project ARA, which wraps it through a thin adapter: the engine measures and returns, the caller persists. Visualization — browsing fits, regression curves, side-by-side ceilings, and the benchmark dashboards — lives in ARA, not here.
Contributions are welcome — especially memory-benchmark results from other Apple Silicon SKUs, which is how the suite becomes trustworthy beyond the reference M4 Pro. See CONTRIBUTING.md. The prime directive applies to every change: never ship something that can crash a machine.
If wmx-suite saved you a kernel panic (or you just like the idea), you can buy me a coffee:
MLX, Apple Silicon, Metal, Mac, and macOS are trademarks of Apple Inc. This project is an independent, community tool — it is not affiliated with, endorsed by, or sponsored by Apple Inc. References to these names are descriptive only, to indicate the technologies the suite works with. All other trademarks are the property of their respective owners.
