Benchmark data, experiment harnesses, and scientific logs for running production AI on local hardware — no cloud required.
This repo accompanies the blog at localfirstai.eu and provides the verifiable evidence backing every claim made there.
| Component | Value |
|---|---|
| Primary machine | Mac Mini M4 Pro (miktam02), 64 GB unified memory |
| Secondary machine | MacBook Pro M5 Max (miktam-mbp) — smart inference tier (Exp 016) |
| Primary model | gemma4:26b (MoE A4B, 25.8B total params / ~4B active per forward pass, Q4_K_M) |
| Router model | gemma4:e4b (fast routing layer) |
| Runtime | Ollama 0.20.2 |
| Orchestration | OpenClaw → Nestor (local AI agent) |
| Operating ceiling (Mini) | > 40,000 tokens on-wire (Exp 008/010 — no cliff at FA=0; Exp 007's 18K was a FA=1 artefact) |
| Operating ceiling (MBP) | > 40,000 tokens on-wire (Exp 007 Phase B — FA=1 cliff at ~45K; FA=0 baseline not yet measured) |
Every claim on the blog is backed by a pre-registered experiment logged in tasks/chronos/scientific_log.md. The methodology: observation → hypothesis → experiment → evidence → conclusion. No retrofitted results.
Roadmap and pending experiments: tasks/chronos/roadmap.md
| # | Name | Status | Key finding |
|---|---|---|---|
| 001 | Verification of Veracity | Complete | Chronos framework activated |
| 002 | Control Plane vs Data Plane | Complete | Thinking mode is flat at ~38 t/s until it isn't — unconstrained prompts trigger runaway |
| 003 | Anonymized Adversarial Memory | Complete | 0/20 source recognition, 0/3 identity leaks on Fight Club corpus — data-sovereignty moat is architectural |
| 004 | Bootstrap Diet | Complete | OpenClaw session hygiene |
| 005 | Router / Reducer Cascade | Phase 0 closed | Working two-model cascade over 8-year Apple Watch corpus; three load-bearing behaviours demonstrated |
| 006 | Redactor Fidelity (GDPR) | Complete | 0/20 × 8 categories — zero true-positive leaks across all pre-registered GDPR categories |
| 007 | Mac Mini vs MacBook Pro M5 Max | Phase A+B complete | Mini cliff ~18K, MBP cliff ~45K (2.5×); MBP gen t/s +200–370%; H1+H2 confirmed ⚠ see Exp 008 |
| 008 | Flash Attention + q8_0 KV Cache | Complete — landmark | FA+q8_0 flags cause the cliff; FA=0/fp16 has no cliff through 40K. Exp 007 ceiling was an artefact |
| 009 | Adversarial Project Critic (Local vs. Frontier) | Complete — FAIL | gemma4:26b matched compliance layer (DPA/DSAR/DPIA) but missed impl-vs-docs gaps; 50% overlap, 50% FP rate |
| 010 | FA vs q8_0 Factorial Isolation | Complete | FA=1 is the sole culprit (cliff at 32.5K alone, 20K combined with q8_0). q8_0 alone: no cliff, +5% gen t/s |
| 011 | MLX Runtime vs Ollama — Context Cliff | Complete | No cliff through 40K on MLX. Prefill matches Ollama FA=0 within 3%. Cliff is an Ollama FA artefact, not a hardware limit. |
| 012 | Cost vs Capability: Where the Curve Breaks | Complete ⚠ see Exp 012-Alpha | gemma4:26b (MoE A4B, ~4B active) 0/8, Haiku/Sonnet/Opus all 5/8 net. Cliff confirmed at the 4B-active/frontier boundary; dense local 12–32B class untested (→ Exp 015). Haiku is cost-dominant; Haiku→Opus is 6.4× cost, 0 score gain. |
| 013 | Local Audit Loop: Can Scaffolding Move gemma4:26b Off Zero? | Complete | H confirmed (partial): decomposition recovered 2/3 findable items (Items 3+4 stable, Item 1 systematic bridge failure). Raw 2/5; context expansion deferred as out of scope for H3. Pre-filter architecture: local audit + Haiku top-up = ~$0.02/audit. |
| 014 | Capability Variance Floor | Complete | H1 CONFIRMED: gemma ≤0/8 all 5 reps (goes negative via systematic DSR FP). H3 CONFIRMED: no overlap (gemma max 0, Haiku min 1). H2 FALSIFIED: Haiku scores 1/1/4/1/2 — not stable at 4–6/8; the 5/8 Exp 012 result was a high-tail sample. Haiku mean ~2/8 across reps; variance driven by DSR/DPA FP penalties and B3 blind spot. |
| 015 | Active-Parameter Ablation: Dense vs MoE | Pre-registered | Does a dense 12–32B local model outperform gemma4:26b (MoE A4B) on the Exp 012 audit rubric? Bottleneck hypothesis: active compute per token, not total params. |
| 016 | Two-Mac Orchestration | Phase B in progress | Phase A: Qwen3-Coder-Next-4bit selected (4.0/5, 100.6 tok/s). Phase B: LAN path mini→MBP confirmed (miktam-mbp.local:8080). TTFT benchmark pending. |
| 017 | Argos Phase 0 — Feed Reconnaissance | Pre-registered | DGT NAP (DATEX II) + OpenChargeMap recon, Málaga–Gibraltar EV corridor. 24h cadence measurement required before Phase 1. |
| 018 | Sovereignty Resilience | Complete | Daemon-down and weights-removed both confirmed: graceful fallback, zero data loss, 1-command recovery. Network-cut was inconclusive by design — the only access to the test machine is remote, so blocking all outbound also blinded the observer, forcing a full reboot. Sovereignty of inference ≠ sovereignty of observability. |
| 019 | Adversarial Legal Review Pipeline | Complete | 0/7 claims survived unchanged across 3 review rounds (Claude adversarial panel + GPT-4o citation check + Gemini barrister review). Two critical issues caught: UK-US BDAA omission; ingestion transfer gap. First application of the Legal Agent pattern. |
| 020 | Hardening and Red-Teaming the Inference Node | Complete | Unauthenticated Ollama on the Tailscale network fixed (loopback-only + token-authenticated proxy); an unrestricted tail NOPASSWD grant confirmed as a genuine root-read primitive; the static listing's worst-looking finding (cp) turned out to be a misread, refuted by direct invocation. |
| 021 | Independent Red-Team Pass on the Inference Node | H3 confirmed | Zero-escalation reachability: without any sudo or exploit, an unpassphrased SSH key and three live credentials are readable by anything running as the automation account — one cat of a shell startup file. A lower bar than exp_020's sudo findings. Rotation declined as a documented accepted risk by the owner; the evidence records names + permissions only, never values. |
| 022 | Adversarial Red-Team of the CasaSol Guide Bot | Complete | H1/H2/H4 (prompt injection, session extraction, guardrail bypass) refuted — the bot resisted every attack, mostly rejected at the intent-router layer. H3 (corpus poisoning via /witness) confirmed: one admin approval is enough to get a fabricated claim indexed and reproduced as fact across every follow-up query. |
| 023 | Generation Efficiency Across the Local Model Family | Complete | Corrected re-run (fixed OLLAMA_CONTEXT_LENGTH, current roster): gemma4:26b posts the best wall-clock of all four models tested — 31.6s, beating even the smaller gemma4:e4b (32.9s) via fewer tokens per answer — while running 4.6x faster than the now-fixed gemma4:31b (144.7s) and 6.7x faster than the brand-new qwen3.8:27b (211.4s). Set against qwen3.8:27b finding more audit-rubric gaps (Exp 015), this is a real speed-vs-thoroughness tradeoff, not a clean win for either model. |
| 024 | Vision Capability: gemma4:26b vs qwen3.8:27b | Inconclusive | gemma4:26b parsed valid JSON on only 1/5 photos on Pharos's exact production vision prompt, with one response degenerating into a token-budget-filling repetition loop — a real failure mode worth flagging to Pharos directly. qwen3.8:27b's run returned 404 on every call because the model had silently vanished from disk mid-experiment — daemon never restarted, disk never filled; cause unexplained. |
| 025 | Context-Allocation Isolation | H1+H2 confirmed, H3 superseded | A clean 2-rep sweep confirms it's a real cliff, not a decline: gemma4:31b holds ~13.3 tok/s from 2048-16384 context, then collapses to ~0.1 tok/s at 262144; gemma4:26b stays flat ~60-64 tok/s across the same full range. Root cause: OLLAMA_CONTEXT_LENGTH was never set, so Ollama auto-selected 256k given this Mac's unified memory — a daemon-wide default hitting every call, not a per-model quirk. Fixed by setting it to 8192 in the launchd plist; verified gemma4:31b now runs at ~13-14 tok/s with zero code changes anywhere. |
| 026 | Contextual Retrieval on the COAPI Corpus, Fully Local | Complete (pilot) | Anthropic's "Contextual Retrieval" technique, tested fully locally on CasaSol's real naive-chunked COAPI corpus. Exact-hit retrieval rate climbs monotonically: unprefixed 0.455 → gemma4:e4b-prefixed 0.545 → gemma4:26b-prefixed 0.636 (n=11, directional). The two models cost almost the same time per chunk here (prompt-eval-bound, not generation-bound), so there's no speed reason to pick the cheaper model over the better one. Extrapolated full-corpus cost: ~4.8h, an easy overnight job. |
| 035 | COAPI Voice — Form in the Weights, Facts in Retrieval, On a Phone | Complete — thesis refuted | Three LoRA runs (Qwen3-1.7B and 4B, adapter strength varied 25×) installed the professional's form and lowered the base's use of retrieved facts every time (with-substance correctness 1.50 → 1.18/1.21, native ES 9/12 → 3/12). Quantisation, epochs, truncation, decoding and adapter strength each ruled out by pre-registered diagnostics. The reference build has no trained weights: untrained Qwen3-4B behind BM25 over our study notes and deterministic post-processing meets every format floor at the base's substance (1.50). Gemma 3 4B probed as a swap: level on substance, better PL/ES fluency, at the iPhone's memory cap — declined. |
Production runs of the Adversarial Watcher — a staged local LLM pipeline that compares documented intent against shipped artefacts and produces annotated gap reports. Each run is an auditable Chronos artefact.
| Run | Project | Confirmed gaps | False positives | Evidence |
|---|---|---|---|---|
| watcher_run_001 | CasaSol | 5 | 3 | 2026-06-06, gemma4:26b, ~270s |
| # | Name | Finding |
|---|---|---|
| 003-Alpha | Memory Bandwidth Cliff | Prefill on gemma4-think:26b goes super-quadratic past ~25K tokens on Apple Silicon. Hard operational ceiling: < 22K tokens on-wire. The bottleneck is memory bandwidth, not VRAM. |
Early benchmarks that preceded the Chronos framework. Scripts and results in tasks/chronos/experiments/.
| Script | What it measures |
|---|---|
nestor-bench-phase1.sh |
Context window (4K–130K) vs generation speed. Finding: gen_tps flat at ~41 t/s. |
nestor-bench-phase1b.sh |
Thinking mode token overhead. Finding: 5–15× token cost for zero quality gain on simple tasks. |
nestor-bench-phase2-compare.sh |
Compressed-memory retrieval vs raw context. |
nestor-bench-phase2-memory.sh |
Memory layer latency at scale. |
nestor-bench-phase2b-retrieval.sh |
Retrieval accuracy across compression levels. |
ollama pull gemma4:26b
chmod +x tasks/chronos/experiments/nestor-bench-phase1.sh
./tasks/chronos/experiments/nestor-bench-phase1.sh
# Results written to tasks/chronos/experiments/results/-
The prefill cliff was an artefact of
OLLAMA_FLASH_ATTENTION=1, not a hardware limit. Exp 007 measured a cliff at ~18K tokens on Mac Mini M4 Pro — but those runs were made with FA=1+q8_0 enabled (Incident 007-Alpha). Exp 008 established the FA=0 baseline: no cliff through 40K tokens. Exp 010's 2×2 factorial confirmed FA=1 is the sole culprit: it alone (without q8_0) produces a cliff at 32.5K and triples prefill latency at 15K tokens (1.774 → 5.405 ms/tok). Under optimal config (FA=0, q8_0), the Mac Mini's true operational ceiling is > 40K tokens on-wire. Flash Attention was designed for discrete GPU SRAM/HBM hierarchies; on Apple Silicon unified memory, its tiling overhead applies without the bandwidth benefit. (Incident 003-Alpha, Exp 007, Incident 007-Alpha, Exp 008, Exp 010) -
Thinking tokens are expensive — cap and name them explicitly.
gemma4:26bis a thinking model: left unconstrained, a simple task generates 10,000–25,000 hidden thinking tokens — at 38 t/s that's 4–11 minutes per response with zero quality gain. The architectural response: agemma4-think:26bOllama alias with a hard 128K context cap, used only for tasks that genuinely need deliberation. The name makes the choice visible; the cap prevents runaway. (Exp 002) -
Data sovereignty is an architectural property, not a policy. An anonymization boundary enforced by the import graph — not by a prompt or a config flag — defeated source recognition (0/20) and identity bridging (0/3) on a corpus the model has memorised. The moat is the architecture. (Exp 003)
-
A two-model cascade extends the operating envelope. Router (
gemma4:e4b) routes in ~3–4s. Reducer (gemma4:26b) synthesises only what fits below the 22K cliff. The cascade made an 8-year health corpus queryable on local hardware without hitting the bandwidth cliff on normal queries. (Exp 005) -
The 22K ceiling is a property of the hardware, not a bug. Memory bandwidth saturates during prefill on the M4 Pro's unified memory architecture. Mitigations: cliff-aware coarsening in the extractor, hard token budgets in the cascade, streaming watchdog for booth/production use.
-
The M5 Max die is in a different performance class for inference. At 25K tokens, MBP gen t/s is 66 vs Mini's 14 — a 4.7× difference on the same model weights and quantisation. MBP at 35K tokens (1.24 ms/tok prefill) is still well below the Mini's baseline at 4K tokens (3.03 ms/tok). The cascade's 22K bundle ceiling — set for the Mini — is comfortably safe on the MBP, which can handle ~40K before hitting its own cliff. (Exp 007)
-
A fixed redaction prompt reliably produces GDPR-clean output. 20 synthetic toxic real estate notes spanning 8 pre-registered GDPR categories — 0 true-positive leaks in any output. The local 26B model with
temperature=0.1and a structured system prompt passes all four pre-registered criteria: zero leaks, full structural compliance (TAGS + DESCRIPTION), all 20 within 300s. (Exp 006) -
The Flash Attention cliff is a runtime artefact, not a hardware limit — confirmed by independent runtime. MLX (Apple's native ML framework) shows no prefill cliff through 40K tokens on the same Mac Mini M4 Pro hardware. MLX prefill at 15K is 1.650 ms/tok — matching Ollama FA=0/q8_0 (1.694) within 3%. The cliff Exp 007 attributed to the Mac Mini's architecture was entirely a product of Ollama's llama.cpp Flash Attention tiling on unified memory. Two independent runtimes; same hardware; same result. The hardware ceiling is memory bandwidth, not attention kernel. (Exp 011)
-
Local models match compliance gaps; frontier models catch implementation gaps. A head-to-head adversarial critic comparison (three fixed personas, fixed JSON schema, same context bundle) found that gemma4:26b matched Claude Sonnet 4.6 on the DPO/compliance layer (DPA template, DSAR procedure, DPIA — 3/3 near-exact matches) but missed the highest-severity engineering finding: a primary moat component described across the BRIEF, deck, and BUILD_LOG had no corresponding code in any commit. gemma4 pattern-matched on documented claims and critiqued their replicability; Claude cross-referenced the BUILD_LOG claim against the git history and flagged the absence. Overlap rate: ~50%. False-positive rate: ~50%. Verdict: FAIL as a drop-in replacement, viable as a zero-cost compliance-layer complement to periodic frontier review. (Exp 009)
-
The cost-capability curve has one step — confirmed at the 4B-active/frontier boundary; dense local models untested. ⚠ Scope correction (Exp 012-Alpha, 2026-06-09): gemma4:26b is MoE A4B (~4B active parameters per forward pass, not 25.8B). Exp 012 tested one local model class; whether the step holds for dense 12–32B local models is open and pre-registered as Exp 015. A four-model sweep (gemma4:26b, Haiku, Sonnet, Opus) on a pre-scored 8-point rubric across two task types (DPO compliance extraction + engineering gap detection) produced: gemma4:26b 0/8, all three frontier models 5/8 net. Haiku→Sonnet (3.1× cost) and Haiku→Opus (6.4× cost) each yield zero additional rubric points. For structured analytical extraction on bounded context (~10K tokens), Haiku is cost-dominant. Qualitative differences exist within the same net score: Haiku and Sonnet have complementary blind spots (Haiku misses concurrency risk; Sonnet misses auth gap). Opus has the highest gross score (6/8) but the highest false positive rate. Two items — a missing DPIA for VLM processing and the Art. 22 model-version accountability gap — evaded all four models. (Exp 012; scope corrected 2026-06-09 — see Exp 012-Alpha in scientific_log.md)
Published at localfirstai.eu:
Technical — benchmarks, experiments, architecture
- We Found the Credentials. We Didn't Rotate Them. — Exp 021: a zero-escalation red-team pass found live credentials readable with one
cat. We chose not to rotate them, wrote down exactly why, and published the risk register — including the part where the audit tooling leaked the secrets itself. - The Model Wasn't Broken — Exp 023/025: a model that looked dead was 400× slower because
OLLAMA_CONTEXT_LENGTHwas never set and Ollama auto-selected 256k context for every call. One config line fixed it, zero code changes; gemma4:26b still wins on wall-clock against the newest Qwen. - We Tested What We Didn't Notice — Exp 018: daemon-down and weights-removed both degraded gracefully with zero data loss. The network-cut test severed the only channel to the machine and forced a reboot — sovereignty of inference is not sovereignty of observability.
- We Red-Teamed Our Own Bot — Exp 022: prompt injection, session extraction, and guardrail bypass all refuted. Corpus poisoning via the
/witnesscommunity channel confirmed — one admin approval turns a fabricated claim into fact for every subsequent user. - Hardening the Inference Node — Exp 020: auditing the machine behind the local-first pitch. A static permission listing's worst-looking finding wasn't real; the one that was real (
tail, NOPASSWD) wasn't the one that looked scary on paper. - We Reviewed Our Own Legal Brief with an Adversarial AI Panel. Zero of Seven Claims Survived Unchanged. — Exp 019: Claude adversarial panel + GPT-4o citation check + Gemini barrister review. 0/7 claims survived. BDAA omission, ingestion transfer gap, Gibraltar-EEA adequacy gap.
- We Didn't Notice — The US government suspended the world's best AI model overnight for all foreign nationals. CasaSol was unaffected. Exp 018 pre-registered.
- The Cost-Capability Curve Has One Step — A four-model sweep at the 4B-active/frontier boundary. One step, not a ramp. Haiku is cost-dominant; Haiku→Opus is 6.4× cost, 0 score gain.
- Same Hardware. Different Runtime. Same Result. — Exp 011: MLX and Ollama FA=0 on the same Mac Mini M4 Pro. Neither cliffs through 40K tokens. Prefill within 3%. The FA cliff was an Ollama/llama.cpp artefact, confirmed by an independent runtime.
- The Cliff That Wasn't — The 20K prefill cliff that shaped six months of cascade architecture was
OLLAMA_FLASH_ATTENTION=1. Removing it tripled the Mac Mini's operational ceiling to >40K tokens. Full 2×2 factorial: FA=1 is the sole culprit, q8_0 alone is benign. - The Adversarial Watcher: When a Local Model Audits Its Own Project — A staged 5-step pipeline that catches documentation drift before every merge. First production run: 5 confirmed gaps, 3 false positives, anatomy of each.
- We Tried to Replace Claude with a Local Critic. Here's Exactly Where It Failed. — Exp 009: head-to-head adversarial review. gemma4:26b matches the compliance layer; only the frontier model caught the impl-vs-docs gap.
- The Silicon Wager: M4 Pro vs M5 Max — Exp 007: every Chronos envelope was measured on one machine. A second arrived. The difference is not incremental.
- The GDPR Canary for Real Estate: 8 Data Categories, 0 Leaks — Exp 006: pre-registered fidelity sweep over 20 synthetic toxic notes. Zero true leaks. The claim becomes evidence.
- The Memory Bandwidth Cliff — Incident 003-Alpha: why local AI is bound by the bus, not the GPU.
- The Architecture of Anonymity — Exp 003: data sovereignty enforced by the import graph, not by policy.
- The Control Plane and the Data Plane — Exp 002: managing the AI thinking tax.
- The Genesis of Chronos — Why Nestor commits to verified, evidence-backed claims.
Essays — strategy, product, philosophy
- Why CasaSol.ai — If every company can be a Palantir now, how do you test that claim? An attempt to answer by building one — on the Costa del Sol.
- Should We Stop Asking Local LLMs to Think? — What Adam Smith, neuroscience, and a melting Mac Mini taught me about the real division of cognitive labour.
- The Sovereign Individual: Why Private Data is the Only Moat Left — As AI becomes commoditised, competitive advantage is private context.
- Every Company Can Be a Palantir Now — Proprietary structured data is the durable moat.
MIT