Skip to content

Latest commit

 

History

131 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Local First AI

Benchmark data, experiment harnesses, and scientific logs for running production AI on local hardware — no cloud required.

This repo accompanies the blog at localfirstai.eu and provides the verifiable evidence backing every claim made there.


Hardware

Component Value
Primary machine Mac Mini M4 Pro (miktam02), 64 GB unified memory
Secondary machine MacBook Pro M5 Max (miktam-mbp) — smart inference tier (Exp 016)
Primary model gemma4:26b (MoE A4B, 25.8B total params / ~4B active per forward pass, Q4_K_M)
Router model gemma4:e4b (fast routing layer)
Runtime Ollama 0.20.2
Orchestration OpenClaw → Nestor (local AI agent)
Operating ceiling (Mini) > 40,000 tokens on-wire (Exp 008/010 — no cliff at FA=0; Exp 007's 18K was a FA=1 artefact)
Operating ceiling (MBP) > 40,000 tokens on-wire (Exp 007 Phase B — FA=1 cliff at ~45K; FA=0 baseline not yet measured)

Project Chronos

Every claim on the blog is backed by a pre-registered experiment logged in tasks/chronos/scientific_log.md. The methodology: observation → hypothesis → experiment → evidence → conclusion. No retrofitted results.

Roadmap and pending experiments: tasks/chronos/roadmap.md

Experiments

# Name Status Key finding
001 Verification of Veracity Complete Chronos framework activated
002 Control Plane vs Data Plane Complete Thinking mode is flat at ~38 t/s until it isn't — unconstrained prompts trigger runaway
003 Anonymized Adversarial Memory Complete 0/20 source recognition, 0/3 identity leaks on Fight Club corpus — data-sovereignty moat is architectural
004 Bootstrap Diet Complete OpenClaw session hygiene
005 Router / Reducer Cascade Phase 0 closed Working two-model cascade over 8-year Apple Watch corpus; three load-bearing behaviours demonstrated
006 Redactor Fidelity (GDPR) Complete 0/20 × 8 categories — zero true-positive leaks across all pre-registered GDPR categories
007 Mac Mini vs MacBook Pro M5 Max Phase A+B complete Mini cliff ~18K, MBP cliff ~45K (2.5×); MBP gen t/s +200–370%; H1+H2 confirmed ⚠ see Exp 008
008 Flash Attention + q8_0 KV Cache Complete — landmark FA+q8_0 flags cause the cliff; FA=0/fp16 has no cliff through 40K. Exp 007 ceiling was an artefact
009 Adversarial Project Critic (Local vs. Frontier) Complete — FAIL gemma4:26b matched compliance layer (DPA/DSAR/DPIA) but missed impl-vs-docs gaps; 50% overlap, 50% FP rate
010 FA vs q8_0 Factorial Isolation Complete FA=1 is the sole culprit (cliff at 32.5K alone, 20K combined with q8_0). q8_0 alone: no cliff, +5% gen t/s
011 MLX Runtime vs Ollama — Context Cliff Complete No cliff through 40K on MLX. Prefill matches Ollama FA=0 within 3%. Cliff is an Ollama FA artefact, not a hardware limit.
012 Cost vs Capability: Where the Curve Breaks Complete ⚠ see Exp 012-Alpha gemma4:26b (MoE A4B, ~4B active) 0/8, Haiku/Sonnet/Opus all 5/8 net. Cliff confirmed at the 4B-active/frontier boundary; dense local 12–32B class untested (→ Exp 015). Haiku is cost-dominant; Haiku→Opus is 6.4× cost, 0 score gain.
013 Local Audit Loop: Can Scaffolding Move gemma4:26b Off Zero? Complete H confirmed (partial): decomposition recovered 2/3 findable items (Items 3+4 stable, Item 1 systematic bridge failure). Raw 2/5; context expansion deferred as out of scope for H3. Pre-filter architecture: local audit + Haiku top-up = ~$0.02/audit.
014 Capability Variance Floor Complete H1 CONFIRMED: gemma ≤0/8 all 5 reps (goes negative via systematic DSR FP). H3 CONFIRMED: no overlap (gemma max 0, Haiku min 1). H2 FALSIFIED: Haiku scores 1/1/4/1/2 — not stable at 4–6/8; the 5/8 Exp 012 result was a high-tail sample. Haiku mean ~2/8 across reps; variance driven by DSR/DPA FP penalties and B3 blind spot.
015 Active-Parameter Ablation: Dense vs MoE Pre-registered Does a dense 12–32B local model outperform gemma4:26b (MoE A4B) on the Exp 012 audit rubric? Bottleneck hypothesis: active compute per token, not total params.
016 Two-Mac Orchestration Phase B in progress Phase A: Qwen3-Coder-Next-4bit selected (4.0/5, 100.6 tok/s). Phase B: LAN path mini→MBP confirmed (miktam-mbp.local:8080). TTFT benchmark pending.
017 Argos Phase 0 — Feed Reconnaissance Pre-registered DGT NAP (DATEX II) + OpenChargeMap recon, Málaga–Gibraltar EV corridor. 24h cadence measurement required before Phase 1.
018 Sovereignty Resilience Complete Daemon-down and weights-removed both confirmed: graceful fallback, zero data loss, 1-command recovery. Network-cut was inconclusive by design — the only access to the test machine is remote, so blocking all outbound also blinded the observer, forcing a full reboot. Sovereignty of inference ≠ sovereignty of observability.
019 Adversarial Legal Review Pipeline Complete 0/7 claims survived unchanged across 3 review rounds (Claude adversarial panel + GPT-4o citation check + Gemini barrister review). Two critical issues caught: UK-US BDAA omission; ingestion transfer gap. First application of the Legal Agent pattern.
020 Hardening and Red-Teaming the Inference Node Complete Unauthenticated Ollama on the Tailscale network fixed (loopback-only + token-authenticated proxy); an unrestricted tail NOPASSWD grant confirmed as a genuine root-read primitive; the static listing's worst-looking finding (cp) turned out to be a misread, refuted by direct invocation.
021 Independent Red-Team Pass on the Inference Node H3 confirmed Zero-escalation reachability: without any sudo or exploit, an unpassphrased SSH key and three live credentials are readable by anything running as the automation account — one cat of a shell startup file. A lower bar than exp_020's sudo findings. Rotation declined as a documented accepted risk by the owner; the evidence records names + permissions only, never values.
022 Adversarial Red-Team of the CasaSol Guide Bot Complete H1/H2/H4 (prompt injection, session extraction, guardrail bypass) refuted — the bot resisted every attack, mostly rejected at the intent-router layer. H3 (corpus poisoning via /witness) confirmed: one admin approval is enough to get a fabricated claim indexed and reproduced as fact across every follow-up query.
023 Generation Efficiency Across the Local Model Family Complete Corrected re-run (fixed OLLAMA_CONTEXT_LENGTH, current roster): gemma4:26b posts the best wall-clock of all four models tested — 31.6s, beating even the smaller gemma4:e4b (32.9s) via fewer tokens per answer — while running 4.6x faster than the now-fixed gemma4:31b (144.7s) and 6.7x faster than the brand-new qwen3.8:27b (211.4s). Set against qwen3.8:27b finding more audit-rubric gaps (Exp 015), this is a real speed-vs-thoroughness tradeoff, not a clean win for either model.
024 Vision Capability: gemma4:26b vs qwen3.8:27b Inconclusive gemma4:26b parsed valid JSON on only 1/5 photos on Pharos's exact production vision prompt, with one response degenerating into a token-budget-filling repetition loop — a real failure mode worth flagging to Pharos directly. qwen3.8:27b's run returned 404 on every call because the model had silently vanished from disk mid-experiment — daemon never restarted, disk never filled; cause unexplained.
025 Context-Allocation Isolation H1+H2 confirmed, H3 superseded A clean 2-rep sweep confirms it's a real cliff, not a decline: gemma4:31b holds ~13.3 tok/s from 2048-16384 context, then collapses to ~0.1 tok/s at 262144; gemma4:26b stays flat ~60-64 tok/s across the same full range. Root cause: OLLAMA_CONTEXT_LENGTH was never set, so Ollama auto-selected 256k given this Mac's unified memory — a daemon-wide default hitting every call, not a per-model quirk. Fixed by setting it to 8192 in the launchd plist; verified gemma4:31b now runs at ~13-14 tok/s with zero code changes anywhere.
026 Contextual Retrieval on the COAPI Corpus, Fully Local Complete (pilot) Anthropic's "Contextual Retrieval" technique, tested fully locally on CasaSol's real naive-chunked COAPI corpus. Exact-hit retrieval rate climbs monotonically: unprefixed 0.455 → gemma4:e4b-prefixed 0.545 → gemma4:26b-prefixed 0.636 (n=11, directional). The two models cost almost the same time per chunk here (prompt-eval-bound, not generation-bound), so there's no speed reason to pick the cheaper model over the better one. Extrapolated full-corpus cost: ~4.8h, an easy overnight job.
035 COAPI Voice — Form in the Weights, Facts in Retrieval, On a Phone Complete — thesis refuted Three LoRA runs (Qwen3-1.7B and 4B, adapter strength varied 25×) installed the professional's form and lowered the base's use of retrieved facts every time (with-substance correctness 1.50 → 1.18/1.21, native ES 9/12 → 3/12). Quantisation, epochs, truncation, decoding and adapter strength each ruled out by pre-registered diagnostics. The reference build has no trained weights: untrained Qwen3-4B behind BM25 over our study notes and deterministic post-processing meets every format floor at the base's substance (1.50). Gemma 3 4B probed as a swap: level on substance, better PL/ES fluency, at the iPhone's memory cap — declined.

Watcher Runs

Production runs of the Adversarial Watcher — a staged local LLM pipeline that compares documented intent against shipped artefacts and produces annotated gap reports. Each run is an auditable Chronos artefact.

Run Project Confirmed gaps False positives Evidence
watcher_run_001 CasaSol 5 3 2026-06-06, gemma4:26b, ~270s

Incidents

# Name Finding
003-Alpha Memory Bandwidth Cliff Prefill on gemma4-think:26b goes super-quadratic past ~25K tokens on Apple Silicon. Hard operational ceiling: < 22K tokens on-wire. The bottleneck is memory bandwidth, not VRAM.

Benchmarks

Early benchmarks that preceded the Chronos framework. Scripts and results in tasks/chronos/experiments/.

Script What it measures
nestor-bench-phase1.sh Context window (4K–130K) vs generation speed. Finding: gen_tps flat at ~41 t/s.
nestor-bench-phase1b.sh Thinking mode token overhead. Finding: 5–15× token cost for zero quality gain on simple tasks.
nestor-bench-phase2-compare.sh Compressed-memory retrieval vs raw context.
nestor-bench-phase2-memory.sh Memory layer latency at scale.
nestor-bench-phase2b-retrieval.sh Retrieval accuracy across compression levels.
ollama pull gemma4:26b
chmod +x tasks/chronos/experiments/nestor-bench-phase1.sh
./tasks/chronos/experiments/nestor-bench-phase1.sh
# Results written to tasks/chronos/experiments/results/

Key findings (cumulative)

  1. The prefill cliff was an artefact of OLLAMA_FLASH_ATTENTION=1, not a hardware limit. Exp 007 measured a cliff at ~18K tokens on Mac Mini M4 Pro — but those runs were made with FA=1+q8_0 enabled (Incident 007-Alpha). Exp 008 established the FA=0 baseline: no cliff through 40K tokens. Exp 010's 2×2 factorial confirmed FA=1 is the sole culprit: it alone (without q8_0) produces a cliff at 32.5K and triples prefill latency at 15K tokens (1.774 → 5.405 ms/tok). Under optimal config (FA=0, q8_0), the Mac Mini's true operational ceiling is > 40K tokens on-wire. Flash Attention was designed for discrete GPU SRAM/HBM hierarchies; on Apple Silicon unified memory, its tiling overhead applies without the bandwidth benefit. (Incident 003-Alpha, Exp 007, Incident 007-Alpha, Exp 008, Exp 010)

  2. Thinking tokens are expensive — cap and name them explicitly. gemma4:26b is a thinking model: left unconstrained, a simple task generates 10,000–25,000 hidden thinking tokens — at 38 t/s that's 4–11 minutes per response with zero quality gain. The architectural response: a gemma4-think:26b Ollama alias with a hard 128K context cap, used only for tasks that genuinely need deliberation. The name makes the choice visible; the cap prevents runaway. (Exp 002)

  3. Data sovereignty is an architectural property, not a policy. An anonymization boundary enforced by the import graph — not by a prompt or a config flag — defeated source recognition (0/20) and identity bridging (0/3) on a corpus the model has memorised. The moat is the architecture. (Exp 003)

  4. A two-model cascade extends the operating envelope. Router (gemma4:e4b) routes in ~3–4s. Reducer (gemma4:26b) synthesises only what fits below the 22K cliff. The cascade made an 8-year health corpus queryable on local hardware without hitting the bandwidth cliff on normal queries. (Exp 005)

  5. The 22K ceiling is a property of the hardware, not a bug. Memory bandwidth saturates during prefill on the M4 Pro's unified memory architecture. Mitigations: cliff-aware coarsening in the extractor, hard token budgets in the cascade, streaming watchdog for booth/production use.

  6. The M5 Max die is in a different performance class for inference. At 25K tokens, MBP gen t/s is 66 vs Mini's 14 — a 4.7× difference on the same model weights and quantisation. MBP at 35K tokens (1.24 ms/tok prefill) is still well below the Mini's baseline at 4K tokens (3.03 ms/tok). The cascade's 22K bundle ceiling — set for the Mini — is comfortably safe on the MBP, which can handle ~40K before hitting its own cliff. (Exp 007)

  7. A fixed redaction prompt reliably produces GDPR-clean output. 20 synthetic toxic real estate notes spanning 8 pre-registered GDPR categories — 0 true-positive leaks in any output. The local 26B model with temperature=0.1 and a structured system prompt passes all four pre-registered criteria: zero leaks, full structural compliance (TAGS + DESCRIPTION), all 20 within 300s. (Exp 006)

  8. The Flash Attention cliff is a runtime artefact, not a hardware limit — confirmed by independent runtime. MLX (Apple's native ML framework) shows no prefill cliff through 40K tokens on the same Mac Mini M4 Pro hardware. MLX prefill at 15K is 1.650 ms/tok — matching Ollama FA=0/q8_0 (1.694) within 3%. The cliff Exp 007 attributed to the Mac Mini's architecture was entirely a product of Ollama's llama.cpp Flash Attention tiling on unified memory. Two independent runtimes; same hardware; same result. The hardware ceiling is memory bandwidth, not attention kernel. (Exp 011)

  9. Local models match compliance gaps; frontier models catch implementation gaps. A head-to-head adversarial critic comparison (three fixed personas, fixed JSON schema, same context bundle) found that gemma4:26b matched Claude Sonnet 4.6 on the DPO/compliance layer (DPA template, DSAR procedure, DPIA — 3/3 near-exact matches) but missed the highest-severity engineering finding: a primary moat component described across the BRIEF, deck, and BUILD_LOG had no corresponding code in any commit. gemma4 pattern-matched on documented claims and critiqued their replicability; Claude cross-referenced the BUILD_LOG claim against the git history and flagged the absence. Overlap rate: ~50%. False-positive rate: ~50%. Verdict: FAIL as a drop-in replacement, viable as a zero-cost compliance-layer complement to periodic frontier review. (Exp 009)

  10. The cost-capability curve has one step — confirmed at the 4B-active/frontier boundary; dense local models untested. ⚠ Scope correction (Exp 012-Alpha, 2026-06-09): gemma4:26b is MoE A4B (~4B active parameters per forward pass, not 25.8B). Exp 012 tested one local model class; whether the step holds for dense 12–32B local models is open and pre-registered as Exp 015. A four-model sweep (gemma4:26b, Haiku, Sonnet, Opus) on a pre-scored 8-point rubric across two task types (DPO compliance extraction + engineering gap detection) produced: gemma4:26b 0/8, all three frontier models 5/8 net. Haiku→Sonnet (3.1× cost) and Haiku→Opus (6.4× cost) each yield zero additional rubric points. For structured analytical extraction on bounded context (~10K tokens), Haiku is cost-dominant. Qualitative differences exist within the same net score: Haiku and Sonnet have complementary blind spots (Haiku misses concurrency risk; Sonnet misses auth gap). Opus has the highest gross score (6/8) but the highest false positive rate. Two items — a missing DPIA for VLM processing and the Art. 22 model-version accountability gap — evaded all four models. (Exp 012; scope corrected 2026-06-09 — see Exp 012-Alpha in scientific_log.md)


Blog posts

Published at localfirstai.eu:

Technical — benchmarks, experiments, architecture

Essays — strategy, product, philosophy


License

MIT

About

Local First AI; Benchmark data, scripts, and configuration for running production AI on local hardware.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages