Skip to content

nagual-qe 0.2.0: build fix, asymmetric reward rule, trained router, working serve, semantic search - #40

Merged
proffesor-for-testing merged 9 commits into
masterfrom
fix/build-sha3-sqlx
Sep 28, 2026
Merged

proffesor-for-testing merged 9 commits into
masterfrom
fix/build-sha3-sqlx

Conversation

@proffesor-for-testing

@proffesor-for-testing proffesor-for-testing commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Why

master has not compiled since the dependabot majors (#26 sha3, #29 sqlx, axum 0.8), and CI only ran --lib. Running the HUSTEF masterclass devcontainer end to end against this branch surfaced more: nagual serve crashed on startup, the dashboard broke on fresh databases, the learning rule differed between CLI and HTTP, and the router tests exercised a mock.

What

Build / CI

  • shake crate for SHAKE-256 (sha3 0.12), sqlx::AssertSqlSafe on four self-owned DDL sites, rand_chacha 0.3.
  • CI runs the whole suite (unit, integration, doc) with kos serve. Dependabot ignores sqlx/sha3/axum majors.

Learning (reward rule, per Profa)

  • One rule for CLI, HTTP and MCP (learning::reward_step). Values: success +0.10, partial +0.05, failure −0.15, security failure −0.30, clamped to [0, 1]. Previously the CLI used an EMA and HTTP its own copy.
  • New failure class security. MCP accepts failure_mode and returns the pattern's real new reward. HTTP accepts partial/neutral and returns 400 on unknown outcomes.

Router

  • tests/router_tests.rs tested a mock defined inside the file. It is rewritten against the production router (49 tests plus property tests).
  • The pretrained FastGRNN scored every query 0.497–0.518: it was trained on random synthetic data with incomplete gradients.
    • Two uninformative inputs are replaced with text features.
    • It is retrained with correct backprop on models/router_queries.jsonl: 120 labelled training queries plus 40 held out.
    • Held-out results: level accuracy 47.5% → 65%, worst miss 2 → 1 level. Tests enforce a bar fixed before training.
  • NaN/inf embeddings are rejected; they used to route to the most expensive tier.

Serve / dashboard

  • The axum 0.8 route syntax ({id}) fixes the startup panic. build_router() is now under test.
  • Local-only mode actually engages (no token, no users, no keys). Before this, a fresh install got 401 from its own dashboard.
  • Dashboard queries detect created_at vs timestamp; they returned 500 on CLI-created databases.
  • PII redaction on the /api read path (port from nagual-rs 1.1.0), and /api/graph/3d?limit=.

CLI

  • knowledge list --domain X --limit N pagination.
  • Logs go to stderr.
  • learn record prints 0.50 -> 0.35 (failure -0.15).
  • Semantic search (--semantic, alias of --hyperbolic) scores all patterns, not just the newest limit×10.
  • The ONNX model is found in $NAGUAL_MODEL_DIR, ./models or ~/.nagual/models.
  • No ONNX warning in hash-only builds.

Tests fixed: a flaky test_high_dimensional_embeddings (failed about 38% of runs) and a non-compiling ml::lora doctest.

Test results

  • cargo test --no-default-features --features "kos serve" on macOS arm64: 3004 passed, 0 failed.
  • Default (onnx-embed) build compiles.
  • Every regression test was checked to fail against the pre-fix code.
  • End to end: the HUSTEF masterclass devcontainer (make smoke) is green against this branch.

API changes

  • ComplexityFeatures::{embedding_norm, pattern_coverage} → {reasoning_demand, structure}.
  • EstimatorConfig::{norm_weight, coverage_weight} → {reasoning_weight, structure_weight}.
  • learn embed --model-path/--tokenizer-path are now optional.
  • Noted, not changed: VendorConfig::cloud_threshold is unused by select().

Generated with Nagual

… pagination, logs to stderr

Build (main did not compile since #26/#29):
- sha3 0.12 no longer ships the SHAKE XOFs; use the RustCrypto `shake` crate for
  Shake256 in ml/hash_embedder.rs and learning/domain_expansion.rs (same digest 0.11 traits).
- sqlx 0.9 requires dynamic SQL to be explicitly asserted safe: wrap the four
  non-literal query sites (db::execute, graph::maintenance prune query, migration up/down
  scripts) in sqlx::AssertSqlSafe. These are all self-owned DDL/migration strings.
- rand_chacha pinned back to 0.3 to match rand 0.8 (0.9/0.10 broke tests/holdout_test.rs;
  CI only ran --lib so it went unnoticed).

Behaviour:
- `nagual knowledge list --domain X --limit N` applied the limit before the domain filter
  and the reward/usage/created sorts, so small limits returned nothing. Fetch the full set
  when a filter or non-recency sort is requested, then filter, sort, paginate.
- Router latency was truncated to whole microseconds; sub-µs decisions recorded 0 and
  avg_latency_us() reported 0.0 (test_vendor_router_metrics failed on release builds).
  Ceil to 1 µs.
- Tracing output goes to stderr (was stdout, polluting every command's output), and
  RUST_LOG now fully controls the filter instead of being overridden by nagual=info.

Process:
- CI runs the whole suite with kos+serve, not just --lib.
- Dependabot ignores majors that need hand-migration (sqlx, sha3, axum) and keeps
  rand_chacha on the rand 0.8 line.

Known, pre-existing, NOT fixed here (fail on 54f6932 too): tests/router_tests.rs
test_text_complexity_estimation and test_vendor_selection_high_complexity — router
behaviour, needs a decision on the estimator thresholds.

Co-Authored-By: claude-flow & Agentic QE fleet
…(port from nagual 1.1.0)

Ported from the private nagual-rs tree (change dated 2026-07-01, "redblue data-exfiltration
HIGH finding"): the local SQLite store is never redacted at rest, but /api search, get and
list responses feed retrieval into agent context, so a stored secret exfiltrated on an
ordinary search. PatternResponse and the list handler now run problem/solution/context
through sync::pii::global_redactor() on the way out; clean text is unchanged.

Also: /api/graph/3d takes ?limit= (default 5000, ceiling 20000) instead of a hard 2000.

Test fixtures use example.com / alice paths — no personal strings in the public edition.

Co-Authored-By: claude-flow & Agentic QE fleet
…lease 0.2.0

- Split `run_list` into `list_from_storage` (fetch) + `filter_sort_paginate` + `print_list`
  so the pagination fix from the previous commit is testable. Three regression tests on a
  SQLite-only store (never init_storage, which may resolve a real Postgres from config):
  domain filter + --limit 2 + reward sort returns the 2 best matches; asc reward sort sees
  the oldest row; plain "recent N" keeps the fast path. Verified they fail on the pre-fix
  fetch (`fetch_n = limit + offset` -> `[]`).
- tests/profdag_search_tests.rs: test_high_dimensional_embeddings used 10 independent random
  1536-d vectors against min_similarity 0.0 (~50% each land below 0), so it failed ~38% of
  runs. Use similar embeddings around the query, like test_topk_returns_correct_count.
- ml::lora module doctest referenced an undefined `pairs` and failed to compile — would have
  turned the new full-suite CI job red.
- `knowledge list` with a domain that matches nothing now says so instead of claiming the
  database is empty.
- Version 0.2.0, CHANGELOG entry, README documents the no-ONNX `kos serve` build path and the
  HTTP read-path redaction.

Full suite (kos+serve) on macOS arm64: 2989 passed, 2 failed — the two in-file router
tests in tests/router_tests.rs, pending a decision (see CHANGELOG "Known issues").

Co-Authored-By: claude-flow & Agentic QE fleet
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-09-28T07:24:22.570934Z 68f4887 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

…rn record output

Found by running the HUSTEF masterclass devcontainer end to end against this branch.

- serve: `nagual serve` panicked on startup ("Path segments must not start with `:`") — axum 0.8
  (dependabot) needs `{id}` path params. Router extracted into `build_router()`; tests build it and
  check that parameterised routes reach their handlers. Verified the tests fail with a `:id` route.
- auth: local-only mode required `key_store.is_none()`, but serve always opens a key store, so a
  fresh local install 401'd its own dashboard. Now: no master token + no dashboard users + zero
  active keys => local-only, evaluated per request, fail-closed. New Send-safe
  ApiKeyStore::has_active_keys(). Three tests (empty store, first key closes it, users disable it).
- learn record: show the pattern's reward before -> after instead of only the outcome's target
  reward (0.20 for failure), which was being read as the new pattern reward.
- main: skip the ORT_DYLIB_PATH check/warning in builds without `onnx-embed`.

Suite (kos+serve, macOS arm64): 2994 passed, 2 failed (in-file router tests, unchanged).

Co-Authored-By: claude-flow & Agentic QE fleet
The dashboard handlers hard-coded `created_at`, but databases created by the CLI's
PatternStorage schema (every fresh local install) have `timestamp` instead, so /api/patterns
and /api/graph/3d returned 500 ("no such column: created_at"), /api/pulse and several
action endpoints failed, and /api/status reported no oldest/newest pattern. Found in the
HUSTEF masterclass devcontainer.

`created_column()` detects the column (like the existing tier/domain detection) and is used in
handlers.rs and the four reasoning_patterns queries in action_handlers.rs. New test builds a DB
through init_storage_sqlite_only + store_pattern and exercises patterns, graph/3d, pulse and
status; verified it fails with the old hard-coded column.

Suite (kos+serve, macOS arm64): 2995 passed, 2 failed (in-file router tests, unchanged).

Co-Authored-By: claude-flow & Agentic QE fleet
…uter tests hit real code

Reward (requested by Profa):
- `learning::reward_step`: success +0.10, partial +0.05, neutral 0, failure -0.15, security
  failure -0.30, clamped to [0,1]. Used by the CLI learner, the HTTP outcome endpoint and MCP.
  Replaces the CLI's EMA toward 0.9/0.2 (failure cost ~= success gain) and the HTTP handler's
  private copy. Effectiveness keeps the EMA.
- FailureMode::SecurityIssue ("security", "security_issue", "vulnerability", "leak").
- SonaLearner::record_outcome_classified carries the failure mode (stored on the pattern and
  used for the penalty); `learn record` prints `0.50 -> 0.35 (failure -0.15)`.
- HTTP outcome: partial/neutral accepted, unknown outcome -> 400, Bayesian update added.
- MCP nagual_record_outcome: optional failure_mode; returns the real new reward/effectiveness.
- Tests: pure rule + clamping, learner on a real SQLite store (0.50 -> 0.35 -> 0.45; security
  0.50 -> 0.20, mode persisted), HTTP handler (security, partial, 400 on typo).

Router:
- tests/router_tests.rs imported nothing from nagual — it tested a mock defined in the file.
  Rewritten against VendorRouter/VendorSelector/ComplexityEstimator/FastGRNN: inference,
  levels, features, thresholds, profiles, fallback + health, latency budget, metrics,
  confidence, properties (range, determinism, cost monotonicity, fallback availability),
  edge cases. 49 tests.
- Reject NaN/inf embeddings (complexity NaN fell through to the Claude tier). Verified the
  test fails without the fix.
- Known issue documented, not asserted: the pretrained FastGRNN scores ~0.47–0.53 for every
  query; needs a decision (retrain vs. simplify).

Suite (kos+serve, macOS arm64): 2998 passed, 0 failed.

Co-Authored-By: claude-flow & Agentic QE fleet
…ality gate on held-out set

The pretrained FastGRNN scored every query 0.497–0.518: it was trained on random synthetic
features with incomplete gradients (no gate updates, no activation derivatives), and two of
its five inputs were uninformative for normalised embeddings.

- Features (all text-derived, [0,1]): log-scaled length, reasoning_demand (design / proof /
  trade-off / diagnosis / planning cues), domain_specificity, structure (code, several
  questions, lists, multi-part asks), historical_accuracy. The embedding is still validated
  (non-empty, finite) but no longer scored.
  API: ComplexityFeatures {embedding_norm, pattern_coverage} -> {reasoning_demand, structure};
  EstimatorConfig {norm_weight, coverage_weight} -> {reasoning_weight, structure_weight}.
- models/router_queries.jsonl: 160 labelled queries, 40 per level, fixed train/test split,
  rubric in models/README.md.
- examples/router_features.rs dumps features with the production estimator;
  models/train_fastgrnn.py rewritten: exact Rust forward pass, full backprop, Adam, seeded
  restarts selected on train loss only; writes weights + metrics + data hash.
- Held-out (40 queries): level 65.0%, local-vs-cloud 90.0%, worst error 1 level. Old weights
  on the same queries: 47.5% / 90% / 2 levels. hello: Claude -> LocalSmall.
- Tests: held-out bar (fixed before training), per-level score spread, anchors, and Rust
  inference == trainer's recorded accuracy. Verified all four fail with the old weights.
  test_pretrained_uses_trained_weights now compares against the embedded JSON.

Suite (kos+serve, macOS arm64): 3002 passed, 0 failed.

Co-Authored-By: claude-flow & Agentic QE fleet
…l outside ./models

For the HUSTEF masterclass switch to ONNX embeddings (verified in the devcontainer: ONNX Runtime
1.24.1 + all-MiniLM-L6-v2 embeds the 520-pattern seed in ~41 s on 8 cores).

- knowledge search --hyperbolic fetched only the `limit * 10` most recent patterns as candidates,
  so older patterns were unfindable by meaning. It now scores all patterns. The substring fallback
  had the same cap; it now loads the full set lazily, only when FTS5 returns nothing.
- ml::resolve_model_paths(): $NAGUAL_MODEL_DIR, then ./models, then ~/.nagual/models. Used by
  learn embed (model/tokenizer paths now optional), knowledge search and patterns.
- --semantic is a visible alias for --hyperbolic; --fts-weight/--vector-weight are marked as
  reserved (they were never read).
- Tests for the resolver precedence.

Suite (kos+serve, macOS arm64): 3004 passed, 0 failed. Default (onnx-embed) build: cargo check OK.

Co-Authored-By: claude-flow & Agentic QE fleet
@proffesor-for-testing proffesor-for-testing changed the title nagual-qe 0.2.0: restore the build, fix list pagination, redact PII on the HTTP read path nagual-qe 0.2.0: build fix, asymmetric reward rule, trained router, working serve, semantic search Sep 28, 2026
Formatting only (rustfmt 1.9.0), so the `fmt + clippy` CI job passes. 260 files; the debt
predates this branch (benches, most of src/). No code changes: `cargo fmt --all -- --check`
clean, clippy (CI flags) exits 0, suite 3004 passed / 0 failed, default-feature build OK.

Co-Authored-By: claude-flow & Agentic QE fleet
@proffesor-for-testing
proffesor-for-testing merged commit 2ddb7fa into master Sep 28, 2026
4 checks passed
@proffesor-for-testing
proffesor-for-testing deleted the fix/build-sha3-sqlx branch September 28, 2026 11:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant