nagual-qe 0.2.0: build fix, asymmetric reward rule, trained router, working serve, semantic search - #40
Merged
Conversation
… pagination, logs to stderr Build (main did not compile since #26/#29): - sha3 0.12 no longer ships the SHAKE XOFs; use the RustCrypto `shake` crate for Shake256 in ml/hash_embedder.rs and learning/domain_expansion.rs (same digest 0.11 traits). - sqlx 0.9 requires dynamic SQL to be explicitly asserted safe: wrap the four non-literal query sites (db::execute, graph::maintenance prune query, migration up/down scripts) in sqlx::AssertSqlSafe. These are all self-owned DDL/migration strings. - rand_chacha pinned back to 0.3 to match rand 0.8 (0.9/0.10 broke tests/holdout_test.rs; CI only ran --lib so it went unnoticed). Behaviour: - `nagual knowledge list --domain X --limit N` applied the limit before the domain filter and the reward/usage/created sorts, so small limits returned nothing. Fetch the full set when a filter or non-recency sort is requested, then filter, sort, paginate. - Router latency was truncated to whole microseconds; sub-µs decisions recorded 0 and avg_latency_us() reported 0.0 (test_vendor_router_metrics failed on release builds). Ceil to 1 µs. - Tracing output goes to stderr (was stdout, polluting every command's output), and RUST_LOG now fully controls the filter instead of being overridden by nagual=info. Process: - CI runs the whole suite with kos+serve, not just --lib. - Dependabot ignores majors that need hand-migration (sqlx, sha3, axum) and keeps rand_chacha on the rand 0.8 line. Known, pre-existing, NOT fixed here (fail on 54f6932 too): tests/router_tests.rs test_text_complexity_estimation and test_vendor_selection_high_complexity — router behaviour, needs a decision on the estimator thresholds. Co-Authored-By: claude-flow & Agentic QE fleet
…(port from nagual 1.1.0) Ported from the private nagual-rs tree (change dated 2026-07-01, "redblue data-exfiltration HIGH finding"): the local SQLite store is never redacted at rest, but /api search, get and list responses feed retrieval into agent context, so a stored secret exfiltrated on an ordinary search. PatternResponse and the list handler now run problem/solution/context through sync::pii::global_redactor() on the way out; clean text is unchanged. Also: /api/graph/3d takes ?limit= (default 5000, ceiling 20000) instead of a hard 2000. Test fixtures use example.com / alice paths — no personal strings in the public edition. Co-Authored-By: claude-flow & Agentic QE fleet
…lease 0.2.0 - Split `run_list` into `list_from_storage` (fetch) + `filter_sort_paginate` + `print_list` so the pagination fix from the previous commit is testable. Three regression tests on a SQLite-only store (never init_storage, which may resolve a real Postgres from config): domain filter + --limit 2 + reward sort returns the 2 best matches; asc reward sort sees the oldest row; plain "recent N" keeps the fast path. Verified they fail on the pre-fix fetch (`fetch_n = limit + offset` -> `[]`). - tests/profdag_search_tests.rs: test_high_dimensional_embeddings used 10 independent random 1536-d vectors against min_similarity 0.0 (~50% each land below 0), so it failed ~38% of runs. Use similar embeddings around the query, like test_topk_returns_correct_count. - ml::lora module doctest referenced an undefined `pairs` and failed to compile — would have turned the new full-suite CI job red. - `knowledge list` with a domain that matches nothing now says so instead of claiming the database is empty. - Version 0.2.0, CHANGELOG entry, README documents the no-ONNX `kos serve` build path and the HTTP read-path redaction. Full suite (kos+serve) on macOS arm64: 2989 passed, 2 failed — the two in-file router tests in tests/router_tests.rs, pending a decision (see CHANGELOG "Known issues"). Co-Authored-By: claude-flow & Agentic QE fleet
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
…rn record output
Found by running the HUSTEF masterclass devcontainer end to end against this branch.
- serve: `nagual serve` panicked on startup ("Path segments must not start with `:`") — axum 0.8
(dependabot) needs `{id}` path params. Router extracted into `build_router()`; tests build it and
check that parameterised routes reach their handlers. Verified the tests fail with a `:id` route.
- auth: local-only mode required `key_store.is_none()`, but serve always opens a key store, so a
fresh local install 401'd its own dashboard. Now: no master token + no dashboard users + zero
active keys => local-only, evaluated per request, fail-closed. New Send-safe
ApiKeyStore::has_active_keys(). Three tests (empty store, first key closes it, users disable it).
- learn record: show the pattern's reward before -> after instead of only the outcome's target
reward (0.20 for failure), which was being read as the new pattern reward.
- main: skip the ORT_DYLIB_PATH check/warning in builds without `onnx-embed`.
Suite (kos+serve, macOS arm64): 2994 passed, 2 failed (in-file router tests, unchanged).
Co-Authored-By: claude-flow & Agentic QE fleet
The dashboard handlers hard-coded `created_at`, but databases created by the CLI's
PatternStorage schema (every fresh local install) have `timestamp` instead, so /api/patterns
and /api/graph/3d returned 500 ("no such column: created_at"), /api/pulse and several
action endpoints failed, and /api/status reported no oldest/newest pattern. Found in the
HUSTEF masterclass devcontainer.
`created_column()` detects the column (like the existing tier/domain detection) and is used in
handlers.rs and the four reasoning_patterns queries in action_handlers.rs. New test builds a DB
through init_storage_sqlite_only + store_pattern and exercises patterns, graph/3d, pulse and
status; verified it fails with the old hard-coded column.
Suite (kos+serve, macOS arm64): 2995 passed, 2 failed (in-file router tests, unchanged).
Co-Authored-By: claude-flow & Agentic QE fleet
…uter tests hit real code
Reward (requested by Profa):
- `learning::reward_step`: success +0.10, partial +0.05, neutral 0, failure -0.15, security
failure -0.30, clamped to [0,1]. Used by the CLI learner, the HTTP outcome endpoint and MCP.
Replaces the CLI's EMA toward 0.9/0.2 (failure cost ~= success gain) and the HTTP handler's
private copy. Effectiveness keeps the EMA.
- FailureMode::SecurityIssue ("security", "security_issue", "vulnerability", "leak").
- SonaLearner::record_outcome_classified carries the failure mode (stored on the pattern and
used for the penalty); `learn record` prints `0.50 -> 0.35 (failure -0.15)`.
- HTTP outcome: partial/neutral accepted, unknown outcome -> 400, Bayesian update added.
- MCP nagual_record_outcome: optional failure_mode; returns the real new reward/effectiveness.
- Tests: pure rule + clamping, learner on a real SQLite store (0.50 -> 0.35 -> 0.45; security
0.50 -> 0.20, mode persisted), HTTP handler (security, partial, 400 on typo).
Router:
- tests/router_tests.rs imported nothing from nagual — it tested a mock defined in the file.
Rewritten against VendorRouter/VendorSelector/ComplexityEstimator/FastGRNN: inference,
levels, features, thresholds, profiles, fallback + health, latency budget, metrics,
confidence, properties (range, determinism, cost monotonicity, fallback availability),
edge cases. 49 tests.
- Reject NaN/inf embeddings (complexity NaN fell through to the Claude tier). Verified the
test fails without the fix.
- Known issue documented, not asserted: the pretrained FastGRNN scores ~0.47–0.53 for every
query; needs a decision (retrain vs. simplify).
Suite (kos+serve, macOS arm64): 2998 passed, 0 failed.
Co-Authored-By: claude-flow & Agentic QE fleet
…ality gate on held-out set
The pretrained FastGRNN scored every query 0.497–0.518: it was trained on random synthetic
features with incomplete gradients (no gate updates, no activation derivatives), and two of
its five inputs were uninformative for normalised embeddings.
- Features (all text-derived, [0,1]): log-scaled length, reasoning_demand (design / proof /
trade-off / diagnosis / planning cues), domain_specificity, structure (code, several
questions, lists, multi-part asks), historical_accuracy. The embedding is still validated
(non-empty, finite) but no longer scored.
API: ComplexityFeatures {embedding_norm, pattern_coverage} -> {reasoning_demand, structure};
EstimatorConfig {norm_weight, coverage_weight} -> {reasoning_weight, structure_weight}.
- models/router_queries.jsonl: 160 labelled queries, 40 per level, fixed train/test split,
rubric in models/README.md.
- examples/router_features.rs dumps features with the production estimator;
models/train_fastgrnn.py rewritten: exact Rust forward pass, full backprop, Adam, seeded
restarts selected on train loss only; writes weights + metrics + data hash.
- Held-out (40 queries): level 65.0%, local-vs-cloud 90.0%, worst error 1 level. Old weights
on the same queries: 47.5% / 90% / 2 levels. hello: Claude -> LocalSmall.
- Tests: held-out bar (fixed before training), per-level score spread, anchors, and Rust
inference == trainer's recorded accuracy. Verified all four fail with the old weights.
test_pretrained_uses_trained_weights now compares against the embedded JSON.
Suite (kos+serve, macOS arm64): 3002 passed, 0 failed.
Co-Authored-By: claude-flow & Agentic QE fleet
…l outside ./models For the HUSTEF masterclass switch to ONNX embeddings (verified in the devcontainer: ONNX Runtime 1.24.1 + all-MiniLM-L6-v2 embeds the 520-pattern seed in ~41 s on 8 cores). - knowledge search --hyperbolic fetched only the `limit * 10` most recent patterns as candidates, so older patterns were unfindable by meaning. It now scores all patterns. The substring fallback had the same cap; it now loads the full set lazily, only when FTS5 returns nothing. - ml::resolve_model_paths(): $NAGUAL_MODEL_DIR, then ./models, then ~/.nagual/models. Used by learn embed (model/tokenizer paths now optional), knowledge search and patterns. - --semantic is a visible alias for --hyperbolic; --fts-weight/--vector-weight are marked as reserved (they were never read). - Tests for the resolver precedence. Suite (kos+serve, macOS arm64): 3004 passed, 0 failed. Default (onnx-embed) build: cargo check OK. Co-Authored-By: claude-flow & Agentic QE fleet
Formatting only (rustfmt 1.9.0), so the `fmt + clippy` CI job passes. 260 files; the debt predates this branch (benches, most of src/). No code changes: `cargo fmt --all -- --check` clean, clippy (CI flags) exits 0, suite 3004 passed / 0 failed, default-feature build OK. Co-Authored-By: claude-flow & Agentic QE fleet
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
masterhas not compiled since the dependabot majors (#26 sha3, #29 sqlx, axum 0.8), and CI only ran--lib. Running the HUSTEF masterclass devcontainer end to end against this branch surfaced more:nagual servecrashed on startup, the dashboard broke on fresh databases, the learning rule differed between CLI and HTTP, and the router tests exercised a mock.What
Build / CI
shakecrate for SHAKE-256 (sha3 0.12),sqlx::AssertSqlSafeon four self-owned DDL sites,rand_chacha0.3.kos serve. Dependabot ignores sqlx/sha3/axum majors.Learning (reward rule, per Profa)
learning::reward_step). Values: success +0.10, partial +0.05, failure −0.15, security failure −0.30, clamped to [0, 1]. Previously the CLI used an EMA and HTTP its own copy.security. MCP acceptsfailure_modeand returns the pattern's real new reward. HTTP acceptspartial/neutraland returns 400 on unknown outcomes.Router
tests/router_tests.rstested a mock defined inside the file. It is rewritten against the production router (49 tests plus property tests).models/router_queries.jsonl: 120 labelled training queries plus 40 held out.Serve / dashboard
{id}) fixes the startup panic.build_router()is now under test.created_atvstimestamp; they returned 500 on CLI-created databases./apiread path (port from nagual-rs 1.1.0), and/api/graph/3d?limit=.CLI
knowledge list --domain X --limit Npagination.learn recordprints0.50 -> 0.35 (failure -0.15).--semantic, alias of--hyperbolic) scores all patterns, not just the newestlimit×10.$NAGUAL_MODEL_DIR,./modelsor~/.nagual/models.Tests fixed: a flaky
test_high_dimensional_embeddings(failed about 38% of runs) and a non-compilingml::loradoctest.Test results
cargo test --no-default-features --features "kos serve"on macOS arm64: 3004 passed, 0 failed.onnx-embed) build compiles.make smoke) is green against this branch.API changes
ComplexityFeatures::{embedding_norm, pattern_coverage}→{reasoning_demand, structure}.EstimatorConfig::{norm_weight, coverage_weight}→{reasoning_weight, structure_weight}.learn embed --model-path/--tokenizer-pathare now optional.VendorConfig::cloud_thresholdis unused byselect().Generated with Nagual