Skip to content

feat: Jev harness on dspy 3.4.0 + ReAnchor, re-derived question registry (CONTRACT EDIT FOR OWNER REVIEW) - #57

Merged
urnlahzer merged 2 commits into
mainfrom
feat/jev-dspy340-reanchor
Sep 27, 2026
Merged

urnlahzer merged 2 commits into
mainfrom
feat/jev-dspy340-reanchor

Conversation

@urnlahzer

Copy link
Copy Markdown
Owner

Summary

Two commits. The first is tooling only; the second is a contract edit.

1. tools/jev-optimize on dspy 3.4.0 native decision types + ReAnchor (5446b7e)

  • Signatures use dspy.experimental.Noul / Choice; the hand-rolled adapter collapses to a JevLM that implements the system-one decision-request protocol and sends the exact production wire shape (flat state, question-id key, tuned docstring in the question's instructions) — checked byte-for-byte for all 19 questions in tests/test_wire_shape.py.
  • Noul questions: accept/escalate cuts come from two ReAnchor passes with mirrored asymmetric metrics (WRONG_DECISION_PENALTY = 9, i.e. ~10% error on automatic decisions). Choice questions keep the confidence sweep. GEPA wording search unchanged.
  • tests/test_offline_end_to_end.py runs the real GEPA + ReAnchor path against a fake Decisions endpoint (realistic answers, malformed reply, HTTP 400, 429 retry) and is the offline gate before any paid run.
  • 92 tests, ruff clean. No [typesafe] extra; OpenRouter transport unchanged.

2. Re-derived contracts/model/decision-questions.json (4ecfce7) — CONTRACT EDIT FOR OWNER REVIEW

  • Live optimize runs on rules, triage and closure (2026-09-27). Eight instruction texts change (4 rules, 4 triage), each with a test-accuracy gain (largest: thread_merge 0.76→0.91, asks_recipient 0.87→0.95, names_time 0.83→0.90). Closure wording unchanged.
  • Thresholds refitted on all 17 questions. Automatic decisions are ≥90% right on every question; weak questions escalate more to the chat model (closure.modified 71%, rules.event_match 59%, closure.fulfilled 29%; others 0–25%).
  • tools/check-model-boundary.ps1 pinned to the checker's own hash. Owner tested the rebuilt release binary.

Follow-ups (pre-existing, not in this PR)

  • The checker's registry hash is timezone-dependent: PowerShell re-emits tuned_at in the local offset. README documents the pinning recipe; the fix belongs in the checker.
  • Closure wording never improves: the proposer's primary-field rule (later.paragraph_text) rejects nearly every GEPA proposal (510 rejections vs ~70 for other sets).

Test plan

  • uv run pytest -q (92 passed) and uv run ruff check . in tools/jev-optimize
  • cargo test -p openloops-inference (98 passed) against the new registry
  • pwsh ./tools/check-model-boundary.ps1 passes with the new pin
  • check-public-repo.ps1 Staged (both commits) and History
  • Release binary rebuilt with the new registry and tested by the owner

🤖 Generated with Claude Code

https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas

urnlahzer and others added 2 commits September 27, 2026 15:10
…and ReAnchor

Replace the hand-rolled Jev adapter and threshold sweep with dspy's native
Noul/Choice output types and the ReAnchor calibrator. JevLM implements the
system-one decision-request protocol and translates DSPy's envelope back to
the production wire shape (flat state, question-id key, tuned docstring in the
question's instructions), verified byte-for-byte for every question.

Noul questions get accept and escalate cuts from two ReAnchor passes with
mirrored asymmetric metrics (one wrong automatic decision costs nine right
ones); Choice questions keep the confidence sweep. GEPA wording search is
unchanged. An offline end-to-end test runs the real GEPA and ReAnchor path
against a fake Decisions endpoint and is the gate before any paid run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas
…ess (CONTRACT EDIT FOR OWNER REVIEW)

Output of `jev-optimize optimize` on the rules, triage and closure synthetic
corpora, merged with `write-registry`. Eight instruction texts change (four
rules, four triage), each with a test-accuracy gain; closure wording is
unchanged. Accept and escalate thresholds are refitted on every question at
the 10% automatic-error operating point, so weak questions (closure.modified,
closure.fulfilled, rules.event_match) escalate more to the chat model.

The model-boundary checker pin is the checker's own ConvertTo-Json hash; it
differs from write-registry's Python hash only in tuned_at, which PowerShell
re-emits in the local offset.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas
@urnlahzer
urnlahzer merged commit 2e2c655 into main Sep 27, 2026
2 checks passed
urnlahzer added a commit that referenced this pull request Oct 2, 2026
…gistry (#58)

PR #57 re-derived contracts/model/decision-questions.json (accept/escalate thresholds
changed; ids, fields and option keys did not) without touching the Rust tests, so 13
tests in review_model::scanning that hard-coded probabilities tuned to the old
thresholds failed on main. Production code reads every threshold from the embedded
registry at run time and is correct for the new values; only the fixtures were stale.

Test module of crates/openloops-desktop/src/review_scan.rs only:
- new helpers registry_accept(id), registry_gray(id) (escalate..accept midpoint) and
  micros(p), following the pattern every_rule_gray_reject_and_error_falls_back already
  used, so the next re-tune does not break these tests;
- FixedDecisionClient::new sends a gray closure.outcome choice (was 0.5, now below the
  new 0.8 escalate cut) so the noul answers still decide;
- each of the 13 tests takes its accept/gray values from the registry; no assertion
  removed. compare_prediction_flags_the_registry_gray_band_for_nouls_and_choices now
  asserts the fixed 0.5 noul label at the gray value and adds a 0.1 false-branch case.

Verified: cargo fmt --check; clippy -D warnings (desktop native-ui,ui-screenshot);
desktop 424 passed, 0 failed, 3 ignored; no Cargo, tools, or contracts changes.

Registry values for the owner to review (unchanged here): closure.outcome accept 1.0 /
escalate 0.8 means a choice below 0.8 ends the pair with no noul fallback; deadline_kind
accept 1.0 almost never auto-applies; three passing tests sit exactly on new cuts.


Claude-Session: https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
urnlahzer added a commit that referenced this pull request Oct 2, 2026
* test(review): derive decision-test probabilities from the question registry

PR #57 re-derived contracts/model/decision-questions.json (accept/escalate thresholds
changed; ids, fields and option keys did not) without touching the Rust tests, so 13
tests in review_model::scanning that hard-coded probabilities tuned to the old
thresholds failed on main. Production code reads every threshold from the embedded
registry at run time and is correct for the new values; only the fixtures were stale.

Test module of crates/openloops-desktop/src/review_scan.rs only:
- new helpers registry_accept(id), registry_gray(id) (escalate..accept midpoint) and
  micros(p), following the pattern every_rule_gray_reject_and_error_falls_back already
  used, so the next re-tune does not break these tests;
- FixedDecisionClient::new sends a gray closure.outcome choice (was 0.5, now below the
  new 0.8 escalate cut) so the noul answers still decide;
- each of the 13 tests takes its accept/gray values from the registry; no assertion
  removed. compare_prediction_flags_the_registry_gray_band_for_nouls_and_choices now
  asserts the fixed 0.5 noul label at the gray value and adds a 0.1 false-branch case.

Verified: cargo fmt --check; clippy -D warnings (desktop native-ui,ui-screenshot);
desktop 424 passed, 0 failed, 3 ignored; no Cargo, tools, or contracts changes.

Registry values for the owner to review (unchanged here): closure.outcome accept 1.0 /
escalate 0.8 means a choice below 0.8 ends the pair with no noul fallback; deadline_kind
accept 1.0 almost never auto-applies; three passing tests sit exactly on new cuts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas

* feat(graph): mail-provider abstraction, shipped-registration hooks, per-provider sessions

Design doc: docs/plans/2026-09-27-google-provider-and-shared-registration.md
(approved 2026-09-27). This is the graph-crate half of PR-1.

- live/provider.rs: MailProvider {Microsoft, Google}, AccountConfig enum dispatch,
  ProviderLoad and load_all (continues after one provider fails).
- live/registration.rs: build-time injected shared registrations via option_env!
  (OPENLOOPS_MS_CLIENT_ID, OPENLOOPS_GOOGLE_CLIENT_ID, OPENLOOPS_GOOGLE_CLIENT_SECRET),
  BYO-over-shipped precedence, and the tenant admin-consent URL. Nothing is committed
  to the repository; a source build without the variables behaves as before.
- live/google/mod.rs: GoogleConfig validation only; every Google call returns
  ProviderUnavailable until PR-2.
- live.rs: SESSIONS has one slot per provider (clear_session_for, clear_all_sessions,
  has_session_for; the Microsoft-named wrappers keep their behaviour); OAuthEndpoints +
  authorize_with parametrise authority, token URI, redirect host, optional installed-app
  client secret, extra params and scope normalisation, with MICROSOFT reproducing the
  current flow exactly; any refresh_token in a token response is removed and zeroed
  before the OAuth library parses it; AADSTS65001/90094 -> AdminConsentRequired and
  AADSTS650052/650056 -> PublisherNotTrusted, content-free; ConnectionError text is
  provider-neutral.
- live/callback.rs: RedirectHost {Localhost, LoopbackIp}; the Host check follows it.
- live/review.rs: provider field on MailItem, SourceReview, UserIdentity (Microsoft
  default, account string unchanged so decision fingerprints are stable);
  fetch_from_origin_with_headers, add_hydrated and cutoff made crate-shareable.
- live/reminders.rs: ReminderRequest.provider; valid_graph_id -> valid_remote_id.
- live/test_support.rs: shared fake servers plus a path-routed server for PR-2.
- desktop: struct literals gain provider: Microsoft; the reminder-status test's three
  expected strings follow the neutral ConnectionError wording (no behaviour change).

Verified: cargo fmt --check; clippy -D warnings (graph live-connection, desktop
native-ui,ui-screenshot); graph 161 passed; desktop 424 passed, 0 failed, 3 ignored (the
branch sits on PR #58's scanning-test fix, 4d969d4, so the suite is green); Cargo.lock
unchanged; the twelve tools/check-*.ps1 checkers match the origin/main baseline
(the four that fail there fail identically here).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas

* feat(desktop): provider-aware review model, account-qualified scan scope, shipped-registration wiring

Desktop half of PR-1 (design: docs/plans/2026-09-27-google-provider-and-shared-registration.md).

- review_scan/review_model/app_model/slint_review: ConversationKey = (account, conversation)
  replaces every bare conversation-id set (ScanScope::Incremental, retry/check outcomes,
  append_sources, merge_scan, retryable_conversations, prior_open_items); closes the
  cross-account collision before a second provider exists. Exchange split-thread merging
  is Microsoft-only: a ThreadGroup with non-Microsoft messages is authoritative and never
  merged. ReviewMessage carries provider.
- loop_state: Reminder::Created/Completed carry provider; wire bytes 0..3 unchanged
  (decode as Microsoft), 4/5 = Google; legacy-bytes test pins the old layout.
- Links: is_trusted_message_link accepts the Outlook hosts and https://mail.google.com/;
  open-external accepts any provider's tasks URL; "Open message in {Outlook|Gmail}" labels
  are bound per card (url-label) instead of literal.
- AccountDisplay::connected(&[MailProvider]) names one or both providers; the title bar
  binds account-text.
- Outcome::Connection(provider, ..), Service::Mailbox; AppModel gains google: Status,
  effective_microsoft_client_id (BYO field > shipped registration > none),
  shared_registration_active, admin_consent_url, exposed to the Sources screen as
  shared-registration-active / admin-consent-url (layout lands in PR-1c).
- settings: fields 13 google_client_id, 14 google_client_secret (zeroizing),
  15 mail_providers tag (ms | google | ms,google; absent -> ms) on the append-only record;
  boundary, round-trip, unknown-tag and size tests.
- graph nits from the PR-1a review: exchange_token_request takes one token URI;
  GoogleConfig exposes client_id()/client_secret() instead of a filler assertion.

- Review fixes (Opus read-only review): the sign-in button gates on the effective client
  ID, not the BYO field; the Microsoft reminder sync/complete paths bind
  provider: Microsoft so a Google record can never reach Graph; apply_reconcile keeps the
  record's provider; conversation_quality / notes / rejection_reasons are keyed by
  ConversationKey; per-message key clones removed; shared borrowed_filter helper;
  AccountDisplay::connected(&[]) keeps the empty-name contract; TODO_URL derives from
  MailProvider::Microsoft.tasks_url(). Deferred to PR-2 (noted in the plan):
  Service::Mailbox(MailProvider) and the microsoft_status() rename.

Microsoft decision fingerprints and relation keys are byte-identical to before.

Verified: cargo fmt --check; clippy -D warnings (graph live-connection, desktop
native-ui,ui-screenshot); graph 161 passed; desktop 436 passed, 0 failed, 3 ignored;
Cargo.lock unchanged; public-repo gate on staged files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCesN8ZHnhqa6n3rWFThas

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant