Skip to content

QUA-1258: author callable duration comparison - #899

Draft
steveya wants to merge 1 commit into
mainfrom
codex/qua-1258-authored-t89-duration
Draft

steveya wants to merge 1 commit into
mainfrom
codex/qua-1258-authored-t89-duration

Conversation

@steveya

@steveya steveya commented Sep 13, 2026

Copy link
Copy Markdown
Owner

T89 now authors an executable duration comparison on the existing named fixed-coupon USD callable-bond fixture. Both targets use Hull-White (a=0.1, sigma=0.01, 200 steps); the separate required output is effective duration in years at symmetric 25 bp discount-zero-rate shifts, holding OAS at zero.

The task analytics bridge delegates OASDuration(None, 25bp) and Duration(25bp) on the bound payoff. It preserves the current holder-PV anchor, checks units/status/finiteness/shock/derivative provenance, and requires the duration outputs even when the prices agree. This is a same-payoff, same-model identity, not an independent oracle or market-price OAS calibration.

Reserved normalized T89 admission rejects missing contracts independently of mutable corpus/manifest labels. A contradictory task kind cannot bypass that validation through the early FpML dispatcher; ordinary FpML requests are unchanged.

Validation:

  • Final T89 + ordinary FpML suites: 70 passed, including 14 new red-to-green admission bypass cases.
  • Manifest/runtime/benchmark regressions: 472 passed with normal settings. An initial diagnosis-suppressed invocation failed three diagnosis-asserting tests; the normal rerun passed.
  • Fresh strict offline T89: 1/1 passed, zero LLM calls, both holder PVs 96.82529728539951 USD and durations 5.533555119264439 years, zero deviation against 0.000001% tolerance.
  • Manifest gate: 0 blocking issues; 585 frozen legacy debt findings. Remediation analysis: zero failures. Focused lint and diff checks passed; task_runtime.py retains 24 verified pre-existing lint findings.

Updated official quant/developer/user docs, L64, the migration map and debt baseline. Generated validation artifacts were archived outside the repository; only the 12 intended files are included.

Draft: full release validation and coordinated branch integration remain pending. QUA-1261 owns generic OAS state preservation, QUA-1264 owns generic comparison-output hardening outside this profile. QUA-1258 remains In Progress.

@steveya
steveya requested a lite review from Copilot September 13, 2026 06:46
@steveya

steveya commented Sep 13, 2026

Copy link
Copy Markdown
Owner Author

@codex review — Please review the authored T89 same-payoff duration contract, required output validation, and reserved-ID admission before ordinary/FpML dispatch. This remains draft pending broader release validation.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-13T06:49:36.490303Z 6ed08b2 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6ed08b255b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines 2444 to +2450
try:
if is_t89_duration_proof:
from trellis.agent.task_manifest_validation import assert_executable_task_selection

# The reserved ID, not mutable provenance labels, identifies this
# proof. Direct callers cannot remove its required-output contract.
assert_executable_task_selection([{

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Validate reserved T89 before deriving its semantic contract

Move the reserved-ID validation ahead of the ordinary task setup. For a direct T89 call whose benchmark_contract is missing or whose call_dates are absent, task_to_semantic_contract() raises from _proof_legacy_semantic_contract() before execution reaches this try, so run_task() propagates an exception instead of returning its normal structured contract-failure result. This leaves part of the exact T89 contract outside the newly added admission boundary and can abort direct task runners rather than recording the rejected run.

Useful? React with 👍 / 👎.

Comment on lines +103 to +108
# OASDuration(None) has one implementation: parallel shifts, no OAS solve.
metadata.update({
"measure": measure_name,
"derivative_method": "finite_difference",
"resolved_derivative_method": "parallel_curve_bump",
"bump_bps": bump_bps,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Require OAS duration provenance instead of synthesizing it

Do not manufacture the finite-difference provenance after accepting an unannotated result. OASDuration.compute() currently returns a plain float, so its empty metadata passes the optional conflict checks and these lines label any finite value as parallel_curve_bump; if that implementation regresses to a constant or delegates to the reference Duration calculation, T89 can still pass whenever the numbers agree, falsely certifying the intended independent OAS-duration lane. Have the measure return authoritative metadata and require it here before admitting the output.

Useful? React with 👍 / 👎.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Require structured measure output validation before adding contract metadata.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Adds an executable T89 callable-bond duration comparison using the existing Hull–White fixture.

Changes:

  • Adds strict T89 admission, runtime routing, and analytics comparison.
  • Adds tests, manifest/baseline updates, and documentation.
  • Moderate finding (2 votes): duration validation may accept unstructured finite results before adding provenance metadata.
File summaries
File Description
trellis/agent/task_runtime.py Integrates T89 execution and comparison.
trellis/agent/task_manifest_validation.py Enforces the T89 contract.
trellis/agent/task_analytics.py Provides duration analytics and validation.
trellis/agent/benchmark_contracts.py Recognizes the callable-bond fixture.
tests/test_tasks/test_t89_callable_duration.py Covers T89 behavior and rejection cases.
TASKS_PROOF_LEGACY.yaml Defines the executable T89 task.
TASKS_PROOF_LEGACY_BASELINE.yaml Updates legacy debt fingerprints.
LIMITATIONS.md Documents T89 scope and limitations.
docs/user_guide/pricing.rst Documents callable duration behavior.
docs/quant/differentiable_pricing.rst Documents duration semantics.
docs/developer/task_and_eval_loops.rst Documents runtime validation and routing.
doc/plan/active__legacy-task-migration-map.md Updates T89 migration status.
Review details
  • Files reviewed: 12/12 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +81 to +85
if (
metadata.get("output_unit", "years") != "years"
or metadata.get("unit", "years") != "years"
or metadata.get("status", "passed") != "passed"
):
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants