diff --git a/LIMITATIONS.md b/LIMITATIONS.md index d3a80f2e..44881c22 100644 --- a/LIMITATIONS.md +++ b/LIMITATIONS.md @@ -69,7 +69,7 @@ ground truth until revalidated. |---|-----------|--------|-------| | L49 | **Agent-cycle and institutional approval evidence are governed internal review surfaces, not external model certification** — Trellis now exposes a stable `agent_cycle` result surface for quant/critic/arbiter/model-validator evidence, model-promotion eligibility, and benchmark trigger rates, and `policy_bundle.production.institutional` can require approval, model-review, snapshot, run-artifact, and audit-bundle evidence before production execution, but these surfaces only certify recorded internal governance evidence for the run or model version | Product and desk review surfaces can show why a cycle passed, failed, was unavailable, or lacked required institutional approval artifacts, but they must not be read as external model approval, regulatory sign-off, xVA/FpML coverage, or correctness beyond the recorded validation scope | `trellis/agent/cycle_surface.py`, `trellis/agent/task_runtime.py`, `trellis/platform/models.py`, `trellis/platform/policies.py`, `trellis/platform/services/pricing_service.py`, `docs/developer/audit_and_observability.rst`, `docs/developer/hosting_and_configuration.rst`, `docs/user_guide/pricing.rst` | | L57 | **Typed comparison-target coherence exposes rather than fills missing numerical composition** — Explicit targets carry canonical method, route/binding, variant, validation, and semantic identities, and the runtime rejects partial explicit target sets, ambiguous references, missing declarations, and unbound shared artifacts before pricing. Explicit semantic axes are now projected onto the per-target `ProductIR` before route selection, which prevents serialized target prose from reclassifying terminal spread baskets and lets T102/T126 bind their independent Stulz, Kirk, Monte Carlo, and Hurd-Zhou lanes. T102 now proves its Monte Carlo variant through authored `n_paths`, `n_steps`, `seed`, and `mc_method` spec overrides while its Stulz reference binds the raw analytical kernel. Other variant execution still must be proven by spec overrides or a canonical full-contract executable declaration; sparse legacy targets still rely on visibly inferred contracts, and `T13` remains an exact-bucket semantic-guard canary whose cached analytical artifact cannot honestly represent both theta-PDE variants and the analytical reference | Coherence can select and prove reusable numerical composition when it exists, but it deliberately does not invent missing methods, infer undeclared semantics for sparse legacy targets, or let one cached artifact impersonate several variants | `trellis/agent/comparison_target_contracts.py`, `trellis/agent/assembly_tools.py`, `trellis/agent/task_runtime.py`, `trellis/agent/executor.py`, `TASKS_PROOF_LEGACY.yaml`, `docs/developer/task_and_eval_loops.rst`, `docs/quant/pricing_stack.rst` | -| L64 | **Legacy proof-task contracts remain broadly incomplete** — the task-manifest gate now validates all modern corpus envelopes strictly and freezes the legacy corpus's exact field-level debt and normalized task content behind a checked baseline. After the authored T02, T17, and T102 repairs, that baseline contains 590 exact issue identities across 122 incomplete retained rows; those rows still lack one or more authored descriptions, economic contracts, market contracts, acceptance criteria, or explicit execution/hold dispositions. Product-specific field sufficiency remains owned by the semantic validators and repair tickets | New or worsened structural manifest debt fails the corpus gate, and the main task runner plus specific-id rerunner reject selected incomplete legacy rows before default market construction or code generation. The exact authored T02, T17, and T102 rows pass that boundary without weakening the remaining debt. The legacy baseline is only a migration guard; it must not be interpreted as evidence that title-only rows are priceable, that every specialized proof harness is governed by the main runner boundary, or that every structurally valid modern product contract is semantically complete | `TASKS_PROOF_LEGACY.yaml`, `TASKS_PROOF_LEGACY_BASELINE.yaml`, `trellis/agent/task_manifest_validation.py`, `scripts/validate_task_manifests.py`, `scripts/run_tasks.py`, `scripts/rerun_ids.py`, `doc/plan/active__task-manifest-integrity.md` | +| L64 | **Legacy proof-task contracts remain broadly incomplete** — the task-manifest gate now validates all modern corpus envelopes strictly and freezes the legacy corpus's exact field-level debt and normalized task content behind a checked baseline. After the authored T02, T17, T89, and T102 repairs, that baseline contains 585 exact issue identities across 121 incomplete retained rows; those rows still lack one or more authored descriptions, economic contracts, market contracts, acceptance criteria, or explicit execution/hold dispositions. Product-specific field sufficiency remains owned by the semantic validators and repair tickets | New or worsened structural manifest debt fails the corpus gate, and the main task runner plus specific-id rerunner reject selected incomplete legacy rows before default market construction or code generation. The exact authored T02, T17, T89, and T102 rows pass that boundary without weakening the remaining debt. T89 is only a same-payoff Hull-White duration identity at constant zero OAS, not independent validation or market-price OAS calibration. The legacy baseline is only a migration guard; it must not be interpreted as evidence that title-only rows are priceable, that every specialized proof harness is governed by the main runner boundary, or that every structurally valid modern product contract is semantically complete | `TASKS_PROOF_LEGACY.yaml`, `TASKS_PROOF_LEGACY_BASELINE.yaml`, `trellis/agent/task_manifest_validation.py`, `scripts/validate_task_manifests.py`, `scripts/run_tasks.py`, `scripts/rerun_ids.py`, `doc/plan/active__task-manifest-integrity.md` | | L65 | **Callable-bond coupons are scalar fixed-rate only** — the checked callable-bond cashflow, compiler, lattice, and PDE routes accept one scalar coupon rate rather than a dated variable-coupon schedule. Legacy task T09 asks for a step-up callable bond but does not author the coupon rates or effective dates, so its validated task contract now fails closed with an exact `variable_coupon_schedule` blocker and zero build attempts instead of using the old title-derived flat 5% fixture | Trellis can price the bounded fixed-coupon callable-bond cohort, but it cannot reasonably price step-up, step-down, floating, or otherwise variable-coupon callable bonds. T09 remains an expected honest block until QUA-1251 adds a reusable dated coupon primitive and an explicit schedule | `trellis/models/short_rate_fixed_income.py`, `trellis/instruments/callable_bond.py`, `trellis/models/callable_bond_pde.py`, `trellis/execution/compiler.py`, `trellis/agent/task_runtime.py`, `TASKS_PROOF_LEGACY.yaml`, `tests/test_agent/test_task_runtime.py` | | L66 | **Physical Bermudan swaption lattice support is a bounded one-factor, static-basis composition** — the strict route preserves explicit co-terminal swap tails, separate named discount/forecast curves, complete supported leg conventions, a provenance-complete named constant-parameter Hull-White set, and authored uniform-grid controls. It supports physical settlement, simple floating coupons with a deterministic additive forward basis, and ACT/365F model time. Stochastic basis, reset/payment convexity, compounded overnight coupons, amortizing or scheduled notionals, seasoned or pre-started fixed tails requiring accrued-settlement treatment, ACT/ACT ICMA coupon accrual without explicit quasi-coupon reference periods, parameterized day-of-month rolls, non-shipped calendar aliases, cash/annuity settlement, term-structured model volatility, multi-factor rates, native Greeks, and production convergence/error governance remain unsupported | Trellis can reasonably price the checked bounded physical dual-curve contract only when each adjusted first fixed accrual start is on or after exercise; it must fail closed rather than value a whole already-started fixed coupon, reinterpret richer Bermudan swaptions through the legacy T04 helper, use a European/Black fallback, or approximate a convention mapping. P005 is the exact executable evidence: its strict lattice lane prices and remains visible even though the paired Monte Carlo lane honestly blocks | `trellis/models/rate_swap_tail.py`, `trellis/models/hull_white_parameters.py`, `trellis/agent/semantic_contracts.py`, `trellis/agent/executor.py`, `trellis/agent/knowledge/canonical/routes.yaml`, `TASKS_EXTENSION.yaml`, `tests/test_models/test_rate_swap_tail.py`, `tests/test_agent/test_physical_bermudan_swaption_semantics.py`, `tests/test_tasks/test_p005_physical_bermudan_swaption.py`, `docs/quant/lattice_algebra.rst`, `docs/user_guide/pricing.rst` | | L58 | **FpML normalization is bounded to fixed-float IRS, physical European swaption, and scheduled cap/floor cohorts** — Support-contract version 1.0.0 distinguishes secure inspection, economic normalization, executable structural lowering, and paired conformance. `make_fpml_request(...)` and `trellis.io.fpml` securely inspect inline UTF-8 FpML 5.13 confirmation `dataDocument` payloads and normalize one regular, single-currency, constant-notional fixed-float swap into `StaticLegContractIR`, one physically settled European payer/receiver swaption into `ContractIR` with the complete swap nested under `underlying_contract`, or one regular single-currency constant-strike cap/floor into the existing signed `PeriodRateOptionStripLeg`. Swaptions reuse the structural resolved Black-76 declaration; cap/floors reuse the existing static strip declaration; historical settled premiums are reported separately and excluded from contract identity. `TASKS_FPML_CONFORMANCE.yaml` pairs all three admitted cohorts with independently specified native contracts and proves identity, projection, structural selection, market binding, price, and non-economic envelope invariance; its negative cohort certifies exact honest blockers with zero agent calls. This evidence does not widen support. Trellis still does not perform complete XSD validation, support other views/versions, resolve external references, bind imported seasoned coupons to historical fixing histories, classify vendor extension children, or normalize amortizing, compounding, stubbed, end-of-month or clamped high-day, cross-currency, OIS, inflation, lifecycle, package, cash-settled/Bermudan/American/partial/automatic/straddle swaption, unsettled-premium, cap/floor collar, stepped-strike, averaged, geared/spread, early-terminable, or other product forms | An admitted swap, physical European swaption, or scheduled cap/floor strip can price deterministically through shared structural execution when the caller declares a valuation party and valuation date and all cohort constraints hold. Unclassified extension children and all other FpML economics remain fail-closed with exact import, clarification, conflict, or unsupported-feature blockers; this is not general FpML pricing coverage | `TASKS_FPML_CONFORMANCE.yaml`, `trellis/io/fpml/`, `trellis/agent/fpml_conformance.py`, `trellis/agent/contract_ir.py`, `trellis/agent/static_leg_contract.py`, `trellis/agent/imported_documents.py`, `trellis/agent/platform_requests.py`, `trellis/platform/executor.py`, `docs/developer/fpml_support_matrix.rst`, `docs/developer/fpml_import.rst`, `docs/quant/contract_ir.rst`, `docs/quant/static_leg_contract_ir.rst`, `doc/plan/draft__fpml-interoperability-roadmap.md` | diff --git a/TASKS_PROOF_LEGACY.yaml b/TASKS_PROOF_LEGACY.yaml index f71c88b5..2461b8ac 100644 --- a/TASKS_PROOF_LEGACY.yaml +++ b/TASKS_PROOF_LEGACY.yaml @@ -1611,15 +1611,68 @@ tasks: analytical: attribution_identity status: pending - id: T89 - title: OAS duration (spread duration) for callable bonds - construct: analytics + title: 'Callable effective duration: constant-zero-OAS same-payoff identity' + description: >- + Compare OASDuration with no market-price anchor and Duration on the named + USD fixed-coupon callable-bond proof fixture using the same Hull-White + tree. Hold OAS at zero and use symmetric 25 bp parallel discount-zero-rate + shifts, with current callable holder PV as denominator. Require effective + duration in years under the authored relative-duration tolerance; this is + a same-payoff same-model identity, not an independent oracle. + task_disposition: named_proof_fixture + proof_fixture_id: usd_fixed_coupon_callable_bond_5pct_2025_2035_v1 + market_scenario_id: usd_callable_fixed_5pct_proof + validation_policy: invariants_and_cross_method + construct: lattice new_component: oas_duration cross_validate: internal: - oas_bump_duration - - modified_duration - external: - - quantlib + - same_payoff_parallel_duration + reference_target: same_payoff_parallel_duration + relations: {oas_bump_duration: within_tolerance} + tolerance_pct: 0.000001 + tolerance_unit: percent_of_reference_price + output_unit: currency_amount + output_currency: USD + output_tolerances_pct: {effective_duration: 0.000001} + analytics_contract: + profile: callable_same_payoff_duration_v1 + output_name: effective_duration + output_unit: years + tolerance_unit: percent_of_reference_duration + bump_bps: 25.0 + market_price: null + constant_oas_bps: 0.0 + risk_coordinate: parallel_continuously_compounded_discount_zero_rate + shock_convention: symmetric_up_down + denominator: 2_times_bump_decimal_times_current_callable_holder_pv + derivative_method: finite_difference + forward_policy: rederive_default_preserve_named_forecast_curves + model_policy: preserve_parameters_recalibrate_tree + reference_role: same_payoff_same_model_identity_not_independent_oracle + measures: + oas_bump_duration: OASDuration + same_payoff_parallel_duration: Duration + target_contracts: + oas_bump_duration: &t89_same_callable_hw_target + method: rate_tree + route_id: exercise_lattice + route_family: rate_lattice + backend_binding_id: trellis.models.trees.algebra.price_on_lattice + variant_parameters: + lattice_model: hull_white + model_parameter_set: callable_fixed_5pct_proof:hull_white + mean_reversion: 0.1 + sigma: 0.01 + tree_steps: 200 + validation_bundle_id: rate_tree:callable_bond + payoff_family: callable_fixed_income + exercise_style: issuer_call + model_family: interest_rate + observation_style: exercise_schedule + equivalence_group: t89_same_callable_hull_white_payoff + same_payoff_parallel_duration: *t89_same_callable_hw_target status: pending - id: T90 title: 'Vega surface: per-expiry per-strike vega bucketing' diff --git a/TASKS_PROOF_LEGACY_BASELINE.yaml b/TASKS_PROOF_LEGACY_BASELINE.yaml index df2021c6..eabe8d32 100644 --- a/TASKS_PROOF_LEGACY_BASELINE.yaml +++ b/TASKS_PROOF_LEGACY_BASELINE.yaml @@ -1,8 +1,8 @@ version: 1 manifest: TASKS_PROOF_LEGACY.yaml -issue_count: 590 -issue_digest: 454a9b098b6d8de5754b9e151004670d2b33c667155ca8eea2bc928b5705acd6 -task_fingerprint: 55b9bcdd3364c0e1c0ca6d96f3cc9caf17e7cf4a78a40af92541804e0c368876 +issue_count: 585 +issue_digest: 819df8b48dc63330687fa7657ecc92a22720c7df7674f1305d1a34f5e8f14b9c +task_fingerprint: 8b09f608fec66515f832a3b60fb66b9a435caaa68bd427fe0b411f50c52af04d policy: exact_issue_identity_and_task_content note: >- This baseline freezes known contract incompleteness in the retained legacy diff --git a/doc/plan/active__legacy-task-migration-map.md b/doc/plan/active__legacy-task-migration-map.md index 845bd8ff..0ddc72f5 100644 --- a/doc/plan/active__legacy-task-migration-map.md +++ b/doc/plan/active__legacy-task-migration-map.md @@ -22,6 +22,11 @@ effective dates, while the checked callable-bond routes accept only one scalar fixed coupon. `QUA-1251` owns the missing variable-coupon primitive and an authored T09 schedule. +`QUA-1258` authors `T89` on the same named fixed-coupon fixture as T02/T17. +It requires constant-zero-OAS effective duration and the same-payoff parallel +duration reference in years at symmetric 25 bp shocks. This internal identity +does not claim an independent oracle or market-price OAS calibration. + T03, T83, and T85 no longer cite resolved limitations as blockers. Their manifest dispositions now match this map: T03 is a `research_hold` pending a reproducible lattice experiment, T83 is a `proof_hold` pending an authored @@ -119,7 +124,7 @@ remain outside executable pricing selection. | `T86` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | Theta (time decay) for options via tree and PDE | | `T87` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | Rho (rate sensitivity) for equity options | | `T88` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | Book P&L attribution: rate + spread + vol decomposition | -| `T89` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | OAS duration (spread duration) for callable bonds | +| `T89` | `named_proof_fixture` | `TASKS_PROOF_LEGACY.yaml` | Callable effective duration: constant-zero-OAS same-payoff identity | | `T90` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | Vega surface: per-expiry per-strike vega bucketing | | `T94` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | FX market bridge: Garman-Kohlhagen vs MC with explicit domestic/foreign curve selection | | `T95` | `proof_only_hold` | `TASKS_PROOF_LEGACY.yaml` | xVA framework: CVA + DVA + FVA on IR swap portfolio | diff --git a/docs/developer/task_and_eval_loops.rst b/docs/developer/task_and_eval_loops.rst index df4aa860..7f773a8b 100644 --- a/docs/developer/task_and_eval_loops.rst +++ b/docs/developer/task_and_eval_loops.rst @@ -856,7 +856,7 @@ payoff, schedule, strike/coupon terms, or settlement rule. Those bridge decisions are task-runner contracts, not general natural-language parser behavior. -Legacy ``T02`` and ``T17`` are executable only through the named +Legacy ``T02``, ``T17``, and ``T89`` are executable only through the named ``usd_fixed_coupon_callable_bond_5pct_2025_2035_v1`` proof fixture and ``usd_callable_fixed_5pct_proof`` market scenario. Manifest loading hydrates the complete economic contract before provenance is recorded; the fixture @@ -868,6 +868,24 @@ targets, and records per-target acceptance. The product-specific validator rejects altered fields, missing units, changed references, or relaxed controls as ``legacy.callable_bond_invalid_contract``. +``T89`` retains holder-PV price reporting in USD and declares its required +``effective_duration`` output separately in ``cross_validate.analytics_contract``. +The two pricing targets explicitly share the same Hull-White execution identity; +their analytics are ``OASDuration(None, 25bp)`` and ``Duration(25bp)``, not different +pricing models. The internal task-analytics bridge executes those public +measures over the bound payoff and validates finite non-boolean values, years, +successful status, the current-price denominator, the 25 bp shock and resolved +finite-difference provenance before extracting scalars. Per-output units and +provenance survive in ``output_metadata`` and target ``output_acceptance``. +Existing ``output_tolerances_pct`` performs the relative-duration comparison. +Missing, failed, non-finite, misunitized, or mismatched duration fails the task +even when both holder prices agree; no external oracle is claimed. +Direct ``run_task`` calls recognize the reserved, whitespace-normalized ``T89`` +ID independently of mutable corpus/manifest labels and validate its exact +contract before market construction or build attempts. Contradictory task kinds +cannot redirect T89 through the early FpML dispatcher; ordinary FpML conformance +requests retain their existing dispatch behavior. + ``T09`` follows the same fail-closed boundary for a different reason. Its title requests a step-up issuer-callable bond but the retained row does not author the dated coupon rates or effective dates, and the checked callable-bond diff --git a/docs/quant/differentiable_pricing.rst b/docs/quant/differentiable_pricing.rst index 5ff3d3e3..7545be7f 100644 --- a/docs/quant/differentiable_pricing.rst +++ b/docs/quant/differentiable_pricing.rst @@ -394,6 +394,24 @@ analytics contract. Runtime ``Duration`` on this curve still reports and preserves the existing market model parameters. It does not make callable tree Greeks autodifferentiable or add bucket shocks or curve rebuilding. +The retained ``T89`` proof compares ``OASDuration(market_price=None, +bump_bps=25)`` with ``Duration(bump_bps=25)`` on the same fixed-coupon +callable-bond economics and Hull-White tree (``a=0.1``, ``sigma=0.01``, +200 steps). Both use the dated flat 5% USD discount curve and reprice the +issuer-call decision after symmetric 25 bp shifts to continuously compounded +discount zero rates. The default derived forward curve is regenerated; +independently named forecast curves and explicit/named model parameters are +preserved, and the tree is recalibrated to each shifted discount curve. + +``effective_duration`` is reported in years as +``(PV_down - PV_up) / (2 * 0.0025 * current_callable_holder_PV)``. OAS remains +zero: no market-price OAS solve is performed. The ``Duration`` lane must resolve +to ``parallel_curve_bump`` (finite difference), not an autodiff result evaluated +at a different shock scale. The reference ``same_payoff_parallel_duration`` is +an internal same-payoff, same-model identity with relative tolerance +``0.000001%`` of reference duration; it is not an independent pricing oracle, +a market-price OAS calibration test, or a Vega proof. + Product-Family Gradient Matrix ------------------------------ diff --git a/docs/user_guide/pricing.rst b/docs/user_guide/pricing.rst index 0718617c..d24ce0a8 100644 --- a/docs/user_guide/pricing.rst +++ b/docs/user_guide/pricing.rst @@ -1307,6 +1307,14 @@ change those economics. These rows demonstrate the bounded fixed-coupon issuer-call route; they do not imply support for step-up/floating coupons, arbitrary callable products, PSOR, or an independent external BDT benchmark. +``T89`` reuses this bond and dated market for a narrower analytics proof: +constant-zero-OAS effective duration versus parallel-curve-bump duration on the +same Hull-White payoff, using symmetric 25 bp shocks. Duration is required in +years and compared within ``0.000001%`` of the reference duration; holder PV +remains a separate USD output. Equal prices cannot substitute for a missing or +failed duration. This is an internal consistency check, not independent +validation, market-price OAS calibration, or a callable Vega claim. + Some retained rows are intentionally executable only as certified honest blocks. For example, T09 asks for a step-up callable bond without supplying the dated coupon schedule, while the checked callable-bond routes currently accept diff --git a/tests/test_tasks/test_t89_callable_duration.py b/tests/test_tasks/test_t89_callable_duration.py new file mode 100644 index 00000000..986475b9 --- /dev/null +++ b/tests/test_tasks/test_t89_callable_duration.py @@ -0,0 +1,363 @@ +"""T89 defends a same-payoff duration identity, not an external price oracle.""" + +from copy import deepcopy +from dataclasses import replace +import math + +import pytest + +from trellis.agent.benchmark_contracts import benchmark_spec_overrides +from trellis.agent.task_manifests import load_task_manifest +from trellis.agent.task_runtime import build_market_state_for_task +from trellis.analytics.measures import Duration, OASDuration +from trellis.instruments.callable_bond import CallableBondPayoff, CallableBondSpec + + +def _task(): + return next(task for task in load_task_manifest("TASKS_PROOF_LEGACY.yaml") if task["id"] == "T89") + + +@pytest.fixture +def isolated_task_artifacts(monkeypatch, tmp_path): + import sys + import trellis.agent.analytical_traces as analytical_traces + import trellis.agent.executor as executor + import trellis.agent.model_audit as model_audit + import trellis.agent.platform_requests as platform_requests + import trellis.agent.platform_traces as platform_traces + + def write_generated_module(module_path, code): + output = tmp_path / "generated" / module_path + output.parent.mkdir(parents=True, exist_ok=True) + output.write_text(code) + return output + + monkeypatch.setattr(executor, "REPO_ROOT", tmp_path) + monkeypatch.setattr(executor, "TRELLIS_PACKAGE_ROOT", tmp_path / "trellis") + monkeypatch.setattr(executor, "_REPO_REVISION", "test") + monkeypatch.setattr(executor, "write_module", write_generated_module) + monkeypatch.setattr(analytical_traces, "TRACE_ROOT", tmp_path / "traces" / "analytical") + monkeypatch.setattr(platform_traces, "TRACE_ROOT", tmp_path / "traces" / "platform") + monkeypatch.setattr(model_audit, "_AUDIT_DIR", tmp_path / "audits") + monkeypatch.setattr(platform_requests, "_record_semantic_extension_artifact", lambda *args, **kwargs: None) + monkeypatch.setenv("TRELLIS_OFFLINE_LOCAL_AGENTS", "1") + monkeypatch.setenv("TRELLIS_SKIP_TASK_DIAGNOSIS_PERSIST", "1") + module_name = "trellis.instruments._agent._fresh.callablebond" + previous = sys.modules.get(module_name) + yield + if previous is None: + sys.modules.pop(module_name, None) + else: + sys.modules[module_name] = previous + + +def test_t89_authors_duration_separately_from_holder_price_and_ignores_title(): + from trellis.agent.task_runtime import _effective_task_description, task_to_semantic_contract + + task = _task() + assert task["proof_fixture_id"] == "usd_fixed_coupon_callable_bond_5pct_2025_2035_v1" + assert task["market_scenario_id"] == "usd_callable_fixed_5pct_proof" + comparison = task["cross_validate"] + assert comparison["reference_target"] == "same_payoff_parallel_duration" + assert comparison["output_tolerances_pct"] == {"effective_duration": 0.000001} + assert comparison["output_unit"] == "currency_amount" + assert comparison["analytics_contract"]["output_unit"] == "years" + assert comparison["analytics_contract"]["bump_bps"] == 25.0 + assert comparison["analytics_contract"]["market_price"] is None + assert comparison["analytics_contract"]["constant_oas_bps"] == 0.0 + renamed = {**task, "title": "Unrelated title containing neither bond nor duration"} + assert _effective_task_description(task) == _effective_task_description(renamed) + assert benchmark_spec_overrides(task) == benchmark_spec_overrides(renamed) + assert task_to_semantic_contract(renamed).product.observation_schedule == ( + "2028-01-15", "2030-01-15", "2032-01-15", + ) + + +def test_duration_bridge_executes_authored_measures_on_dated_curve(monkeypatch): + from trellis.agent.task_analytics import evaluate_comparison_analytics + from trellis.curves.yield_curve import YieldCurve + + task = _task() + market, _ = build_market_state_for_task(task) + market = replace(market, forecast_curves={"independent_forecast": YieldCurve.flat(0.07)}) + observed_markets = [] + + class ObservedPayoff(CallableBondPayoff): + def evaluate(self, market_state): + observed_markets.append(market_state) + return super().evaluate(market_state) + + payoff = ObservedPayoff(CallableBondSpec(**benchmark_spec_overrides(task))) + anchor = payoff.evaluate(market) + up = payoff.evaluate(replace(market, discount=market.discount.shift(25.0), forward_curve=None)) + down = payoff.evaluate(replace(market, discount=market.discount.shift(-25.0), forward_curve=None)) + expected = (down - up) / (2 * 0.0025 * anchor) + observed_markets.clear() + def forbidden_oas_solve(*args, **kwargs): + raise AssertionError("constant-zero-OAS proof must not solve a market-price OAS") + monkeypatch.setattr("trellis.analytics.measures._callable_market_price_oas_bps", forbidden_oas_solve) + calls = [] + for measure_cls in (Duration, OASDuration): + original = measure_cls.compute + def record(self, actual_payoff, actual_market, _original=original, **ctx): + calls.append((type(self).__name__, self.bump_bps, actual_payoff, actual_market, ctx["base_price"])) + return _original(self, actual_payoff, actual_market, **ctx) + monkeypatch.setattr(measure_cls, "compute", record) + for target in task["cross_validate"]["internal"]: + output = evaluate_comparison_analytics( + payoff, market, target_id=target, + contract=task["cross_validate"]["analytics_contract"], anchor_price=anchor, + )["effective_duration"] + assert output["value"] == pytest.approx(expected, rel=1e-12) + assert output["unit"] == "years" + assert output["metadata"]["derivative_method"] == "finite_difference" + assert output["metadata"]["resolved_derivative_method"] == "parallel_curve_bump" + assert output["metadata"]["anchor_price"] == pytest.approx(anchor) + assert [(name, bump) for name, bump, *_ in calls] == [("OASDuration", 25.0), ("Duration", 25.0)] + assert all(actual is payoff and ms is market and base == anchor for _, _, actual, ms, base in calls) + assert expected == pytest.approx(5.533555119264439, abs=1e-9) + assert len(observed_markets) == 4 + assert sorted(ms.discount.flat_rate for ms in observed_markets) == pytest.approx([0.0475, 0.0475, 0.0525, 0.0525]) + for shifted in observed_markets: + assert shifted.forecast_curves == market.forecast_curves + assert shifted.model_parameters == market.model_parameters + assert shifted.model_parameter_sets == market.model_parameter_sets + assert shifted.selected_curve_names == market.selected_curve_names + assert shifted.discount.value_date == market.discount.value_date + assert shifted.discount.curve_day_count is market.discount.curve_day_count + assert shifted.forward_curve.discount_date(payoff.spec.end_date) == pytest.approx( + shifted.discount.discount_date(payoff.spec.end_date), + ) + + +@pytest.mark.parametrize("mutation", [ + lambda task: task["cross_validate"]["analytics_contract"].pop("output_unit"), + lambda task: task["cross_validate"]["analytics_contract"].update(output_unit="percent"), + lambda task: task["cross_validate"]["analytics_contract"].update(bump_bps=1.0), + lambda task: task["cross_validate"]["analytics_contract"].update(market_price=100.0), + lambda task: task["cross_validate"]["analytics_contract"].update(constant_oas_bps=10.0), + lambda task: task["cross_validate"].pop("output_tolerances_pct"), + lambda task: task["cross_validate"].update(reference_target="modified_duration"), + lambda task: task["benchmark_contract"].update(call_dates=["2028-01-15"]), +]) +def test_t89_rejects_contract_drift(mutation): + from trellis.agent.task_manifest_validation import _validate_legacy_task + + task = deepcopy(_task()) + mutation(task) + issues = _validate_legacy_task("TASKS_PROOF_LEGACY.yaml", task, "tasks[0]") + assert any(issue.code == "legacy.callable_bond_invalid_contract" for issue in issues) + + +@pytest.mark.global_workflow +def test_t89_fresh_offline_run_requires_duration_and_authored_tree_controls(monkeypatch, tmp_path, isolated_task_artifacts): + from trellis.agent.task_runtime import run_task + from trellis.models.trees import lattice + + monkeypatch.setenv("TRELLIS_OFFLINE_LOCAL_AGENTS", "1") + observed = [] + original = lattice.build_generic_lattice + + def observe(*args, **kwargs): + observed.append(dict(kwargs)) + return original(*args, **kwargs) + + monkeypatch.setattr(lattice, "build_generic_lattice", observe) + task = {**_task(), "title": "Unrelated renamed proof"} + result = run_task(task, None, fresh_build=True, task_run_storage_root=tmp_path) + assert result["success"], result + report = result["cross_validation"] + assert report["status"] == "passed" + assert report["output_validation"]["effective_duration"]["output_unit"] == "years" + for target in task["cross_validate"]["internal"]: + assert report["outputs"][target]["price"] == pytest.approx(96.8252972854, abs=1e-7) + assert report["outputs"][target]["effective_duration"] == pytest.approx(5.533555119264439, abs=1e-8) + assert report["output_metadata"][target]["effective_duration"]["bump_bps"] == 25.0 + assert report["target_acceptance"][target]["output_acceptance"]["effective_duration"]["output_unit"] == "years" + assert result["method_results"][target]["generated_artifact"]["is_fresh_build"] + assert observed + assert all(call["n_steps"] == 200 for call in observed) + assert all(call["a"] == 0.1 and call["sigma"] == 0.01 for call in observed) + observed_rates = { + round(-math.log(float(call["discount_curve"].discount(1.0))), 4) + for call in observed + } + assert {0.0475, 0.05, 0.0525} <= observed_rates + + +@pytest.mark.parametrize("value, metadata", [ + (True, {}), (float("inf"), {}), (float("nan"), {}), + (5.5, {"status": "failed"}), (5.5, {"output_unit": "days"}), +]) +def test_duration_bridge_rejects_raw_bad_measure_before_float_conversion(monkeypatch, value, metadata): + from trellis.agent.task_analytics import evaluate_comparison_analytics + from trellis.analytics.result import ScalarRiskMeasureOutput + + raw = ScalarRiskMeasureOutput(value, metadata=metadata) if metadata else value + monkeypatch.setattr(OASDuration, "compute", lambda *args, **kwargs: raw) + with pytest.raises(ValueError): + evaluate_comparison_analytics( + object(), object(), target_id="oas_bump_duration", + contract=_task()["cross_validate"]["analytics_contract"], anchor_price=96.0, + ) + + +@pytest.mark.global_workflow +@pytest.mark.parametrize("mode", ["mismatch", "failed_status", "missing", "exception"]) +def test_t89_real_run_fails_duration_defect_while_prices_agree(monkeypatch, tmp_path, isolated_task_artifacts, mode): + from trellis.agent.task_runtime import run_task + from trellis.analytics.result import ScalarRiskMeasureOutput + + def bad_measure(*args, **kwargs): + if mode == "exception": + raise ValueError("injected duration failure") + if mode == "missing": + return None + if mode == "failed_status": + return ScalarRiskMeasureOutput(5.533555119264439, metadata={"status": "failed"}) + return 6.0 + + monkeypatch.setattr(OASDuration, "compute", bad_measure) + result = run_task(_task(), None, fresh_build=True, task_run_storage_root=tmp_path) + assert result["success"] is False + assert result["passed_expectation"] is False + report = result["cross_validation"] + assert len(set(report["prices"].values())) == 1 + assert report["status"] == "failed" + assert report["output_validation"]["effective_duration"]["status"] != "passed" + assert all(method["success"] for method in result["method_results"].values()) + + +@pytest.mark.parametrize("task_id", ["T89", " T89 "]) +@pytest.mark.parametrize("field", ["analytics_contract", "output_tolerances_pct"]) +@pytest.mark.parametrize("provenance", ["authored", "removed", "spoofed"]) +def test_t89_direct_run_rejects_missing_duration_requirement_before_build( + field, task_id, provenance, monkeypatch, tmp_path, +): + from trellis.agent.task_runtime import run_task + + task = deepcopy(_task()) + task["id"] = task_id + task["cross_validate"].pop(field) + if provenance == "removed": + task.pop("task_corpus", None) + task.pop("task_definition_manifest", None) + elif provenance == "spoofed": + task.update(task_corpus="extension", task_definition_manifest="TASKS_EXTENSION.yaml") + calls = [] + def forbidden_market(*args, **kwargs): + calls.append("market") + raise AssertionError("invalid T89 must fail before market construction") + monkeypatch.setattr("trellis.agent.task_runtime.build_market_state_for_task", forbidden_market) + result = run_task(task, None, build_fn=lambda **kwargs: calls.append(kwargs), task_run_storage_root=tmp_path) + assert result["success"] is False + assert not calls + assert "legacy.callable_bond_invalid_contract" in result["error"] + + +@pytest.mark.parametrize("task_id", ["T89", " T89 "]) +@pytest.mark.parametrize("removed_field", [None, "analytics_contract", "output_tolerances_pct"]) +def test_t89_cannot_bypass_pricing_contract_through_fpml_dispatch( + task_id, removed_field, monkeypatch, tmp_path, +): + from trellis.agent.task_runtime import run_task + + task = deepcopy(_task()) + task.update(id=task_id, task_kind="fpml_conformance") + if removed_field: + task["cross_validate"].pop(removed_field) + calls = [] + def forbidden_market(*args, **kwargs): + calls.append("market") + raise AssertionError("reserved T89 must not reach FpML market construction") + monkeypatch.setattr("trellis.agent.task_runtime.build_market_state_for_task", forbidden_market) + result = run_task(task, None, build_fn=lambda **kwargs: calls.append(kwargs), task_run_storage_root=tmp_path) + assert result["success"] is False + assert not calls + assert "legacy.callable_bond_invalid_contract" in result["error"] + + +@pytest.mark.parametrize("mutation", [ + lambda output: output.clear(), + lambda output: output["effective_duration"].update(value=float("inf")), + lambda output: output["effective_duration"].update(value=float("nan")), + lambda output: output["effective_duration"].update(value=True), + lambda output: output["effective_duration"].update(value=6.0), + lambda output: output["effective_duration"].update(value=5.50000011), + lambda output: output["effective_duration"].update(unit="days"), + lambda output: output["effective_duration"].update(status="failed"), + lambda output: output["effective_duration"]["metadata"].update(bump_bps=1.0), + lambda output: output["effective_duration"]["metadata"].update(unit="days"), + lambda output: output["effective_duration"]["metadata"].update(output_unit="days"), + lambda output: output["effective_duration"]["metadata"].update(status="failed"), + lambda output: output["effective_duration"]["metadata"].update(derivative_method_category="autograd"), +]) +def test_duration_comparison_fails_bad_output_even_when_prices_agree(monkeypatch, mutation): + from types import SimpleNamespace + from trellis.agent import task_analytics + from trellis.agent.task_runtime import _cross_validate_comparison_task, _task_comparison_targets + + task = _task() + targets = _task_comparison_targets(task, ["rate_tree"]) + market, _ = build_market_state_for_task(task) + payoff = CallableBondPayoff(CallableBondSpec(**benchmark_spec_overrides(task))) + # A reported native price cannot replace the actual bound-payoff anchor. + payoff.benchmark_outputs = lambda _: {"price": 1.0, "effective_duration": 5.5} + live = {target.target_id: SimpleNamespace(success=True, payoff_cls=type(payoff)) for target in targets} + monkeypatch.setattr("trellis.agent.task_runtime._comparison_target_binding_report", lambda target, _: { + "status": "bound_unique_artifact", "artifact_identity": target.target_id, "failures": [], + }) + def outputs(actual_payoff, actual_market, *, target_id, contract, anchor_price): + output = {"effective_duration": { + "value": 5.5, "unit": "years", "status": "passed", "metadata": { + "measure": contract["measures"][target_id], "bump_bps": 25.0, + "derivative_method": "finite_difference", "resolved_derivative_method": "parallel_curve_bump", + "market_price": None, "constant_oas_bps": 0.0, "anchor_price": anchor_price, + "denominator": contract["denominator"], "risk_coordinate": contract["risk_coordinate"], + "reference_role": contract["reference_role"], + }, + }} + # Mutate the reference as well: +inf must not make the tolerance infinite. + if target_id == "same_payoff_parallel_duration": + mutation(output) + return output + monkeypatch.setattr(task_analytics, "evaluate_comparison_analytics", outputs) + report = _cross_validate_comparison_task( + targets, live, market, configured_targets=task["cross_validate"], + payoff_factory=lambda *args: payoff, price_fn=lambda *args: 96.0, + ) + assert report["prices"] == {target.target_id: 96.0 for target in targets} + assert report["status"] == "failed" + assert report["output_validation"]["effective_duration"]["status"] != "passed" + if report["output_validation"]["effective_duration"]["status"] == "failed": + assert report["output_validation"]["effective_duration"]["deviations_pct"]["oas_bump_duration"] > 0.000001 + + +@pytest.mark.parametrize("value, metadata", [ + (True, {}), (float("inf"), {}), (float("nan"), {}), (0.0, {}), + (96.0, {"status": "failed"}), (96.0, {"output_unit": "years"}), +]) +def test_t89_rejects_invalid_evaluator_anchor_before_analytics(monkeypatch, value, metadata): + from types import SimpleNamespace + from trellis.agent import task_analytics + from trellis.agent.task_runtime import _cross_validate_comparison_task, _task_comparison_targets + from trellis.analytics.result import ScalarRiskMeasureOutput + + task = _task() + targets = _task_comparison_targets(task, ["rate_tree"]) + market, _ = build_market_state_for_task(task) + payoff = CallableBondPayoff(CallableBondSpec(**benchmark_spec_overrides(task))) + live = {target.target_id: SimpleNamespace(success=True, payoff_cls=type(payoff)) for target in targets} + monkeypatch.setattr("trellis.agent.task_runtime._comparison_target_binding_report", lambda target, _: { + "status": "bound_unique_artifact", "artifact_identity": target.target_id, "failures": [], + }) + analytics_calls = [] + monkeypatch.setattr(task_analytics, "evaluate_comparison_analytics", lambda *args, **kwargs: analytics_calls.append(kwargs)) + raw = ScalarRiskMeasureOutput(value, metadata=metadata) if metadata else value + report = _cross_validate_comparison_task( + targets, live, market, configured_targets=task["cross_validate"], + payoff_factory=lambda *args: payoff, price_fn=lambda *args: raw, + ) + assert not analytics_calls + assert not report["prices"] + assert report["status"] != "passed" diff --git a/trellis/agent/benchmark_contracts.py b/trellis/agent/benchmark_contracts.py index b31aba9e..be5eb895 100644 --- a/trellis/agent/benchmark_contracts.py +++ b/trellis/agent/benchmark_contracts.py @@ -270,7 +270,7 @@ def benchmark_request_description( product = str(contract.get("product") or "").strip().lower() if ( product == "callable_bond" - and str(task.get("id") or "").strip() in {"T02", "T17"} + and str(task.get("id") or "").strip() in {"T02", "T17", "T89"} and str(task.get("proof_fixture_id") or "").strip() == "usd_fixed_coupon_callable_bond_5pct_2025_2035_v1" ): diff --git a/trellis/agent/task_analytics.py b/trellis/agent/task_analytics.py new file mode 100644 index 00000000..e5812bfd --- /dev/null +++ b/trellis/agent/task_analytics.py @@ -0,0 +1,164 @@ +"""Bounded, declared analytics over an already-bound comparison payoff. + +This is task orchestration, not a pricing adapter or an independent oracle. +The callable duration identity delegates both computations to public measures. +""" + +from __future__ import annotations + +from collections.abc import Mapping +import math +from typing import Any + + +def callable_duration_contract() -> dict[str, Any]: + """Return the admitted same-payoff, constant-zero-OAS duration contract.""" + return { + "profile": "callable_same_payoff_duration_v1", + "output_name": "effective_duration", + "output_unit": "years", + "tolerance_unit": "percent_of_reference_duration", + "bump_bps": 25.0, + "market_price": None, + "constant_oas_bps": 0.0, + "risk_coordinate": "parallel_continuously_compounded_discount_zero_rate", + "shock_convention": "symmetric_up_down", + "denominator": "2_times_bump_decimal_times_current_callable_holder_pv", + "derivative_method": "finite_difference", + "forward_policy": "rederive_default_preserve_named_forecast_curves", + "model_policy": "preserve_parameters_recalibrate_tree", + "reference_role": "same_payoff_same_model_identity_not_independent_oracle", + "measures": { + "oas_bump_duration": "OASDuration", + "same_payoff_parallel_duration": "Duration", + }, + } + + +def validated_callable_anchor_price(value: Any) -> float: + """Validate the authoritative evaluator output before scalar conversion.""" + if ( + isinstance(value, bool) or not isinstance(value, (int, float)) + or not math.isfinite(value) or value <= 0.0 + ): + raise ValueError("Duration anchor must be a finite positive non-boolean holder PV") + metadata = dict(getattr(value, "metadata", {}) or {}) + if any( + key in metadata and metadata[key] != expected + for key, expected in ( + ("status", "passed"), ("unit", "currency_amount"), + ("output_unit", "currency_amount"), ("output_currency", "USD"), + ) + ): + raise ValueError("Duration anchor has invalid price units or failed status") + return float(value) + + +def evaluate_comparison_analytics( + payoff, market_state, *, target_id: str, contract: Mapping[str, Any], + anchor_price: float, +) -> dict[str, dict[str, Any]]: + """Execute the declared measure without substituting a price comparison.""" + from trellis.analytics.measures import Duration, OASDuration + + if dict(contract) != callable_duration_contract(): + raise ValueError("Unsupported comparison analytics contract") + anchor_price = validated_callable_anchor_price(anchor_price) + measure_name = contract["measures"][target_id] + bump_bps = float(contract["bump_bps"]) + measure = ( + OASDuration(market_price=None, bump_bps=bump_bps) + if measure_name == "OASDuration" + else Duration(bump_bps=bump_bps) + ) + value = measure.compute( + payoff, market_state, base_price=anchor_price, + _cache={"base_price": anchor_price}, + ) + metadata = dict(getattr(value, "metadata", {}) or {}) + if isinstance(value, bool) or not isinstance(value, (int, float)) or not math.isfinite(value): + raise ValueError("effective_duration must be finite and numeric") + if ( + metadata.get("output_unit", "years") != "years" + or metadata.get("unit", "years") != "years" + or metadata.get("status", "passed") != "passed" + ): + raise ValueError("effective_duration has invalid units or failed status") + if any( + key in metadata and metadata[key] != expected + for key, expected in ( + ("resolved_derivative_method", "parallel_curve_bump"), + ("derivative_method_category", "finite_difference_bump"), + ("derivative_method", "finite_difference"), + ("bump_bps", bump_bps), + ) + ): + raise ValueError("effective_duration reports conflicting derivative provenance") + if measure_name == "Duration" and ( + metadata.get("resolved_derivative_method") != "parallel_curve_bump" + or metadata.get("derivative_method_category") != "finite_difference_bump" + or metadata.get("bump_bps") != bump_bps + ): + raise ValueError("Callable Duration did not resolve the authored finite difference") + # OASDuration(None) has one implementation: parallel shifts, no OAS solve. + metadata.update({ + "measure": measure_name, + "derivative_method": "finite_difference", + "resolved_derivative_method": "parallel_curve_bump", + "bump_bps": bump_bps, + "constant_oas_bps": 0.0, + "market_price": None, + "anchor_price": anchor_price, + "denominator": contract["denominator"], + "risk_coordinate": contract["risk_coordinate"], + "reference_role": contract["reference_role"], + }) + return {"effective_duration": { + "value": float(value), "unit": "years", "status": "passed", "metadata": metadata, + }} + + +def validated_comparison_analytics_values( + payload: Mapping[str, Any], *, target_id: str, contract: Mapping[str, Any], + anchor_price: float, +) -> tuple[dict[str, float], dict[str, dict[str, Any]]]: + """Check required value, unit and execution evidence before scalar extraction.""" + if not isinstance(payload, Mapping) or set(payload) != {"effective_duration"}: + raise ValueError("Missing required effective_duration output") + output = payload["effective_duration"] + if not isinstance(output, Mapping) or output.get("unit") != "years": + raise ValueError("effective_duration must be reported in years") + if output.get("status") != "passed": + raise ValueError("effective_duration output did not pass") + value = output.get("value") + if isinstance(value, bool) or not isinstance(value, (int, float)) or not math.isfinite(value): + raise ValueError("effective_duration must be finite and numeric") + metadata = output.get("metadata") + expected = { + "measure": contract["measures"][target_id], + "derivative_method": "finite_difference", + "resolved_derivative_method": "parallel_curve_bump", + "bump_bps": contract["bump_bps"], + "constant_oas_bps": 0.0, + "market_price": None, + "anchor_price": anchor_price, + "denominator": contract["denominator"], + "risk_coordinate": contract["risk_coordinate"], + "reference_role": contract["reference_role"], + } + if not isinstance(metadata, Mapping) or any( + key not in metadata or metadata[key] != expected_value + for key, expected_value in expected.items() + ): + raise ValueError("effective_duration execution metadata does not match its contract") + if any( + key in metadata and metadata[key] != expected_value + for key, expected_value in ( + ("unit", "years"), ("output_unit", "years"), ("status", "passed"), + ("derivative_method_category", "finite_difference_bump"), + ) + ): + raise ValueError("effective_duration has conflicting nested output metadata") + return {"effective_duration": float(value)}, {"effective_duration": { + **dict(metadata), "unit": "years", "status": "passed", + }} diff --git a/trellis/agent/task_manifest_validation.py b/trellis/agent/task_manifest_validation.py index 66c5c55f..5dc3056d 100644 --- a/trellis/agent/task_manifest_validation.py +++ b/trellis/agent/task_manifest_validation.py @@ -1562,7 +1562,7 @@ def _validate_legacy_task( ) ) - if task_id in {"T02", "T17"}: + if task_id in {"T02", "T17", "T89"}: issues.extend( _validate_legacy_callable_bond_comparison_contract( manifest_name, @@ -1652,7 +1652,7 @@ def _validate_legacy_callable_bond_comparison_contract( *, root: Path | None = None, ) -> list[TaskManifestIssue]: - """Keep T02/T17 on one exact, reusable fixed-coupon proof fixture.""" + """Keep callable price/duration proofs on one exact, reusable fixture.""" task_id = _text(task.get("id")) contract = task.get("benchmark_contract") cross_validate = task.get("cross_validate") @@ -1749,6 +1749,48 @@ def _validate_legacy_callable_bond_comparison_contract( }, "external": ["financepy", "quantlib"], } + elif task_id == "T89": + from trellis.agent.task_analytics import callable_duration_contract + + expected_description = ( + "Compare OASDuration with no market-price anchor and Duration on the named " + "USD fixed-coupon callable-bond proof fixture using the same Hull-White " + "tree. Hold OAS at zero and use symmetric 25 bp parallel discount-zero-rate " + "shifts, with current callable holder PV as denominator. Require effective " + "duration in years under the authored relative-duration tolerance; this is " + "a same-payoff same-model identity, not an independent oracle." + ) + expected_construct = "lattice" + duration_target = { + "method": "rate_tree", + "route_id": "exercise_lattice", + "route_family": "rate_lattice", + "backend_binding_id": "trellis.models.trees.algebra.price_on_lattice", + "variant_parameters": { + "lattice_model": "hull_white", + "model_parameter_set": "callable_fixed_5pct_proof:hull_white", + "mean_reversion": 0.1, + "sigma": 0.01, + "tree_steps": 200, + }, + **common_target, + "equivalence_group": "t89_same_callable_hull_white_payoff", + } + expected_cross_validate = { + "internal": ["oas_bump_duration", "same_payoff_parallel_duration"], + "reference_target": "same_payoff_parallel_duration", + "relations": {"oas_bump_duration": "within_tolerance"}, + "tolerance_pct": 0.000001, + "tolerance_unit": "percent_of_reference_price", + "output_unit": "currency_amount", + "output_currency": "USD", + "output_tolerances_pct": {"effective_duration": 0.000001}, + "analytics_contract": callable_duration_contract(), + "target_contracts": { + target_id: dict(duration_target) + for target_id in ("oas_bump_duration", "same_payoff_parallel_duration") + }, + } else: expected_description = ( "Price the named USD fixed-coupon callable-bond proof fixture with " @@ -1875,6 +1917,7 @@ def _validate_legacy_callable_bond_comparison_contract( valid = all( ( + task_id != "T89" or _text(task.get("task_kind")) in {"", "pricing"}, _text(task.get("task_disposition")) == "named_proof_fixture", _text(task.get("proof_fixture_id")) == "usd_fixed_coupon_callable_bond_5pct_2025_2035_v1", @@ -1909,7 +1952,7 @@ def _validate_legacy_callable_bond_comparison_contract( _issue( manifest_name, "legacy.callable_bond_invalid_contract", - "T02/T17 require the exact named fixed-coupon callable-bond proof contract", + "T02/T17/T89 require the exact named fixed-coupon callable-bond proof contract", task_id=task_id, path=path, ) diff --git a/trellis/agent/task_runtime.py b/trellis/agent/task_runtime.py index 2a447032..03606ea4 100644 --- a/trellis/agent/task_runtime.py +++ b/trellis/agent/task_runtime.py @@ -1542,7 +1542,7 @@ def _proof_legacy_semantic_contract(task: dict, description: str): option_type="put", ) - if task_id in {"T02", "T17"}: + if task_id in {"T02", "T17", "T89"}: from trellis.agent.semantic_contracts import make_callable_bond_contract contract = task.get("benchmark_contract") @@ -2302,7 +2302,11 @@ def run_task( ) -> dict: """Execute one task with separate artifact-freshness and source-origin policy.""" assert_executable_task_disposition([task]) - if str(task.get("task_kind") or "").strip() == "fpml_conformance": + is_t89_duration_proof = str(task.get("id") or "").strip() == "T89" + if ( + str(task.get("task_kind") or "").strip() == "fpml_conformance" + and not is_t89_duration_proof + ): from trellis.agent.fpml_conformance import run_fpml_conformance_task market_state, _market_context = build_market_state_for_task(task, market_state) @@ -2438,6 +2442,14 @@ def run_task( result_data["llm_cassette"] = llm_cassette_payload try: + if is_t89_duration_proof: + from trellis.agent.task_manifest_validation import assert_executable_task_selection + + # The reserved ID, not mutable provenance labels, identifies this + # proof. Direct callers cannot remove its required-output contract. + assert_executable_task_selection([{ + **task, "task_definition_manifest": "TASKS_PROOF_LEGACY.yaml", + }]) expected_honest_block = _expected_honest_block_for_task(task) if expected_honest_block is not None: raise ExpectedTaskHonestBlock(expected_honest_block) @@ -5031,6 +5043,8 @@ def _cross_validate_comparison_task( priced: dict[str, float] = {} outputs: dict[str, dict[str, float]] = {} output_errors: dict[str, str] = {} + output_metadata: dict[str, dict[str, Any]] = {} + analytics_contract = configured_targets.get("analytics_contract") output_tolerances = { str(name): float(tolerance) for name, tolerance in dict( @@ -5129,7 +5143,7 @@ def _cross_validate_comparison_task( continue try: target_outputs: dict[str, float] = {} - if output_tolerances: + if output_tolerances and analytics_contract is None: benchmark_outputs_fn = getattr(payoff, "benchmark_outputs", None) if callable(benchmark_outputs_fn): try: @@ -5143,7 +5157,38 @@ def _cross_validate_comparison_task( except Exception as exc: output_errors[target.target_id] = str(exc) if "price" not in target_outputs: - target_outputs["price"] = float(price_fn(payoff, market_state)) + raw_price = price_fn(payoff, market_state) + if analytics_contract is not None: + from trellis.agent.task_analytics import validated_callable_anchor_price + + target_outputs["price"] = validated_callable_anchor_price(raw_price) + else: + target_outputs["price"] = float(raw_price) + if analytics_contract is not None: + from trellis.agent.task_analytics import ( + evaluate_comparison_analytics, + validated_comparison_analytics_values, + ) + + # Required analytics are computed here, not inferred from + # equal prices or untyped adapter benchmark outputs. + for name in output_tolerances: + target_outputs.pop(name, None) + try: + analytics_payload = evaluate_comparison_analytics( + payoff, market_state, target_id=target.target_id, + contract=analytics_contract, + anchor_price=target_outputs["price"], + ) + values, metadata = validated_comparison_analytics_values( + analytics_payload, target_id=target.target_id, + contract=analytics_contract, + anchor_price=target_outputs["price"], + ) + target_outputs.update(values) + output_metadata[target.target_id] = metadata + except Exception as exc: + output_errors[target.target_id] = str(exc) priced[target.target_id] = float(target_outputs["price"]) outputs[target.target_id] = target_outputs except Exception as exc: @@ -5244,7 +5289,10 @@ def _cross_validate_comparison_task( / denominator * 100.0 ) - output_deviations[target_id] = round(deviation_pct, 4) + output_deviations[target_id] = ( + deviation_pct if analytics_contract is not None + else round(deviation_pct, 4) + ) if abs(values[target_id] - output_reference_value) <= tolerance_amount: output_passed_targets.append(target_id) else: @@ -5267,6 +5315,11 @@ def _cross_validate_comparison_task( "failed_targets": output_failed_targets, "missing_targets": missing_targets, } + if analytics_contract is not None: + output_validation[output_name].update({ + "output_unit": analytics_contract["output_unit"], + "tolerance_unit": analytics_contract["tolerance_unit"], + }) output_validation_failed = any( report["status"] != "passed" for report in output_validation.values() @@ -5303,6 +5356,12 @@ def _cross_validate_comparison_task( "value": report["values"].get(target_id), "deviation_pct": report["deviations_pct"].get(target_id), } + if analytics_contract is not None: + output_acceptance[output_name].update({ + "output_unit": report["output_unit"], + "tolerance_unit": report["tolerance_unit"], + "metadata": output_metadata.get(target_id, {}).get(output_name), + }) output_statuses = { report["status"] for report in output_acceptance.values() } @@ -5369,6 +5428,7 @@ def _cross_validate_comparison_task( "prices": priced, "outputs": outputs, "output_errors": output_errors, + "output_metadata": output_metadata, "output_validation": output_validation, "price_errors": price_errors, "artifact_coherence": artifact_coherence,