diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..9e572e7 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,462 @@ +# Alphaforge — Multi-Agent Collaboration Guide + +## Repository Overview + +Alphaforge is a point-in-time data and feature engineering library for +systematic research. + +**Language:** Python 3.10+ +**Core dependencies:** pandas, duckdb, pyarrow, PyYAML +**Python:** `/Users/steveyang/miniforge3/bin/python` +**Tests:** `pytest` +**Lint:** `ruff check .` +**Type check:** `mypy alphaforge` +**Docs:** `mkdocs build --strict` + +Prefer the miniforge interpreter above unless a repo-local virtualenv exists and +has already been adopted for the current task. + +## Agent Roles + +### Role: Data Source Developer +**Scope:** `alphaforge/data/public_web/`, `alphaforge/data/sources/`, +`alphaforge/data/transforms/`, `alphaforge/data/registries/` +**Task:** Build, refactor, or extend public-web loaders, source adapters, +registry-backed metadata, and source-specific PIT transforms. +**Rules:** +- Read the source module, its matching test file, and shared helpers before + editing. +- Preserve table names, schema contracts, entity-id semantics, sorting, column + projection, and `asof_utc` behavior unless the ticket explicitly changes the + contract. +- Keep shared abstractions shallow until at least three sources clearly benefit. +- Migrate only a small family of loaders at a time. +- Update targeted source tests in `tests/public_web/` and broader adapter or + regression tests in `tests/` when routing or contracts change. + +### Role: Test & Validation Agent +**Scope:** `tests/`, `tests/public_web/` +**Task:** Add or update targeted tests, broader regression coverage, and review +gates for changed behavior. +**Rules:** +- Prefer the smallest targeted failing test first, then broaden to subsystem + regression coverage. +- Review code and tests together; missing interaction coverage is a finding, not + a note for later. +- When compatibility-only behavior must remain, keep the compatibility boundary + explicit in tests instead of mixing it into ordinary happy-path coverage. +- Run the narrowest useful validation first, then the broader commands required + by the ticket or changed subsystem. + +### Role: Documentation Agent +**Scope:** `docs/api/`, `docs/getting-started/`, `docs/guides/` +**Task:** Keep API reference, onboarding, workflow, and conceptual docs aligned +with shipped behavior. +**Rules:** +- Update docs after implementation and tests stabilize. +- Treat behavior, API, workflow, source-coverage, and validation changes as + documentation work unless the ticket is explicitly doc-free. +- Do not document unsupported source behavior or time semantics that code and + tests do not prove. + +### Role: Planning Agent +**Scope:** `doc/plan/` +**Task:** Maintain mirrored ticket tables, implementation order, and review or +cleanup backlog notes that future agents rely on. +**Rules:** +- Keep mirrored plan rows aligned with Linear, not ahead of it. +- Keep ticket tables sorted in implementation order. +- When a shared abstraction or program plan changes scope, update the relevant + plan doc in the same slice. + +## Module Ownership + +| Path | Owner | Notes | +|------|-------|-------| +| `alphaforge/data/public_web/` | Data Source Developer | Public web source loaders, parsing, HTTP helpers | +| `alphaforge/data/sources/` | Data Source Developer | Unified source adapters and cache-aware wrappers | +| `alphaforge/data/transforms/` | Data Source Developer | Source-specific PIT transforms | +| `alphaforge/data/registries/` | Data Source Developer | Registry-backed source metadata | +| `tests/public_web/` | Test & Validation Agent | Public web loader coverage | +| `tests/` | Test & Validation Agent | Broader regression coverage | +| `docs/api/` | Documentation Agent | API reference updates | +| `docs/getting-started/` | Documentation Agent | Onboarding and quickstart guidance | +| `docs/guides/` | Documentation Agent | Workflow and conceptual guides | +| `doc/plan/` | Planning Agent | Repo-local implementation plans and mirrored ticket tables | + +## Coordination Protocol + +### When changing public-web loaders + +1. Read the source module, its matching test file, and any shared helper modules + it relies on. +2. Preserve table names, schema contracts, and entity-id semantics unless the + ticket explicitly calls for contract changes. +3. Update or add targeted tests in `tests/public_web/`. +4. If the change affects higher-level routing, also update the relevant adapter + tests in `tests/`. +5. Update docs if behavior, supported sources, or developer workflow changes. + +### When adding or changing shared abstractions + +1. Keep the abstraction shallow until at least three sources clearly benefit. +2. Migrate only a small family of loaders at a time. +3. Verify that empty-frame behavior, column projection, sorting, and + `asof_utc` semantics stay stable. +4. Update the relevant plan doc in `doc/plan/` if the abstraction changes + implementation order or scope. + +### When validation or compatibility boundaries change + +1. Review code and tests together. +2. Add or update targeted tests before broadening the regression scope. +3. Keep compatibility-only tests clearly separated from ordinary API-surface + coverage. +4. Update docs or plan notes when the validation strategy or migration boundary + changes. + +## Linear Issue Writing Spec + +Linear is the shared work ledger for cross-agent work. Issues should be +specific, searchable, dependency-aware, and tied to an observable outcome. +Treat an issue as a short engineering spec, not as a note or a chat summary. + +### When to create or update an issue + +- Create a Linear issue when work must survive beyond the current chat, spans + more than one file or module, introduces a blocker, or needs durable tracking. +- Update an existing issue instead of creating a duplicate when the scope is + the same. +- Split a ticket when it contains more than one independent reviewable outcome. + Use an umbrella issue for the broad objective and child issues for delivery + slices. +- Do not mirror every trivial local task into Linear. Use Linear for durable + planning, blockers, coordination, and user-visible work. + +### Title and naming + +- Use outcome-first titles of the form `: `. +- Keep titles short, concrete, and searchable. +- Put the domain noun in the title, not just the implementation verb. +- If the work concerns temporal semantics, PIT APIs, dataset contracts, + compatibility shims, or public-web source families, name that domain + explicitly in the title. +- Avoid vague titles such as `Cleanup`, `Refactor`, or `Improve module` + unless paired with the exact target. +- Good examples: + - `Public web: extract shared source finalization helpers` + - `Registry APIs: add base class for entity-driven public sources` + - `Archive loaders: unify historical-batch URL selection` + +### Issue body shape + +Use a compact spec structure: + +- Objective: what should exist when the issue is done. +- Why now: why this matters now. +- Scope: the exact modules, files, or docs in scope. +- Non-goals: what is explicitly out of scope. +- Dependencies: hard blockers with issue IDs. +- Acceptance criteria: observable conditions that define success. +- Validation: tests, lint, docs, or review gates. +- Follow-on work: separate issues for future slices, if needed. + +### Dependency rules + +- Use parent/child relationships for umbrella work and implementation slices. +- Use `blockedBy` and `blocks` only for hard prerequisites. +- Use `relatedTo` for adjacent work that does not prevent completion. +- Keep blocker chains shallow. +- If the blocker does not yet exist, create it first or state the missing + prerequisite explicitly. +- Do not block on anticipated future reuse alone; keep that as a scoped note + unless the dependency is already real. + +### Priority rules + +- Priority 1 / Urgent: broken build, release blocker, or active outage. +- Priority 2 / High: foundational platform work or an item that unlocks + multiple other tickets. +- Priority 3 / Normal: planned implementation slices and most feature work. +- Priority 4 / Low: docs-only work, cleanup, exploratory refactors. +- Default to Priority 3 unless there is a concrete reason to raise it. + +### Blocker handling + +- If a task is blocked, say so explicitly in the issue and when communicating + with the user. +- State the blocking issue ID(s), the missing prerequisite, and the next + unblock step. +- Never present a blocked issue as complete. +- If the user asks to complete work but a Linear blocker remains, surface the + blocker before claiming success. + +### Done criteria + +- Mark an issue Done only when the implementation slice is landed, validation + passes, and acceptance criteria are satisfied. +- If behavior, APIs, or supported-source coverage changed, update the relevant + docs: + - `docs/api/` for API and reference behavior + - `docs/getting-started/` for onboarding or examples + - `docs/guides/` for workflows and conceptual docs + - `doc/plan/` for mirrored ticket tables and implementation plans +- Close the issue with a short note summarizing the result, validation, docs + changes, and any follow-on issue IDs. +- If useful work remains, split it into follow-on issues instead of leaving the + original issue ambiguous. + +### Umbrella closeout + +- Treat an umbrella issue as a maintenance checkpoint, not just a delivery + milestone. +- Before closing an umbrella, do the cleanup pass, documentation maintenance, + and mirrored-plan updates that the completed slices imply. +- If an umbrella is intentionally doc-free, record that explicitly in the + closeout note and explain why no docs changed. + +### Engineering plan quality + +- Every implementation issue should include a concrete plan with small phases. +- Prefer stable scaffolds plus surgical deltas over whole-module rewrites. +- If the plan cannot be explained as a few reviewable phases, the issue is too + large and should be split. + +## Ticket Implementation Workflow + +All coding agents working from Linear must follow this workflow for every +implementation ticket unless the user explicitly overrides it. + +### Required execution order + +1. Review upstream context before coding. +2. Announce the current ticket number and its plain-English goal on screen. +3. Implement the current ticket with test-driven development. +4. Update the relevant docs after the implementation and tests pass. +5. Leave a handoff note in Linear describing what changed and any caveats. +6. Mark the ticket `Done`, then update the mirrored ticket table in the + relevant plan doc. + +### Step 1: Review upstream context + +Before writing code, the implementer must: + +- read the current ticket body in full +- read all hard-blocking upstream tickets and their completion notes +- read recent comments on the parent issue when the parent is an active + umbrella ticket +- inspect the referenced plan docs and the current code paths in scope +- identify the exact files, tests, and docs that are likely to change + +If an upstream ticket is not done, do not start implementation unless the user +explicitly approves working around the blocker. + +### Step 2: Announce the current ticket on screen + +Before coding, print the ticket number and a short plain-English explanation +of what the ticket aims to do in the current terminal/chat session. + +Minimum expectation: + +- include the Linear ticket id, for example `ALP-123` +- explain the ticket goal in one or two plain-English sentences +- do this after reading upstream tickets and before writing code + +### Step 3: Implement with TDD + +For code-changing tickets: + +- start by adding or updating the tests that define the target behavior +- run the tests and confirm they fail for the expected reason +- implement the smallest coherent code change that makes the tests pass +- rerun the targeted tests, then rerun the broader validation required by the + ticket or module owner rules + +Minimum expectation: + +- targeted tests for the changed behavior +- broader regression coverage for the touched subsystem when practical + +For documentation-only or planning-only tickets, state explicitly in the ticket +that TDD does not apply. + +### Step 4: Update docs after tests pass + +When behavior, APIs, runtime flow, source coverage, governance, validation, or +developer workflow changes, update the relevant docs after the implementation is +stable: + +- `docs/api/` for API and source reference behavior +- `docs/getting-started/` for quickstart and setup +- `docs/guides/` for workflows and design guidance +- `doc/plan/` for mirrored ticket tables and implementation plans + +Doc updates are part of completing the ticket, not optional follow-up work, +unless the ticket is explicitly scoped as doc-free. + +### Step 5: Leave a Linear handoff note + +Before marking the ticket done, add a Linear comment with: + +- what was implemented +- which tests were added or updated +- which test commands were run and whether they passed +- which docs were updated +- any caveats, deferred work, follow-on risks, or compatibility notes + +If the ticket is blocked or only partially complete, leave the same note but do +not mark it done. + +### Step 6: Mark done and update the mirrored plan table + +Linear is the source of truth for ticket state. The plan tables in `doc/plan/` +are the repo-local mirror for subsequent coding agents. + +Mark the ticket `Done` only when: + +- the code is landed +- the agreed validation passed +- the relevant docs are updated or the ticket explicitly records why no docs + changed +- the Linear handoff note is written + +After the ticket is moved to `Done` in Linear: + +- update the corresponding plan-table row in the relevant plan doc +- keep the table sorted in implementation order +- skip tickets already marked `Done` when selecting the next ticket +- do not mark the plan-table row `Done` before the Linear ticket is actually + closed + +### Recommended Linear closeout template + +Use this structure for the final implementation note: + +- Implemented: +- Tests: +- Docs: +- Caveats: +- Follow-ons: + +## Code Review Workflow + +When the task is a code review rather than a feature implementation, use the +repo-wide review program under `doc/plan/` when one exists and the linked +Linear review workstream for ticket state. If Alphaforge does not yet have a +dedicated review program doc or review-ticket queue, treat the user-directed +scope or the current branch diff as the review slice and follow the same +dossier, severity, and closeout standards. + +### Review ticket selection + +- Linear is the source of truth for review-ticket state when review tickets + exist. +- Use the mirrored queue in `doc/plan/` when a dedicated review-program doc is + present. +- Pick the earliest review ticket in the ordered queue whose status is not + `Done`, unless the user explicitly redirects to a different slice. +- If no dedicated review queue exists yet, review the current branch diff or the + user-directed scope instead of inventing tickets. +- Tickets in the same wave may run in parallel only when their write scopes do + not conflict. + +### Review scope and posture + +- Treat review tickets as review-and-fix slices, not as read-only audits. +- Review code and tests together. +- Check local, regional, and global behavior: + - Local: single function, class, or module behavior. + - Regional: bounded interactions across modules, adapters, registries, + transforms, or docs. + - Global: user-visible or workflow-visible behavior across major layers. +- For important behavior, missing regional or global interaction coverage is a + review finding, not a note for later. +- Keep compatibility-only coverage isolated from the ordinary API-surface + suites. +- Use the available test strata explicitly during review and remediation: + - targeted source or module tests for local behavior + - adapter, contract, or regression tests for bounded interactions + - broader workflow or docs validation when the slice affects user-visible + behavior + +### Required review workflow + +1. Read the review ticket or selected scope, the mirrored plan section if one + exists, and any upstream blockers. +2. Build a review dossier: + - files in scope + - tests in scope + - contracts being defended + - inbound and outbound module interactions + - current local / regional / global coverage map + - relevant docs, plan notes, and compatibility or migration notes +3. Announce the ticket number or review slice and the plain-English review goal + before editing. +4. Perform the static review of code and tests. +5. Run the targeted tests for the slice and the broader regression required by + the ticket when practical. +6. Land low-risk in-scope fixes, test additions, doc updates, or plan updates + inside the same ticket when the user asked for review-and-fix work. +7. Split larger remediations into follow-on Linear issues instead of letting + the review ticket sprawl. + +### Findings and severity + +- Findings are the primary output of a review ticket. +- Use this severity rubric: + - `P0`: wrong result, silent corruption, or broken governance on a critical + path + - `P1`: high-confidence correctness, runtime, or contract bug + - `P2`: meaningful maintainability, test, or observability gap with real risk + - `P3`: lower-risk cleanup or consistency gap + +### Review closeout requirements + +Before marking a review ticket `Done`: + +- leave a Linear handoff note with: + - Reviewed: + - Findings: + - Tests checked: + - Coverage assessment: + - Docs / plan mismatches: + - Compatibility or shim removal candidates: + - Follow-on tickets: + - Disposition: +- update the mirrored ticket row in the review-program doc when one exists +- do not mark the ticket `Done` if a hard blocker remains unresolved + +### Code review deliverables + +Every completed review ticket should leave behind: + +- prioritized findings with file references +- a local / regional / global coverage assessment +- missing-test and misleading-test notes +- docs and plan mismatches +- compatibility or shim-removal notes where relevant +- follow-on issues for out-of-scope or larger remediations + +## Anti-Hallucination Rules + +1. Never invent table names, entity ids, registry keys, or source names. Verify + from code or tests. +2. Never claim a public-web source supports a format or time semantics unless + the loader or tests prove it. +3. Never silently broaden a shared abstraction across unrelated source + families. +4. Check file targets before patching. Resolve paths from repo root and verify + with `rg --files` or `git status`. +5. Check docs targets before updating them. Use only existing doc trees unless + the ticket explicitly introduces a new one. + +## Key Files to Read First + +1. `README.md` +2. `pyproject.toml` +3. `doc/plan/` +4. `alphaforge/data/public_web/registry.py` +5. `alphaforge/data/public_web/http.py` +6. `alphaforge/data/public_web/parsing.py` +7. `alphaforge/data/public_web/utils.py` +8. `tests/public_web/test_source_test_mapping.py` diff --git a/CHANGELOG.md b/CHANGELOG.md index 0fe2afa..b776c34 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,9 @@ ## Unreleased +- Fixed `mypy alphaforge` regressions across ref-period PIT query normalization, `DataContext` adapter signatures, PIT adapter wrappers, cache manifest aggregation, futures loaders, and optional-store handling so the existing type-check CI gate passes on this branch. +- Declared explicit `pytz` runtime dependency and `types-pytz` development stub so fresh installs and docs-example CI no longer rely on pandas pulling timezone support transitively. +- Clarified the PIT ref-query public contract: `RefSnapshotQuery` / `RefRevisionQuery` plus `snapshot_ref(...)` / `revisions_ref(...)` are the canonical surface, while `get_snapshot_ref(...)` and `get_revision_timeline_ref(...)` remain compatibility wrappers. - Added unified data layer with `SourceAdapter` protocol, `SourceAdapterBase` mixin, `FetchResult`/`CacheManifest` value types, and `CacheLayer` (DuckDB-backed PIT/market cache). - Added built-in source adapters: `TiingoAdapter` (market OHLCV), `FREDSourceAdapter` (macro PIT), `CFTCAdapter` (CoT positioning), `DTCCAdapter` (swap derivatives). - Added `alphaforge.source_adapters` entry-point group and `discover_adapters()` for plugin-style adapter registration. diff --git a/alphaforge/__init__.py b/alphaforge/__init__.py index 838fdd0..ea78162 100644 --- a/alphaforge/__init__.py +++ b/alphaforge/__init__.py @@ -10,6 +10,8 @@ BCBSGSDataSource, BEADataSource, BLSDataSource, + CFTCCoTSource, + CFTCDisaggregatedCoTSource, CFTCWeeklySwapsSource, CMEProductSlateSource, DestatisGenesisDataSource, @@ -24,7 +26,9 @@ FRBTermStructureBenchmarkSource, IBGESidraDataSource, LCHCDSClearDailySource, + MOFJGBYieldCurveSource, PhiladelphiaSPFMeanLevelSource, + default_public_web_sources, ) from .data.query import Query from .data.schema import TableSchema @@ -37,6 +41,7 @@ ) from .data.universe import EntityMetadata, Universe from .features.frame import Artifact, FeatureFrame +from .features.market import LagReturnsTemplate, RollingVolatilityTemplate from .features.ops import join_feature_frames, materialize from .features.realization import FeatureRealization, FitState from .features.template import FeatureTemplate, ParamSpec, SliceSpec @@ -85,6 +90,12 @@ PITPipelineStep, coerce_pipeline_spec, ) +from .pit.queries import ( + RefRevisionQuery, + RefSnapshotQuery, + coerce_ref_revision_query, + coerce_ref_snapshot_query, +) from .pit.ref_entity import make_ref_entity_id, parse_ref_entity_id from .pit.tasks import ( build_snapshot_tape, @@ -110,7 +121,25 @@ from .time.align import AlignedPanel, AlignSpec, AvailabilityState, align_panel from .time.calendar import TradingCalendar from .time.grids import EventGrid, Grid, NativeGrid, SessionGrid -from .time.ref_period import RefFreq, RefPeriod +from .time.missingness import MissingnessReason, classify_missingness +from .time.ref_period import ( + ObsDateAnchor, + RefFreq, + RefPeriod, + coerce_ref_period, + normalize_obs_date_anchor, + normalize_ref_freq, +) +from .time.release_rules import ( + CalendarDay, + CustomRule, + FixedLagMonths, + NthBusinessDay, + NthWeekday, + QuarterlyRelease, + ReleaseRule, + WeeklyRelease, +) __all__ = [ "DataContext", @@ -132,6 +161,8 @@ "load_first_rate_futures_metadata", "BLSDataSource", "BEADataSource", + "CFTCCoTSource", + "CFTCDisaggregatedCoTSource", "EIADataSource", "EurostatDataSource", "ECBSDMXDataSource", @@ -149,7 +180,9 @@ "LCHCDSClearDailySource", "EzoicAdRevenueDailySource", "FRBTermStructureBenchmarkSource", + "MOFJGBYieldCurveSource", "PhiladelphiaSPFMeanLevelSource", + "default_public_web_sources", "build_kim_orphanides_dataset", "build_policy_rule_dataset", "build_duan_weekly_dataset", @@ -167,9 +200,25 @@ "align_panel", "RefFreq", "RefPeriod", + "ObsDateAnchor", + "coerce_ref_period", + "normalize_ref_freq", + "normalize_obs_date_anchor", + "ReleaseRule", + "NthBusinessDay", + "NthWeekday", + "CalendarDay", + "FixedLagMonths", + "QuarterlyRelease", + "WeeklyRelease", + "CustomRule", + "MissingnessReason", + "classify_missingness", "FeatureFrame", "Artifact", + "LagReturnsTemplate", "ParamSpec", + "RollingVolatilityTemplate", "SliceSpec", "FeatureTemplate", "FeatureRealization", @@ -193,6 +242,10 @@ "PITPipelineSpec", "PITPipelineResult", "coerce_pipeline_spec", + "RefSnapshotQuery", + "RefRevisionQuery", + "coerce_ref_snapshot_query", + "coerce_ref_revision_query", "PITExpressionNode", "PITExpressionGraphSpec", "PITExpressionGraphResult", diff --git a/alphaforge/data/cache_layer.py b/alphaforge/data/cache_layer.py index 76f9634..3c47836 100644 --- a/alphaforge/data/cache_layer.py +++ b/alphaforge/data/cache_layer.py @@ -113,10 +113,11 @@ def store( # Update manifest — query aggregate stats from the full table # Get total row count for this dataset+source in the table - total_rows = self._conn.execute( + total_rows_row = self._conn.execute( f"SELECT COUNT(*) FROM {table} WHERE dataset = ? AND source = ?", [dataset, source], - ).fetchone()[0] + ).fetchone() + total_rows = int(total_rows_row[0]) if total_rows_row is not None else 0 # Get all entity keys for this dataset+source all_keys = self._conn.execute( @@ -124,14 +125,16 @@ def store( [dataset, source], ).fetchdf()["series_key"].tolist() - all_min = self._conn.execute( + all_min_row = self._conn.execute( f"SELECT MIN(obs_date) FROM {table} WHERE dataset = ? AND source = ?", [dataset, source], - ).fetchone()[0] - all_max = self._conn.execute( + ).fetchone() + all_min = all_min_row[0] if all_min_row is not None else None + all_max_row = self._conn.execute( f"SELECT MAX(obs_date) FROM {table} WHERE dataset = ? AND source = ?", [dataset, source], - ).fetchone()[0] + ).fetchone() + all_max = all_max_row[0] if all_max_row is not None else None entity_keys_str = ",".join(all_keys) diff --git a/alphaforge/data/context.py b/alphaforge/data/context.py index bf7851d..3d2d7e9 100644 --- a/alphaforge/data/context.py +++ b/alphaforge/data/context.py @@ -1,8 +1,8 @@ from __future__ import annotations from dataclasses import dataclass, field -from datetime import timedelta -from typing import TYPE_CHECKING, Mapping, Optional +from datetime import date, timedelta +from typing import TYPE_CHECKING, Mapping, Optional, Sequence import pandas as pd @@ -11,7 +11,7 @@ from ..store.store import Store from ..time.calendar import TradingCalendar from .panel import PanelFrame -from .query import Query +from .query import Query, VintageMode from .source import DataSource from .universe import EntityMetadata, Universe @@ -27,20 +27,23 @@ class DataContext: Parameters ---------- sources : Mapping[str, DataSource] - Legacy data source mapping (backward compatibility). + Legacy data source mapping kept for backward compatibility and + raw-loader workflows. calendars : Mapping[str, TradingCalendar] Trading calendar lookup. store : Store Backing store for persistence. adapters : dict[str, SourceAdapter] | None - Unified source adapters keyed by source_name (e.g. ``"cftc"``). + Canonical public data-loading surface keyed by source_name + (e.g. ``"cftc"``). default_sources : dict[str, str] | None - Maps dataset → default source_name (e.g. ``{"cot.tff": "cftc"}``). + Maps dataset → default source_name for canonical adapter routing + (e.g. ``{"cot.tff": "cftc"}``). """ sources: Mapping[str, DataSource] calendars: Mapping[str, TradingCalendar] - store: Store + store: Store | None universe: Optional[Universe] = None entity_meta: Optional[EntityMetadata] = None adapters: Optional[dict[str, "SourceAdapter"]] = field(default=None) @@ -52,6 +55,47 @@ class DataContext: init=False, default_factory=dict, repr=False ) + @classmethod + def from_adapters( + cls, + *adapters: "SourceAdapter", + calendars: Mapping[str, TradingCalendar] | None = None, + store: Store | None = None, + universe: Optional[Universe] = None, + entity_meta: Optional[EntityMetadata] = None, + default_sources: Optional[dict[str, str]] = None, + ) -> "DataContext": + """Build a DataContext from adapters without manual mapping boilerplate.""" + adapter_map: dict[str, "SourceAdapter"] = {} + dataset_to_sources: dict[str, list[str]] = {} + + for adapter in adapters: + if adapter.source_name in adapter_map: + raise ValueError( + f"Duplicate adapter source_name: {adapter.source_name!r}" + ) + adapter_map[adapter.source_name] = adapter + for dataset in adapter.datasets: + dataset_to_sources.setdefault(dataset, []).append(adapter.source_name) + + derived_defaults = { + dataset: sources[0] + for dataset, sources in dataset_to_sources.items() + if len(sources) == 1 + } + merged_defaults = dict(derived_defaults) + merged_defaults.update(default_sources or {}) + + return cls( + sources={}, + calendars=dict(calendars or {}), + store=store, + universe=universe, + entity_meta=entity_meta, + adapters=adapter_map, + default_sources=merged_defaults or None, + ) + def __post_init__(self) -> None: if isinstance(self.store, DuckDBParquetStore): self.pit = PITAccessor(self.store.conn()) @@ -63,6 +107,7 @@ def __post_init__(self) -> None: self._dataset_to_sources.setdefault(ds, []).append(source_name) def fetch_panel(self, source: str, q: Query) -> PanelFrame: + """Legacy panel-building path for DataSource-backed loaders.""" df = self.sources[source].fetch(q) try: @@ -145,12 +190,29 @@ def _resolve_source( # Use default_sources mapping if self.default_sources and dataset in self.default_sources: default_src = self.default_sources[dataset] - if default_src in self.adapters: - return self.adapters[default_src] + if default_src not in self.adapters: + raise KeyError( + f"Default source '{default_src}' for dataset '{dataset}' is not " + f"registered. Available: {sorted(self.adapters.keys())}" + ) + adapter = self.adapters[default_src] + if dataset not in adapter.datasets: + raise KeyError( + f"Default source '{default_src}' does not serve dataset " + f"'{dataset}'. It serves: {sorted(adapter.datasets)}" + ) + return adapter # Fallback: find any adapter that serves this dataset if dataset in self._dataset_to_sources: - src_name = self._dataset_to_sources[dataset][0] + source_names = self._dataset_to_sources[dataset] + if len(source_names) > 1: + raise KeyError( + f"Multiple adapters serve dataset '{dataset}': " + f"{sorted(source_names)}. Configure default_sources or " + "pass source= explicitly." + ) + src_name = source_names[0] return self.adapters[src_name] raise KeyError( @@ -165,10 +227,42 @@ def fetch( source: Optional[str] = None, max_staleness: Optional[timedelta] = None, ) -> "FetchResult": - """Unified fetch: resolve adapter and delegate.""" + """Canonical fetch path: resolve an adapter and delegate.""" adapter = self._resolve_source(query.table, source) return adapter.fetch(query, max_staleness=max_staleness) + def load( + self, + dataset: str, + *, + columns: Sequence[str], + start: Optional[pd.Timestamp | str] = None, + end: Optional[pd.Timestamp | str] = None, + entities: Optional[Sequence[str]] = None, + asof: Optional[pd.Timestamp | str] = None, + vintage: VintageMode = "latest", + vintage_id: Optional[str] = None, + grid: Optional[str] = None, + source: Optional[str] = None, + max_staleness: Optional[timedelta] = None, + ) -> "FetchResult": + """Happy-path source load without explicit Query construction.""" + return self.fetch( + Query( + table=dataset, + columns=list(columns), + start=start, + end=end, + entities=list(entities) if entities is not None else None, + asof=asof, + vintage=vintage, + vintage_id=vintage_id, + grid=grid, + ), + source=source, + max_staleness=max_staleness, + ) + def fetch_many( self, queries: list[Query], @@ -176,19 +270,43 @@ def fetch_many( source: Optional[str] = None, max_staleness: Optional[timedelta] = None, ) -> list["FetchResult"]: - """Fetch multiple queries, routing each to the correct adapter.""" - results = [] - for q in queries: - adapter = self._resolve_source(q.table, source) - results.append(adapter.fetch(q, max_staleness=max_staleness)) - return results + """Canonical batch fetch path, grouped by resolved adapter.""" + if not queries: + return [] + + grouped_queries: dict[int, tuple["SourceAdapter", list[tuple[int, Query]]]] = {} + for idx, query in enumerate(queries): + adapter = self._resolve_source(query.table, source) + adapter_key = id(adapter) + if adapter_key not in grouped_queries: + grouped_queries[adapter_key] = (adapter, []) + grouped_queries[adapter_key][1].append((idx, query)) + + results: list[Optional["FetchResult"]] = [None] * len(queries) + for adapter, indexed_queries in grouped_queries.values(): + batch_queries = [query for _, query in indexed_queries] + batch_results = adapter.fetch_many( + batch_queries, + max_staleness=max_staleness, + ) + if len(batch_results) != len(batch_queries): + raise ValueError( + f"Adapter '{adapter.source_name}' returned {len(batch_results)} " + f"results for {len(batch_queries)} queries." + ) + for (idx, _), result in zip(indexed_queries, batch_results): + results[idx] = result + + if any(result is None for result in results): + raise ValueError("fetch_many() did not populate every requested result.") + return [result for result in results if result is not None] def prefetch( self, dataset: str, *, source: Optional[str] = None, - asof_range: tuple = None, + asof_range: tuple[date, date] | None = None, ) -> "CacheManifest": """Warm cache for a dataset via the resolved adapter.""" adapter = self._resolve_source(dataset, source) diff --git a/alphaforge/data/fred_source.py b/alphaforge/data/fred_source.py index 572a52a..a1318c3 100644 --- a/alphaforge/data/fred_source.py +++ b/alphaforge/data/fred_source.py @@ -9,7 +9,10 @@ class FREDDataSource(DataSource): """ - Data source for fetching data from FRED. + Legacy/raw-loader FRED DataSource kept for compatibility. + + New code should prefer ``alphaforge.data.sources.fred.FREDSourceAdapter`` + via ``DataContext.fetch(...)``. """ name: str = "fred" diff --git a/alphaforge/data/pit_source.py b/alphaforge/data/pit_source.py index b5aaa9f..1268a69 100644 --- a/alphaforge/data/pit_source.py +++ b/alphaforge/data/pit_source.py @@ -15,7 +15,7 @@ @dataclass class PITDataSource(DataSource): - """Expose PIT snapshots/observations through the DataSource contract.""" + """Expose PIT rows through the legacy/raw-loader DataSource contract.""" pit: PITAccessor lag_policy: ReleaseLagPolicy | None = None diff --git a/alphaforge/data/public_web/__init__.py b/alphaforge/data/public_web/__init__.py index 6560846..f301b8f 100644 --- a/alphaforge/data/public_web/__init__.py +++ b/alphaforge/data/public_web/__init__.py @@ -3,6 +3,7 @@ from .bcb_sgs import BCBSGSDataSource from .bea import BEADataSource from .bls import BLSDataSource +from .cftc_cot import CFTCCoTSource, CFTCDisaggregatedCoTSource from .cftc_swaps_weekly import CFTCWeeklySwapsSource from .cme_productslate_reference import CMEProductSlateSource from .destatis_genesis import DestatisGenesisDataSource @@ -24,6 +25,8 @@ __all__ = [ "BLSDataSource", "BEADataSource", + "CFTCCoTSource", + "CFTCDisaggregatedCoTSource", "EIADataSource", "EurostatDataSource", "ECBSDMXDataSource", diff --git a/alphaforge/data/public_web/archive.py b/alphaforge/data/public_web/archive.py new file mode 100644 index 0000000..a5d673b --- /dev/null +++ b/alphaforge/data/public_web/archive.py @@ -0,0 +1,185 @@ +from __future__ import annotations + +import io +import re +import zipfile +from collections.abc import Iterable +from dataclasses import dataclass +from pathlib import Path +from urllib.parse import urljoin, urlparse + + +@dataclass(frozen=True) +class ArchiveFetchPlanEntry: + url: str + artifact_name: str + year: int | None = None + + +def _path_from_url(url: str) -> str: + return urlparse(str(url)).path + + +def _artifact_name_from_url(url: str, fallback: str) -> str: + path = Path(_path_from_url(url)) + return path.name or fallback + + +def _infer_year(url: str) -> int | None: + match = re.search(r"(19|20)\d{2}", _path_from_url(url)) + return int(match.group(0)) if match else None + + +def discover_archive_links( + html: str, + *, + base_url: str, + suffixes: Iterable[str], +) -> list[str]: + allowed = tuple(str(suffix).lower() for suffix in suffixes) + hrefs = re.findall(r'href=["\']([^"\']+)["\']', html, flags=re.IGNORECASE) + urls = [ + urljoin(base_url, href) + for href in hrefs + if _path_from_url(href).lower().endswith(allowed) + ] + return sorted(set(urls)) + + +def filter_urls_for_years(urls: Iterable[str], years: Iterable[int]) -> list[str]: + url_list = list(urls) + year_tokens = {str(year) for year in years} + if not year_tokens: + return url_list + filtered = [ + url for url in url_list if any(token in str(url) for token in year_tokens) + ] + return filtered or url_list + + +def iter_yearly_archive_urls( + *, + start_year: int, + end_year: int, + url_template: str, + first_year: int, + yearly_first_year: int | None = None, + historical_url: str | None = None, + historical_last_year: int | None = None, + file_urls: list[str] | None = None, +) -> list[str]: + if file_urls: + return list(file_urls) + + urls: list[str] = [] + requested_start = max(start_year, first_year) + effective_yearly_first = yearly_first_year or first_year + + if ( + historical_url is not None + and historical_last_year is not None + and requested_start < effective_yearly_first + and end_year >= first_year + ): + urls.append(historical_url) + requested_start = max(requested_start, historical_last_year + 1) + + urls.extend( + url_template.format(year=year) + for year in range(max(requested_start, effective_yearly_first), end_year + 1) + ) + return urls + + +def plan_archive_fetches( + urls: Iterable[str], + *, + years: Iterable[int] | None = None, + fallback_artifact_prefix: str = "archive", +) -> list[ArchiveFetchPlanEntry]: + planned_urls = ( + filter_urls_for_years(urls, years or []) if years is not None else list(urls) + ) + + entries: list[ArchiveFetchPlanEntry] = [] + seen: set[str] = set() + for index, url in enumerate(planned_urls, start=1): + if url in seen: + continue + seen.add(url) + entries.append( + ArchiveFetchPlanEntry( + url=url, + artifact_name=_artifact_name_from_url( + url, f"{fallback_artifact_prefix}_{index}" + ), + year=_infer_year(url), + ) + ) + return entries + + +def discover_archive_fetches( + html: str, + *, + base_url: str, + suffixes: Iterable[str], + years: Iterable[int] | None = None, + fallback_artifact_prefix: str = "archive", +) -> list[ArchiveFetchPlanEntry]: + return plan_archive_fetches( + discover_archive_links(html, base_url=base_url, suffixes=suffixes), + years=years, + fallback_artifact_prefix=fallback_artifact_prefix, + ) + + +def iter_yearly_archive_fetches( + *, + start_year: int, + end_year: int, + url_template: str, + first_year: int, + yearly_first_year: int | None = None, + historical_url: str | None = None, + historical_last_year: int | None = None, + file_urls: list[str] | None = None, + fallback_artifact_prefix: str = "archive", +) -> list[ArchiveFetchPlanEntry]: + return plan_archive_fetches( + iter_yearly_archive_urls( + start_year=start_year, + end_year=end_year, + url_template=url_template, + first_year=first_year, + yearly_first_year=yearly_first_year, + historical_url=historical_url, + historical_last_year=historical_last_year, + file_urls=file_urls, + ), + years=None, + fallback_artifact_prefix=fallback_artifact_prefix, + ) + + +def read_zip_members( + payload: bytes, + *, + suffixes: Iterable[str], +) -> list[tuple[str, bytes]]: + allowed = tuple(str(suffix).lower() for suffix in suffixes) + members: list[tuple[str, bytes]] = [] + with zipfile.ZipFile(io.BytesIO(payload)) as zf: + for name in zf.namelist(): + if str(name).lower().endswith(allowed): + members.append((name, zf.read(name))) + return members + + +def read_first_zip_member( + payload: bytes, + *, + suffixes: Iterable[str], +) -> tuple[str, bytes] | None: + members = read_zip_members(payload, suffixes=suffixes) + return members[0] if members else None diff --git a/alphaforge/data/public_web/b3_historical_quotes.py b/alphaforge/data/public_web/b3_historical_quotes.py index 36d5b4d..9ade937 100644 --- a/alphaforge/data/public_web/b3_historical_quotes.py +++ b/alphaforge/data/public_web/b3_historical_quotes.py @@ -1,22 +1,18 @@ from __future__ import annotations -import io -import re -import zipfile from pathlib import Path -from urllib.parse import urljoin import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .archive import discover_archive_fetches, read_first_zip_member +from .base import PublicWebSourceBase from .http import CachedHttpClient -from .utils import apply_query_filters, project_columns +from .schema_helpers import table_schema -class B3HistoricalQuotesDataSource(DataSource): +class B3HistoricalQuotesDataSource(PublicWebSourceBase): name = "b3_historical_quotes" TABLE = "b3_equity_quotes_daily" @@ -27,13 +23,13 @@ def __init__( cache_dir: str | Path | None = None, page_url: str = "https://www.b3.com.br/en_us/market-data-and-indices/data-services/market-data/historical-data/equities/historical-quote-data/", ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._page_url = page_url - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["open", "high", "low", "close", "volume"], canonical_columns=["open", "high", "low", "close", "volume"], entity_column="ticker", @@ -42,25 +38,23 @@ def schemas(self) -> dict[str, TableSchema]: ) } - def _discover_links(self, q: Query) -> list[str]: + def _discover_links(self, q: Query): payload = self._http.get_bytes( url=self._page_url, source="b3_quotes", artifact_name="landing.html" ) html = payload.decode(errors="ignore") - hrefs = re.findall(r'href=["\']([^"\']+)["\']', html, flags=re.IGNORECASE) - links = [ - urljoin(self._page_url, h) - for h in hrefs - if h.lower().endswith((".zip", ".txt", ".csv")) - ] years = set() if q.start is not None: years.add(q.start.year) if q.end is not None: years.add(q.end.year) - if years: - links = [u for u in links if any(str(y) in u for y in years)] or links - return sorted(set(links)) + return discover_archive_fetches( + html, + base_url=self._page_url, + suffixes=(".zip", ".txt", ".csv"), + years=years, + fallback_artifact_prefix="b3_quotes", + ) @staticmethod def _parse_fixed_width(text: str) -> pd.DataFrame: @@ -93,62 +87,34 @@ def _parse_fixed_width(text: str) -> pd.DataFrame: return pd.DataFrame(rows) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for link in self._discover_links(q): + for planned in self._discover_links(q): payload = self._http.get_bytes( - url=link, + url=planned.url, source="b3_quotes", - artifact_name=Path(link.split("?")[0]).name, + artifact_name=planned.artifact_name, ) text = "" - if link.lower().endswith(".zip"): - with zipfile.ZipFile(io.BytesIO(payload)) as zf: - names = [ - n for n in zf.namelist() if n.lower().endswith((".txt", ".csv")) - ] - if not names: - continue - text = zf.read(names[0]).decode("latin-1", errors="ignore") + if planned.url.lower().split("?", 1)[0].endswith(".zip"): + member = read_first_zip_member(payload, suffixes=(".txt", ".csv")) + if member is None: + continue + _member_name, member_payload = member + text = member_payload.decode("latin-1", errors="ignore") else: text = payload.decode("latin-1", errors="ignore") parsed = self._parse_fixed_width(text) if not parsed.empty: - parsed["asof_utc"] = q.asof or pd.Timestamp.now(tz="UTC") + parsed["asof_utc"] = asof_utc rows.append(parsed) out = ( pd.concat(rows, ignore_index=True) if rows - else pd.DataFrame( - columns=[ - "date", - "ticker", - "asof_utc", - "open", - "high", - "low", - "close", - "volume", - ] - ) - ) - if out.empty: - return out - out = apply_query_filters( - out.rename(columns={"ticker": "entity_id"}), - q=q, - time_col="date", - entity_col="entity_id", - ).rename(columns={"entity_id": "ticker"}) - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="ticker", + else self._empty_frame(schema, entity_col="ticker") ) - return out.sort_values(["ticker", "date"]).reset_index(drop=True) + return self._finalize(out, q=q, schema=schema, entity_col="ticker") diff --git a/alphaforge/data/public_web/base.py b/alphaforge/data/public_web/base.py new file mode 100644 index 0000000..318cf69 --- /dev/null +++ b/alphaforge/data/public_web/base.py @@ -0,0 +1,103 @@ +from __future__ import annotations + +from collections.abc import Callable, Iterable, Sequence +from pathlib import Path + +import pandas as pd + +from alphaforge.data.query import Query +from alphaforge.data.schema import TableSchema + +from .finalize import empty_frame_for_schema, finalize_public_frame, frame_from_records +from .http import CachedHttpClient + + +class PublicWebSourceBase: + name: str + + def __init__( + self, + *, + http_client: CachedHttpClient | None = None, + cache_dir: str | Path | None = None, + now_fn: Callable[[], pd.Timestamp] | None = None, + ) -> None: + self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + self._now_fn = now_fn or (lambda: pd.Timestamp.now(tz="UTC")) + + def _default_table(self) -> str: + table = getattr(self, "TABLE", None) + if not isinstance(table, str) or not table: + raise AttributeError(f"{type(self).__name__} must define TABLE") + return table + + def schemas(self) -> dict[str, TableSchema]: + raise NotImplementedError + + def _schema(self, table: str | None = None) -> TableSchema: + resolved_table = table or self._default_table() + return self.schemas()[resolved_table] + + def _require_table(self, q: Query, expected: str | None = None) -> str: + resolved_expected = expected or self._default_table() + if q.table != resolved_expected: + raise ValueError(f"Unknown table: {q.table}") + return resolved_expected + + def _require_entities(self, q: Query, *, error_message: str) -> list[str]: + entities = [str(entity) for entity in (q.entities or [])] + if not entities: + raise ValueError(error_message) + return entities + + def _now_utc(self) -> pd.Timestamp: + now = pd.Timestamp(self._now_fn()) + if now.tzinfo is None: + return now.tz_localize("UTC") + return now.tz_convert("UTC") + + def _asof_utc(self, q: Query) -> pd.Timestamp: + return q.asof or self._now_utc() + + def _empty_frame( + self, + schema: TableSchema, + *, + time_col: str | None = None, + entity_col: str | None = None, + ) -> pd.DataFrame: + return empty_frame_for_schema(schema, time_col=time_col, entity_col=entity_col) + + def _frame_from_records( + self, + records: Sequence[dict] | Iterable[dict], + *, + schema: TableSchema, + time_col: str | None = None, + entity_col: str | None = None, + ) -> pd.DataFrame: + return frame_from_records( + records, + schema=schema, + time_col=time_col, + entity_col=entity_col, + ) + + def _finalize( + self, + df: pd.DataFrame, + *, + q: Query, + schema: TableSchema, + time_col: str | None = None, + entity_col: str | None = None, + sort_by: Sequence[str] | None = None, + ) -> pd.DataFrame: + return finalize_public_frame( + df, + q=q, + schema=schema, + time_col=time_col, + entity_col=entity_col, + sort_by=sort_by, + ) diff --git a/alphaforge/data/public_web/bcb_sgs.py b/alphaforge/data/public_web/bcb_sgs.py index 7cbae2c..f0b373f 100644 --- a/alphaforge/data/public_web/bcb_sgs.py +++ b/alphaforge/data/public_web/bcb_sgs.py @@ -7,14 +7,14 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .base import PublicWebSourceBase from .http import CachedHttpClient -from .utils import apply_query_filters, make_entity_id, project_columns +from .schema_helpers import single_value_schema +from .utils import make_entity_id -class BCBSGSDataSource(DataSource): +class BCBSGSDataSource(PublicWebSourceBase): name = "bcb_sgs" TABLE = "bcb_sgs_series" @@ -25,19 +25,11 @@ def __init__( cache_dir: str | Path | None = None, base_url: str = "https://api.bcb.gov.br/dados/serie/bcdata.sgs", ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._base_url = base_url.rstrip("/") - def schemas(self) -> dict[str, TableSchema]: - return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - ) - } + def schemas(self): + return {self.TABLE: single_value_schema(self.TABLE)} def _call(self, code: str, q: Query) -> pd.DataFrame: params = {"formato": "csv"} @@ -54,11 +46,13 @@ def _call(self, code: str, q: Query) -> pd.DataFrame: return frame def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - codes = [str(x) for x in (q.entities or [])] - if not codes: - raise ValueError("BCBSGSDataSource requires q.entities with SGS codes") + self._require_table(q) + schema = self._schema() + codes = self._require_entities( + q, + error_message="BCBSGSDataSource requires q.entities with SGS codes", + ) + asof_utc = self._asof_utc(q) rows = [] for code in codes: @@ -77,7 +71,7 @@ def fetch(self, q: Query) -> pd.DataFrame: "date": dates, "entity_id": make_entity_id(code), "value": vals, - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) rows.append(tmp) @@ -85,17 +79,6 @@ def fetch(self, q: Query) -> pd.DataFrame: out = ( pd.concat(rows, ignore_index=True) if rows - else pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - ) - if out.empty: - return out - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", + else self._empty_frame(schema) ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/bea.py b/alphaforge/data/public_web/bea.py index 4d384ff..0dd1ebc 100644 --- a/alphaforge/data/public_web/bea.py +++ b/alphaforge/data/public_web/bea.py @@ -8,15 +8,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .registry_loader import load_registry_entries, map_registry -from .utils import apply_query_filters, project_columns +from .registry_api import RegistryApiSourceBase +from .schema_helpers import single_value_schema -class BEADataSource(DataSource): +class BEADataSource(RegistryApiSourceBase): name = "bea" TABLE = "bea_series" @@ -30,24 +28,18 @@ def __init__( registry_entries: list[dict] | None = None, registry_path: str | Path | None = None, ) -> None: + super().__init__(http_client=http_client, cache_dir=cache_dir) self._api_key = api_key or os.getenv("BEA_API_KEY") self._api_url = api_url - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) - entries = load_registry_entries( - "bea_series.yaml", entries=registry_entries, registry_path=registry_path + self._init_registry( + "bea_series.yaml", + registry_entries=registry_entries, + registry_path=registry_path, ) - self._registry = map_registry(entries) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - native_freq="M", - ) + self.TABLE: single_value_schema(self.TABLE, native_freq="M") } @staticmethod @@ -77,19 +69,17 @@ def _call(self, params: dict) -> dict: return json.loads(payload.decode("utf-8")) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) if not self._api_key: raise ValueError("BEA API key required via BEA_API_KEY or constructor arg") - entities = list(q.entities or []) - if not entities: - raise ValueError("BEADataSource requires q.entities registry keys") + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for entity in entities: - config = self._registry.get(str(entity)) - if config is None: - continue + for entity, config in self._iter_entity_configs( + q, + error_message="BEADataSource requires q.entities registry keys", + ): params = dict(config.get("params", {})) payload = self._call(params) data_rows = payload.get("BEAAPI", {}).get("Results", {}).get("Data", []) @@ -105,20 +95,9 @@ def fetch(self, q: Query) -> pd.DataFrame: str(row.get("DataValue", "")).replace(",", ""), errors="coerce", ), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/bls.py b/alphaforge/data/public_web/bls.py index 29f5d00..bb9aea0 100644 --- a/alphaforge/data/public_web/bls.py +++ b/alphaforge/data/public_web/bls.py @@ -7,13 +7,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource -from .utils import apply_query_filters, make_entity_id, project_columns +from .base import PublicWebSourceBase +from .schema_helpers import single_value_schema +from .utils import make_entity_id -class BLSDataSource(DataSource): +class BLSDataSource(PublicWebSourceBase): name = "bls" TABLE = "bls_series" @@ -25,19 +25,16 @@ def __init__( chunk_size: int = 25, response_provider=None, ) -> None: + super().__init__() self._api_key = api_key or os.getenv("BLS_API_KEY") self._api_url = api_url self._chunk_size = chunk_size self._response_provider = response_provider - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", + self.TABLE: single_value_schema( + self.TABLE, native_freq="M", time_semantics="point", ) @@ -75,11 +72,13 @@ def _month_end(year: int, period: str) -> pd.Timestamp | None: ) + pd.offsets.MonthEnd(0) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - entities = list(q.entities or []) - if not entities: - raise ValueError("BLSDataSource requires q.entities with BLS series ids") + self._require_table(q) + schema = self._schema() + entities = self._require_entities( + q, + error_message="BLSDataSource requires q.entities with BLS series ids", + ) + asof_utc = self._asof_utc(q) start_year = ( q.start.year @@ -105,21 +104,9 @@ def fetch(self, q: Query) -> pd.DataFrame: "date": date, "entity_id": make_entity_id(sid), "value": pd.to_numeric(point.get("value"), errors="coerce"), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/cftc_cot.py b/alphaforge/data/public_web/cftc_cot.py index b9f14e2..53669cf 100644 --- a/alphaforge/data/public_web/cftc_cot.py +++ b/alphaforge/data/public_web/cftc_cot.py @@ -1,23 +1,18 @@ -"""CFTC Commitments of Traders — Traders in Financial Futures (TFF). +"""CFTC Commitments of Traders public-web sources. -Downloads and normalises the disaggregated Traders in Financial Futures -report from the CFTC bulk-file archive. Each year's data comes as a ZIP -containing a single CSV with positions as of Tuesday, published the -following Friday. +Includes: + +- ``CFTCCoTSource`` for Traders in Financial Futures (TFF; futures only) +- ``CFTCDisaggregatedCoTSource`` for disaggregated commodity futures Canonical entity-id pattern:: futures.{contract_code}.{trader_category}.cftc - -Example entity IDs:: - - futures.vix.lev_money.cftc - futures.vix.dealer.cftc - futures.sp500.asset_mgr.cftc """ from __future__ import annotations +import io import re from pathlib import Path from typing import Sequence @@ -28,6 +23,11 @@ from alphaforge.data.schema import TableSchema from alphaforge.data.source import DataSource +from .archive import ( + iter_yearly_archive_fetches, + iter_yearly_archive_urls, + read_zip_members, +) from .http import CachedHttpClient from .parsing import normalize_headers from .utils import ( @@ -38,43 +38,59 @@ to_float, ) -# ── Column name mapping (snake_case after normalisation) ─────────────── +_TraderCategorySpec = tuple[str, str, str | None, str | None, str | None, str] + +# Common column names after header normalisation _COL_REPORT_DATE = "report_date_as_yyyy_mm_dd" +_COL_REPORT_DATE_LEGACY = "report_date_as_mm_dd_yyyy" _COL_MARKET = "market_and_exchange_names" _COL_CFTC_CODE = "cftc_contract_market_code" +_COL_OI = "open_interest_all" -# Leveraged-money columns +# TFF columns _COL_LEV_LONG = "lev_money_positions_long_all" _COL_LEV_SHORT = "lev_money_positions_short_all" _COL_LEV_SPREAD = "lev_money_positions_spread_all" _COL_CHG_LEV_LONG = "change_in_lev_money_long_all" _COL_CHG_LEV_SHORT = "change_in_lev_money_short_all" -# Dealer columns _COL_DEALER_LONG = "dealer_positions_long_all" _COL_DEALER_SHORT = "dealer_positions_short_all" _COL_DEALER_SPREAD = "dealer_positions_spread_all" _COL_CHG_DEALER_LONG = "change_in_dealer_long_all" _COL_CHG_DEALER_SHORT = "change_in_dealer_short_all" -# Asset-manager columns _COL_AM_LONG = "asset_mgr_positions_long_all" _COL_AM_SHORT = "asset_mgr_positions_short_all" _COL_AM_SPREAD = "asset_mgr_positions_spread_all" _COL_CHG_AM_LONG = "change_in_asset_mgr_long_all" _COL_CHG_AM_SHORT = "change_in_asset_mgr_short_all" -# Other-reportable columns _COL_OTHER_LONG = "other_rept_positions_long_all" _COL_OTHER_SHORT = "other_rept_positions_short_all" _COL_OTHER_SPREAD = "other_rept_positions_spread_all" _COL_CHG_OTHER_LONG = "change_in_other_rept_long_all" _COL_CHG_OTHER_SHORT = "change_in_other_rept_short_all" -_COL_OI = "open_interest_all" - -# Each trader category → (long, short, spread, chg_long, chg_short, label) -_TRADER_CATEGORIES: dict[str, tuple[str, str, str, str, str, str]] = { +# Disaggregated columns +_COL_PROD_MERC_LONG = "prod_merc_positions_long_all" +_COL_PROD_MERC_SHORT = "prod_merc_positions_short_all" +_COL_CHG_PROD_MERC_LONG = "change_in_prod_merc_long_all" +_COL_CHG_PROD_MERC_SHORT = "change_in_prod_merc_short_all" + +_COL_SWAP_LONG = "swap_positions_long_all" +_COL_SWAP_SHORT = "swap_positions_short_all" +_COL_SWAP_SPREAD = "swap_positions_spread_all" +_COL_CHG_SWAP_LONG = "change_in_swap_long_all" +_COL_CHG_SWAP_SHORT = "change_in_swap_short_all" + +_COL_M_MONEY_LONG = "m_money_positions_long_all" +_COL_M_MONEY_SHORT = "m_money_positions_short_all" +_COL_M_MONEY_SPREAD = "m_money_positions_spread_all" +_COL_CHG_M_MONEY_LONG = "change_in_m_money_long_all" +_COL_CHG_M_MONEY_SHORT = "change_in_m_money_short_all" + +_TFF_TRADER_CATEGORIES: dict[str, _TraderCategorySpec] = { "lev_money": ( _COL_LEV_LONG, _COL_LEV_SHORT, @@ -109,35 +125,66 @@ ), } -# Well-known CFTC contract codes → short name. -_CONTRACT_CODES: dict[str, str] = { +_DISAGG_TRADER_CATEGORIES: dict[str, _TraderCategorySpec] = { + "prod_merc": ( + _COL_PROD_MERC_LONG, + _COL_PROD_MERC_SHORT, + None, + _COL_CHG_PROD_MERC_LONG, + _COL_CHG_PROD_MERC_SHORT, + "prod_merc", + ), + "swap": ( + _COL_SWAP_LONG, + _COL_SWAP_SHORT, + _COL_SWAP_SPREAD, + _COL_CHG_SWAP_LONG, + _COL_CHG_SWAP_SHORT, + "swap", + ), + "m_money": ( + _COL_M_MONEY_LONG, + _COL_M_MONEY_SHORT, + _COL_M_MONEY_SPREAD, + _COL_CHG_M_MONEY_LONG, + _COL_CHG_M_MONEY_SHORT, + "m_money", + ), + "other_rept": ( + _COL_OTHER_LONG, + _COL_OTHER_SHORT, + _COL_OTHER_SPREAD, + _COL_CHG_OTHER_LONG, + _COL_CHG_OTHER_SHORT, + "other_rept", + ), +} + +_TFF_CONTRACT_CODES: dict[str, str] = { "1170E1": "vix", "13874+": "sp500", "13874A": "sp500_e_mini", "33874E": "sp500_micro", - "209742": "vix", # VIX futures (alternative code) - # G10 FX futures (CME) - "099741": "eur", # Euro FX - "096742": "gbp", # British Pound - "097741": "jpy", # Japanese Yen - "092741": "chf", # Swiss Franc - "090741": "cad", # Canadian Dollar - "232741": "aud", # Australian Dollar - "112741": "nzd", # New Zealand Dollar - "095741": "mxn", # Mexican Peso (not G10 but heavily traded) - "089741": "sek", # Swedish Krona - "088741": "nok", # Norwegian Krone - # US rates futures - "13874P": "sofr_3m", # Three-Month SOFR (CME) - "134741": "ust_10y", # 10-Year T-Note - "020601": "ust_30y", # T-Bond (30Y) - "044601": "ust_5y", # 5-Year T-Note - "042601": "ust_2y", # 2-Year T-Note - "043602": "fed_funds", # 30-Day Federal Funds + "209742": "vix", + "099741": "eur", + "096742": "gbp", + "097741": "jpy", + "092741": "chf", + "090741": "cad", + "232741": "aud", + "112741": "nzd", + "095741": "mxn", + "089741": "sek", + "088741": "nok", + "13874P": "sofr_3m", + "134741": "ust_10y", + "020601": "ust_30y", + "044601": "ust_5y", + "042601": "ust_2y", + "043602": "fed_funds", } -# Regex fallbacks for market-name-based contract detection. -_MARKET_NAME_PATTERNS: list[tuple[re.Pattern[str], str]] = [ +_TFF_MARKET_NAME_PATTERNS: list[tuple[re.Pattern[str], str]] = [ (re.compile(r"\bVIX\b", re.IGNORECASE), "vix"), (re.compile(r"\bCBOE VOLATILITY INDEX\b", re.IGNORECASE), "vix"), (re.compile(r"\bS&P 500\b", re.IGNORECASE), "sp500"), @@ -152,39 +199,99 @@ (re.compile(r"\b10.YEAR\b.*\bT.NOTE\b", re.IGNORECASE), "ust_10y"), ] +_DISAGG_MARKET_NAME_PATTERNS: list[tuple[re.Pattern[str], str]] = [ + (re.compile(r"\bWHEAT[- ]SRW\b", re.IGNORECASE), "wheat_srw"), + (re.compile(r"\bWHEAT[- ]HRW\b", re.IGNORECASE), "wheat_hrw"), + (re.compile(r"\bCORN\b", re.IGNORECASE), "corn"), + (re.compile(r"\bSOYBEAN(S)?\b", re.IGNORECASE), "soybeans"), + (re.compile(r"\bSOYBEAN OIL\b", re.IGNORECASE), "soybean_oil"), + (re.compile(r"\bSOYBEAN MEAL\b", re.IGNORECASE), "soybean_meal"), + (re.compile(r"\bLIGHT SWEET CRUDE OIL\b|\bWTI\b", re.IGNORECASE), "wti"), + (re.compile(r"\bBRENT CRUDE OIL\b", re.IGNORECASE), "brent"), + (re.compile(r"\bNATURAL GAS\b", re.IGNORECASE), "natgas"), + (re.compile(r"\bGOLD\b", re.IGNORECASE), "gold"), + (re.compile(r"\bSILVER\b", re.IGNORECASE), "silver"), + (re.compile(r"\bCOPPER\b", re.IGNORECASE), "copper"), + (re.compile(r"\bCOFFEE\b", re.IGNORECASE), "coffee"), + (re.compile(r"\bSUGAR\b", re.IGNORECASE), "sugar"), + (re.compile(r"\bCOTTON\b", re.IGNORECASE), "cotton"), +] + -def _infer_contract_code(cftc_code: str, market_name: str) -> str: - """Map CFTC contract code or market name to a short identifier.""" - code = cftc_code.strip() - if code in _CONTRACT_CODES: - return _CONTRACT_CODES[code] - for pat, name in _MARKET_NAME_PATTERNS: +def _slugify_contract_name(market_name: str) -> str: + """Turn the human-readable market name into a stable identifier.""" + primary = (market_name or "").split(" - ", 1)[0].strip() + slug = re.sub(r"[^a-z0-9]+", "_", primary.lower()).strip("_") + slug = re.sub(r"_+", "_", slug) + return slug or "unknown" + + +def _infer_contract_code( + cftc_code: str, + market_name: str, + *, + contract_codes: dict[str, str], + market_name_patterns: Sequence[tuple[re.Pattern[str], str]], + prefer_market_slug: bool = False, +) -> str: + """Map a CFTC contract code or market name to a short identifier.""" + code = str(cftc_code).strip() + if code in contract_codes: + return contract_codes[code] + for pat, name in market_name_patterns: if pat.search(market_name or ""): return name - # Fallback: use the raw CFTC code, cleaned. + if prefer_market_slug: + market_slug = _slugify_contract_name(market_name) + if market_slug != "unknown": + return market_slug return re.sub(r"[^a-z0-9]", "", code.lower()) or "unknown" def _publication_date(report_date: pd.Series) -> pd.Series: - """Compute publication date (Friday) from report date (Tuesday). - - The CFTC publishes COT data on Friday afternoon for positions reported - as of the preceding Tuesday — a 3 *business-day* lag. We use - ``pd.tseries.offsets.BDay(3)`` so that Tuesday → Friday even across - holidays (it simply rolls forward through any intervening non-business - days). - """ + """Compute publication date (Friday) from report date (Tuesday).""" offset = pd.tseries.offsets.BDay(3) return report_date.map(lambda d: d + offset if pd.notna(d) else pd.NaT) -class CFTCCoTSource(DataSource): - """CFTC Commitments of Traders — Traders in Financial Futures (TFF).""" +def _parse_yymmdd_dates(series: pd.Series) -> pd.Series: + """Parse CFTC YYMMDD date fields reliably, preserving leading zeros.""" + text = ( + series.astype(str) + .str.strip() + .str.replace(r"\.0$", "", regex=True) + .str.replace(r"[^0-9]", "", regex=True) + .str.zfill(6) + ) + parsed = pd.to_datetime(text, format="%y%m%d", errors="coerce", utc=True) + normalized = parsed.map(lambda ts: ts.normalize() if pd.notna(ts) else pd.NaT) + return pd.to_datetime(normalized, utc=True) + + +def _parse_report_dates(frame: pd.DataFrame) -> pd.Series: + """Parse the report-date column across CFTC's header variants.""" + for candidate in (_COL_REPORT_DATE, _COL_REPORT_DATE_LEGACY, "report_date"): + if candidate in frame.columns: + return ensure_date_utc(frame[candidate]) + if "as_of_date_in_form_yymmdd" in frame.columns: + return _parse_yymmdd_dates(frame["as_of_date_in_form_yymmdd"]) + return pd.to_datetime(pd.Series([pd.NaT] * len(frame), index=frame.index), utc=True) + + +class _BaseCFTCCoTSource(DataSource): + """Shared implementation for CFTC CoT ZIP archives.""" name: str = "cftc_cot" - TABLE = "cftc.cot.tff" - URL_TEMPLATE = "https://www.cftc.gov/files/dea/history/fut_fin_txt_{year}.zip" + TABLE = "cftc.cot" + URL_TEMPLATE = "" FIRST_YEAR = 2006 + YEARLY_FIRST_YEAR = 2006 + HISTORICAL_URL: str | None = None + HISTORICAL_LAST_YEAR: int | None = None + TRADER_CATEGORIES: dict[str, _TraderCategorySpec] = {} + CONTRACT_CODES: dict[str, str] = {} + MARKET_NAME_PATTERNS: Sequence[tuple[re.Pattern[str], str]] = () + PREFER_MARKET_SLUG_FALLBACK = False def __init__( self, @@ -199,10 +306,9 @@ def __init__( self._url_template = url_template or self.URL_TEMPLATE self._file_urls = file_urls self._trader_categories = ( - list(trader_categories) if trader_categories else list(_TRADER_CATEGORIES) + list(trader_categories) if trader_categories else list(self.TRADER_CATEGORIES) ) - # ── schema ────────────────────────────────────────────────────────── def schemas(self) -> dict[str, TableSchema]: return { self.TABLE: TableSchema( @@ -229,57 +335,52 @@ def schemas(self) -> dict[str, TableSchema]: ) } - # ── internal helpers ──────────────────────────────────────────────── def _year_urls(self, start_year: int, end_year: int) -> list[str]: - """Build download URLs for the requested year range.""" - if self._file_urls: - return list(self._file_urls) - return [ - self._url_template.format(year=y) - for y in range(max(start_year, self.FIRST_YEAR), end_year + 1) - ] + return iter_yearly_archive_urls( + start_year=start_year, + end_year=end_year, + url_template=self._url_template, + first_year=self.FIRST_YEAR, + yearly_first_year=self.YEARLY_FIRST_YEAR, + historical_url=self.HISTORICAL_URL, + historical_last_year=self.HISTORICAL_LAST_YEAR, + file_urls=self._file_urls, + ) + + def _year_fetches(self, start_year: int, end_year: int): + return iter_yearly_archive_fetches( + start_year=start_year, + end_year=end_year, + url_template=self._url_template, + first_year=self.FIRST_YEAR, + yearly_first_year=self.YEARLY_FIRST_YEAR, + historical_url=self.HISTORICAL_URL, + historical_last_year=self.HISTORICAL_LAST_YEAR, + file_urls=self._file_urls, + fallback_artifact_prefix=self.name, + ) - def _read_zip(self, url: str) -> pd.DataFrame: + def _read_zip(self, planned) -> pd.DataFrame: payload = self._http.get_bytes( - url=url, - source="cftc_cot", - artifact_name=Path(url).stem + ".zip", + url=planned.url, + source=self.name, + artifact_name=planned.artifact_name, ) - # CFTC TFF ZIPs contain .txt files (CSV-formatted), not .csv, - # so parse_zip_csv_bytes (which filters for .csv) won't work. - # Read the first CSV-like file (.csv or .txt) from the archive. - import io - import zipfile frames: list[pd.DataFrame] = [] - with zipfile.ZipFile(io.BytesIO(payload)) as zf: - for name in zf.namelist(): - ext = name.rsplit(".", 1)[-1].lower() if "." in name else "" - if ext in ("csv", "txt"): - with zf.open(name) as fh: - frame = pd.read_csv(fh) - frames.append(normalize_headers(frame)) + for _name, member_payload in read_zip_members(payload, suffixes=(".csv", ".txt")): + frames.append(normalize_headers(pd.read_csv(io.BytesIO(member_payload)))) if not frames: return pd.DataFrame() return pd.concat(frames, ignore_index=True) def _to_long(self, frame: pd.DataFrame) -> pd.DataFrame: - """Melt wide TFF CSV into long format with one row per (date, entity).""" if frame.empty: return pd.DataFrame() - # Report date - date_col = None - for candidate in [_COL_REPORT_DATE, "report_date", "as_of_date_in_form_yymmdd"]: - if candidate in frame.columns: - date_col = candidate - break - if date_col is None: + report_dates = _parse_report_dates(frame) + if report_dates.isna().all(): return pd.DataFrame() - - report_dates = ensure_date_utc(frame[date_col]) - - # Contract identification cftc_codes = ( frame[_COL_CFTC_CODE].astype(str) if _COL_CFTC_CODE in frame.columns @@ -291,11 +392,16 @@ def _to_long(self, frame: pd.DataFrame) -> pd.DataFrame: else pd.Series("", index=frame.index) ) contract_codes = [ - _infer_contract_code(code, name) + _infer_contract_code( + code, + name, + contract_codes=self.CONTRACT_CODES, + market_name_patterns=self.MARKET_NAME_PATTERNS, + prefer_market_slug=self.PREFER_MARKET_SLUG_FALLBACK, + ) for code, name in zip(cftc_codes, market_names) ] - # Open interest (shared across categories) oi = ( to_float(frame[_COL_OI]) if _COL_OI in frame.columns @@ -304,13 +410,12 @@ def _to_long(self, frame: pd.DataFrame) -> pd.DataFrame: rows: list[pd.DataFrame] = [] for cat_key in self._trader_categories: - if cat_key not in _TRADER_CATEGORIES: + if cat_key not in self.TRADER_CATEGORIES: continue + col_long, col_short, col_spread, col_chg_long, col_chg_short, label = ( - _TRADER_CATEGORIES[cat_key] + self.TRADER_CATEGORIES[cat_key] ) - - # Skip category if columns are missing from the CSV if col_long not in frame.columns or col_short not in frame.columns: continue @@ -323,28 +428,26 @@ def _to_long(self, frame: pd.DataFrame) -> pd.DataFrame: "short_positions": to_float(frame[col_short]), "spread_positions": ( to_float(frame[col_spread]) - if col_spread in frame.columns + if col_spread is not None and col_spread in frame.columns else float("nan") ), "open_interest": oi, "change_long": ( to_float(frame[col_chg_long]) - if col_chg_long in frame.columns + if col_chg_long is not None and col_chg_long in frame.columns else float("nan") ), "change_short": ( to_float(frame[col_chg_short]) - if col_chg_short in frame.columns + if col_chg_short is not None and col_chg_short in frame.columns else float("nan") ), } ) - chunk["entity_id"] = [ - make_entity_id("futures", cc, label, "cftc") - for cc in chunk["contract_code"] + make_entity_id("futures", contract_code, label, "cftc") + for contract_code in chunk["contract_code"] ] - rows.append(chunk) if not rows: @@ -352,28 +455,27 @@ def _to_long(self, frame: pd.DataFrame) -> pd.DataFrame: out = pd.concat(rows, ignore_index=True) out["publication_date"] = _publication_date(out["report_date"]) - # Use publication_date as the primary time column for alignment out["date"] = out["publication_date"] out["asof_utc"] = out["publication_date"] return out - # ── public API ────────────────────────────────────────────────────── def fetch(self, q: Query) -> pd.DataFrame: if q.table != self.TABLE: raise ValueError(f"Unknown table: {q.table}") schema = self.schemas()[self.TABLE] - - # Determine year range from query start_year = q.start.year if q.start is not None else self.FIRST_YEAR end_year = q.end.year if q.end is not None else pd.Timestamp.now(tz="UTC").year frames: list[pd.DataFrame] = [] - for url in self._year_urls(start_year, end_year): + for planned in self._year_fetches(start_year, end_year): try: - raw = self._read_zip(url) - except Exception: - continue + raw = self._read_zip(planned) + except Exception as exc: + raise RuntimeError( + f"Failed to load CFTC archive {planned.url} " + f"for source '{self.name}'." + ) from exc long = self._to_long(raw) if not long.empty: frames.append(long) @@ -389,6 +491,7 @@ def fetch(self, q: Query) -> pd.DataFrame: ) out = pd.concat(frames, ignore_index=True) + out = out.drop_duplicates(ignore_index=True) out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") out = project_columns( out, @@ -398,3 +501,33 @@ def fetch(self, q: Query) -> pd.DataFrame: entity_col=schema.entity_column, ) return out.reset_index(drop=True) + + +class CFTCCoTSource(_BaseCFTCCoTSource): + """CFTC Commitments of Traders: Traders in Financial Futures (futures only).""" + + name = "cftc_cot" + TABLE = "cftc.cot.tff" + URL_TEMPLATE = "https://www.cftc.gov/files/dea/history/fut_fin_txt_{year}.zip" + FIRST_YEAR = 2006 + YEARLY_FIRST_YEAR = 2010 + HISTORICAL_URL = "https://www.cftc.gov/files/dea/history/fin_fut_txt_2006_2016.zip" + HISTORICAL_LAST_YEAR = 2016 + TRADER_CATEGORIES = _TFF_TRADER_CATEGORIES + CONTRACT_CODES = _TFF_CONTRACT_CODES + MARKET_NAME_PATTERNS = _TFF_MARKET_NAME_PATTERNS + + +class CFTCDisaggregatedCoTSource(_BaseCFTCCoTSource): + """CFTC Commitments of Traders: disaggregated commodity futures.""" + + name = "cftc_cot_disagg" + TABLE = "cftc.cot.disagg" + URL_TEMPLATE = "https://www.cftc.gov/files/dea/history/fut_disagg_txt_{year}.zip" + FIRST_YEAR = 2006 + YEARLY_FIRST_YEAR = 2010 + HISTORICAL_URL = "https://www.cftc.gov/files/dea/history/fut_disagg_txt_hist_2006_2016.zip" + HISTORICAL_LAST_YEAR = 2016 + TRADER_CATEGORIES = _DISAGG_TRADER_CATEGORIES + MARKET_NAME_PATTERNS = _DISAGG_MARKET_NAME_PATTERNS + PREFER_MARKET_SLUG_FALLBACK = True diff --git a/alphaforge/data/public_web/cftc_swaps_weekly.py b/alphaforge/data/public_web/cftc_swaps_weekly.py index fb4e41b..7072368 100644 --- a/alphaforge/data/public_web/cftc_swaps_weekly.py +++ b/alphaforge/data/public_web/cftc_swaps_weekly.py @@ -1,28 +1,29 @@ from __future__ import annotations -import re from pathlib import Path -from urllib.parse import urljoin import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .archive import ( + discover_archive_fetches, + plan_archive_fetches, + read_first_zip_member, +) +from .base import PublicWebSourceBase from .http import CachedHttpClient from .parsing import parse_csv_bytes, parse_xlsx_bytes +from .schema_helpers import table_schema from .utils import ( - apply_query_filters, ensure_date_utc, first_existing, make_entity_id, - project_columns, to_float, ) -class CFTCWeeklySwapsSource(DataSource): +class CFTCWeeklySwapsSource(PublicWebSourceBase): name: str = "cftc_swaps_weekly" TABLE = "cftc.swaps.weekly" ARCHIVE_URL = "https://www.cftc.gov/MarketReports/SwapsReports/Archive/index.htm" @@ -35,14 +36,14 @@ def __init__( archive_url: str | None = None, file_urls: list[str] | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._archive_url = archive_url or self.ARCHIVE_URL self._file_urls = file_urls - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["value"], canonical_columns=[ "value", @@ -52,16 +53,18 @@ def schemas(self) -> dict[str, TableSchema]: "maturity_bucket", "participant_type", ], - entity_column="entity_id", - time_column="date", native_freq="W", time_semantics="interval_end", ) } - def _discover_file_urls(self) -> list[str]: + def _discover_file_urls(self, q: Query): if self._file_urls: - return self._file_urls + return plan_archive_fetches( + self._file_urls, + years=None, + fallback_artifact_prefix="cftc_swaps_weekly", + ) payload = self._http.get_bytes( url=self._archive_url, @@ -69,20 +72,35 @@ def _discover_file_urls(self) -> list[str]: artifact_name="archive_index.html", ) html = payload.decode(errors="ignore") - hrefs = re.findall(r'href=["\']([^"\']+)["\']', html, flags=re.IGNORECASE) - files: list[str] = [] - for href in hrefs: - if href.lower().endswith((".xlsx", ".xls", ".csv")): - files.append(urljoin(self._archive_url, href)) - return sorted(set(files)) - - def _read_file(self, url: str) -> pd.DataFrame: - ext = url.lower().split("?")[0] + years: set[int] = set() + if q.start is not None: + years.add(q.start.year) + if q.end is not None: + years.add(q.end.year) + return discover_archive_fetches( + html, + base_url=self._archive_url, + suffixes=(".xlsx", ".xls", ".csv", ".zip"), + years=years, + fallback_artifact_prefix="cftc_swaps_weekly", + ) + + def _read_file(self, planned) -> pd.DataFrame: + ext = planned.url.lower().split("?", 1)[0] payload = self._http.get_bytes( - url=url, + url=planned.url, source="cftc_swaps_weekly", - artifact_name=Path(ext).name or "weekly_file", + artifact_name=planned.artifact_name, ) + if ext.endswith(".zip"): + member = read_first_zip_member(payload, suffixes=(".csv", ".xlsx", ".xls")) + if member is None: + return pd.DataFrame() + member_name, member_payload = member + member_ext = member_name.lower() + if member_ext.endswith(".csv"): + return parse_csv_bytes(member_payload) + return parse_xlsx_bytes(member_payload) if ext.endswith(".csv"): return parse_csv_bytes(payload) return parse_xlsx_bytes(payload) @@ -154,34 +172,18 @@ def _to_long(self, frame: pd.DataFrame, report_name: str) -> pd.DataFrame: return out def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) frames: list[pd.DataFrame] = [] - for url in self._discover_file_urls(): - parsed = self._read_file(url) - long_df = self._to_long(parsed, Path(url).name) + for planned in self._discover_file_urls(q): + parsed = self._read_file(planned) + long_df = self._to_long(parsed, planned.artifact_name) if not long_df.empty: frames.append(long_df) - schema = self.schemas()[self.TABLE] + schema = self._schema() if not frames: - return pd.DataFrame( - columns=[ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - ) + return self._empty_frame(schema) out = pd.concat(frames, ignore_index=True) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + return self._finalize(out, q=q, schema=schema, sort_by=[]) diff --git a/alphaforge/data/public_web/cme_productslate_reference.py b/alphaforge/data/public_web/cme_productslate_reference.py index dee920b..5b3e400 100644 --- a/alphaforge/data/public_web/cme_productslate_reference.py +++ b/alphaforge/data/public_web/cme_productslate_reference.py @@ -5,20 +5,17 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .parsing import parse_csv_bytes -from .utils import ( - apply_query_filters, - first_existing, - make_entity_id, - project_columns, +from .schema_helpers import table_schema +from .tabular import ( + TabularDocumentSourceBase, + resolved_text_series, ) +from .utils import first_existing, make_entity_id -class CMEProductSlateSource(DataSource): +class CMEProductSlateSource(TabularDocumentSourceBase): name: str = "cme_productslate" TABLE = "cme.productslate.reference" @@ -31,13 +28,13 @@ def __init__( cache_dir: str | Path | None = None, csv_url: str | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._csv_url = csv_url or self.URL - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=[ "exchange", "product_code", @@ -55,20 +52,17 @@ def schemas(self) -> dict[str, TableSchema]: "clearing_code", "mic", ], - entity_column="entity_id", - time_column="date", native_freq="D", time_semantics="point", ) } def _download(self) -> pd.DataFrame: - payload = self._http.get_bytes( + return self._read_csv_frame( url=self._csv_url, source="cme_productslate", artifact_name="productslate.csv", ) - return parse_csv_bytes(payload) def _build_entity_id(self, row: pd.Series) -> str: product_code = str(row.get("product_code") or "unk").lower() @@ -80,25 +74,27 @@ def _build_entity_id(self, row: pd.Series) -> str: return make_entity_id(domain, instrument, product_code, "cme") def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) - schema = self.schemas()[q.table] + schema = self._schema() raw = self._download() - raw_columns = {c.lower(): c for c in raw.columns} out = pd.DataFrame() - out["exchange"] = raw[ - first_existing(raw, "exchange") or raw_columns.get("exchange", "exchange") - ] - out["product_code"] = raw[ - first_existing(raw, "product_code", "product", "code") - or raw_columns.get("product_code", "product_code") - ].astype(str) - out["product_name"] = raw[ - first_existing(raw, "product_name", "name", "description") - or raw_columns.get("product_name", "product_name") - ].astype(str) + out["exchange"] = resolved_text_series( + raw, + ["exchange"], + default="", + ) + out["product_code"] = resolved_text_series( + raw, + ["product_code", "product", "code"], + default="", + ) + out["product_name"] = resolved_text_series( + raw, + ["product_name", "name", "description"], + default="", + ) asset_col = first_existing(raw, "asset_class", "assetclass", "asset") sub_asset_col = first_existing( @@ -111,17 +107,9 @@ def fetch(self, q: Query) -> pd.DataFrame: src_col = first_existing(raw, optional_col) out[optional_col] = raw[src_col].astype(str) if src_col else "" - now_utc = pd.Timestamp.now(tz="UTC") + now_utc = self._asof_utc(q) out["date"] = now_utc.normalize() out["asof_utc"] = now_utc out["entity_id"] = out.apply(self._build_entity_id, axis=1) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + return self._finalize(out, q=q, schema=schema, sort_by=[]) diff --git a/alphaforge/data/public_web/destatis_genesis.py b/alphaforge/data/public_web/destatis_genesis.py index fe98f19..94b9c89 100644 --- a/alphaforge/data/public_web/destatis_genesis.py +++ b/alphaforge/data/public_web/destatis_genesis.py @@ -8,15 +8,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .registry_loader import load_registry_entries, map_registry -from .utils import apply_query_filters, project_columns +from .registry_api import RegistryApiSourceBase +from .schema_helpers import single_value_schema -class DestatisGenesisDataSource(DataSource): +class DestatisGenesisDataSource(RegistryApiSourceBase): name = "destatis_genesis" TABLE = "destatis_series" @@ -32,28 +30,20 @@ def __init__( registry_entries: list[dict] | None = None, registry_path: str | Path | None = None, ) -> None: + super().__init__(http_client=http_client, cache_dir=cache_dir) self._user = user or os.getenv("DESTATIS_GENESIS_USER") self._password = password or os.getenv("DESTATIS_GENESIS_PASS") self._api_key = api_key or os.getenv("DESTATIS_GENESIS_KEY") - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) self._base_url = base_url.rstrip("/") - entries = load_registry_entries( + self._init_registry( "destatis_series.yaml", - entries=registry_entries, + registry_entries=registry_entries, registry_path=registry_path, ) - self._registry = map_registry(entries) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - native_freq="M", - ) + self.TABLE: single_value_schema(self.TABLE, native_freq="M") } @staticmethod @@ -85,19 +75,15 @@ def _call(self, cfg: dict) -> dict: return json.loads(payload.decode("utf-8", errors="ignore")) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - entities = list(q.entities or []) - if not entities: - raise ValueError( - "DestatisGenesisDataSource requires q.entities registry keys" - ) + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for entity in entities: - cfg = self._registry.get(str(entity)) - if cfg is None: - continue + for entity, cfg in self._iter_entity_configs( + q, + error_message="DestatisGenesisDataSource requires q.entities registry keys", + ): payload = self._call(cfg) values = ( payload.get("Object", {}).get("Value", []) @@ -117,20 +103,9 @@ def fetch(self, q: Query) -> pd.DataFrame: "date": date, "entity_id": str(entity), "value": pd.to_numeric(value, errors="coerce"), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/dtcc_ppd.py b/alphaforge/data/public_web/dtcc_ppd.py index b80ff4a..4ec50db 100644 --- a/alphaforge/data/public_web/dtcc_ppd.py +++ b/alphaforge/data/public_web/dtcc_ppd.py @@ -8,25 +8,23 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .base import PublicWebSourceBase from .http import CachedHttpClient from .parsing import parse_zip_csv_bytes +from .schema_helpers import event_table_schema, table_schema from .utils import ( - apply_query_filters, bucket_tenor, coalesce_columns, ensure_date_utc, ensure_utc, make_entity_id, normalize_ccy, - project_columns, to_float, ) -class DTCCPPDSource(DataSource): +class DTCCPPDSource(PublicWebSourceBase): name: str = "dtcc_ppd" EVENTS_TABLE = "dtcc.ppd.events" @@ -45,7 +43,11 @@ def __init__( artifact_provider: Callable[[str], bytes] | None = None, now_fn: Callable[[], pd.Timestamp] | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__( + http_client=http_client, + cache_dir=cache_dir, + now_fn=now_fn, + ) self._api_base_url = api_base_url.rstrip("/") self._jurisdiction = jurisdiction.upper() self._jurisdiction_lower = jurisdiction.lower() @@ -53,12 +55,11 @@ def __init__( self._asset_codes = asset_codes self._list_provider = list_provider self._artifact_provider = artifact_provider - self._now_fn = now_fn or (lambda: pd.Timestamp.now(tz="UTC")) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.EVENTS_TABLE: TableSchema( - name=self.EVENTS_TABLE, + self.EVENTS_TABLE: event_table_schema( + self.EVENTS_TABLE, required_columns=[ "asset_class", "product", @@ -83,15 +84,12 @@ def schemas(self) -> dict[str, TableSchema]: "maturity_date", "reported_at_utc", ], - entity_column="entity_id", - time_column="ts_utc", native_freq="D", time_semantics="point", - event_time_column="ts_utc", release_time_column="asof_utc", ), - self.DAILY_TABLE: TableSchema( - name=self.DAILY_TABLE, + self.DAILY_TABLE: table_schema( + self.DAILY_TABLE, required_columns=[ "trade_count", "notional_sum", @@ -113,8 +111,6 @@ def schemas(self) -> dict[str, TableSchema]: "price_p90", "notional_p90", ], - entity_column="entity_id", - time_column="date", native_freq="D", time_semantics="interval_end", ), @@ -321,7 +317,7 @@ def _normalize_events(self, raw: pd.DataFrame) -> pd.DataFrame: ) out["cleared"] = cleared.isin({"1", "true", "yes", "y"}) - out["asof_utc"] = out["reported_at_utc"].fillna(self._now_fn()) + out["asof_utc"] = out["reported_at_utc"].fillna(self._now_utc()) asset_norm = ( out["asset_class"] @@ -391,42 +387,33 @@ def fetch(self, q: Query) -> pd.DataFrame: if q.table not in {self.EVENTS_TABLE, self.DAILY_TABLE}: raise ValueError(f"Unknown table: {q.table}") - start = q.start or (self._now_fn().normalize() - pd.Timedelta(days=7)) - end = q.end or self._now_fn() + now_utc = self._now_utc() + start = q.start or (now_utc.normalize() - pd.Timedelta(days=7)) + end = q.end or now_utc events = self._load_events_for_range(start=start, end=end, table=q.table) + schema = self._schema(q.table) if events.empty: - schema = self.schemas()[q.table] - cols = [ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - return pd.DataFrame(columns=cols) - - if q.table == self.EVENTS_TABLE: - schema = self.schemas()[self.EVENTS_TABLE] - out = apply_query_filters( - events, q=q, time_col="ts_utc", entity_col="entity_id" - ) - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, + return self._empty_frame( + schema, time_col=schema.time_column, entity_col=schema.entity_column, ) - return out.reset_index(drop=True) - schema = self.schemas()[self.DAILY_TABLE] + if q.table == self.EVENTS_TABLE: + return self._finalize( + events, + q=q, + schema=schema, + time_col="ts_utc", + entity_col="entity_id", + sort_by=[], + ) + daily = self._aggregate_daily(events) - daily = apply_query_filters(daily, q=q, time_col="date", entity_col="entity_id") - daily = project_columns( + return self._finalize( daily, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, + q=q, + schema=schema, + sort_by=[], ) - return daily.reset_index(drop=True) diff --git a/alphaforge/data/public_web/ec_weekly_oil_bulletin.py b/alphaforge/data/public_web/ec_weekly_oil_bulletin.py index 6f7c4d9..71d0f4f 100644 --- a/alphaforge/data/public_web/ec_weekly_oil_bulletin.py +++ b/alphaforge/data/public_web/ec_weekly_oil_bulletin.py @@ -7,21 +7,20 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient from .parsing import parse_xlsx_bytes -from .utils import ( - apply_query_filters, - ensure_date_utc, - make_entity_id, - project_columns, - to_float, +from .schema_helpers import table_schema +from .tabular import ( + TabularDocumentSourceBase, + resolved_date_series, + resolved_numeric_series, + resolved_text_series, ) +from .utils import make_entity_id -class ECWeeklyOilBulletinDataSource(DataSource): +class ECWeeklyOilBulletinDataSource(TabularDocumentSourceBase): name = "ec_weekly_oil_bulletin" TABLE = "ec_oil_bulletin_weekly" @@ -32,17 +31,15 @@ def __init__( cache_dir: str | Path | None = None, bulletin_url: str = "https://energy.ec.europa.eu/data-and-analysis/weekly-oil-bulletin_en", ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._bulletin_url = bulletin_url - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["value"], canonical_columns=["value", "product", "country", "tax_flag"], - entity_column="entity_id", - time_column="date", native_freq="W", ) } @@ -64,8 +61,10 @@ def _discover_links(self) -> list[tuple[str, str]]: return out def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) + snapshot_date = self._snapshot_date(q) rows = [] for url, tax_flag in self._discover_links(): @@ -75,56 +74,35 @@ def fetch(self, q: Query) -> pd.DataFrame: artifact_name=Path(url.split("?")[0]).name, ) frame = parse_xlsx_bytes(payload) - frame.columns = [str(c).lower() for c in frame.columns] - date_col = next( - (c for c in frame.columns if "date" in c or "week" in c), None - ) - product_col = next( - (c for c in frame.columns if "product" in c or "fuel" in c), None - ) - country_col = next( - (c for c in frame.columns if "country" in c or c in {"ms", "geo"}), None - ) - value_col = next( - (c for c in frame.columns if "price" in c or "value" in c), None - ) - if value_col is None: + value_columns = [c for c in frame.columns if "price" in c or "value" in c] + if not value_columns: continue tmp = pd.DataFrame(index=frame.index) - tmp["date"] = ( - ensure_date_utc(frame[date_col]) - if date_col - else pd.Timestamp.now(tz="UTC").normalize() + tmp["date"] = resolved_date_series( + frame, + [c for c in frame.columns if "date" in c or "week" in c], + default_date=snapshot_date, ) - tmp["product"] = ( - frame[product_col].astype(str).str.upper() if product_col else "UNKNOWN" + tmp["product"] = resolved_text_series( + frame, + [c for c in frame.columns if "product" in c or "fuel" in c], + default="UNKNOWN", + case="upper", ) - tmp["country"] = ( - frame[country_col].astype(str).str.upper() if country_col else "EU" + tmp["country"] = resolved_text_series( + frame, + [c for c in frame.columns if "country" in c or c in {"ms", "geo"}], + default="EU", + case="upper", ) - tmp["value"] = to_float(frame[value_col]) + tmp["value"] = resolved_numeric_series(frame, value_columns) tmp["tax_flag"] = tax_flag tmp["entity_id"] = [ make_entity_id("ec_oil", prod, ctry, tax_flag) for prod, ctry in zip(tmp["product"], tmp["country"]) ] - tmp["asof_utc"] = q.asof or pd.Timestamp.now(tz="UTC") + tmp["asof_utc"] = asof_utc rows.append(tmp) - out = ( - pd.concat(rows, ignore_index=True) - if rows - else pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - ) - if out.empty: - return out - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = pd.concat(rows, ignore_index=True) if rows else self._empty_frame(schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/ecb_sdmx.py b/alphaforge/data/public_web/ecb_sdmx.py index 4d4cc1f..23b20fe 100644 --- a/alphaforge/data/public_web/ecb_sdmx.py +++ b/alphaforge/data/public_web/ecb_sdmx.py @@ -7,15 +7,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .registry_loader import load_registry_entries, map_registry -from .utils import apply_query_filters, project_columns +from .registry_api import RegistryApiSourceBase +from .schema_helpers import single_value_schema -class ECBSDMXDataSource(DataSource): +class ECBSDMXDataSource(RegistryApiSourceBase): name = "ecb_sdmx" TABLE = "ecb_sdmx_series" @@ -28,25 +26,17 @@ def __init__( registry_entries: list[dict] | None = None, registry_path: str | Path | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._base_url = base_url.rstrip("/") - entries = load_registry_entries( + self._init_registry( "ecb_sdmx_series.yaml", - entries=registry_entries, + registry_entries=registry_entries, registry_path=registry_path, ) - self._registry = map_registry(entries) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - native_freq="M", - ) + self.TABLE: single_value_schema(self.TABLE, native_freq="M") } @staticmethod @@ -78,17 +68,15 @@ def _call(self, cfg: dict, q: Query) -> pd.DataFrame: return frame def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - entities = list(q.entities or []) - if not entities: - raise ValueError("ECBSDMXDataSource requires q.entities registry keys") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for entity in entities: - cfg = self._registry.get(str(entity)) - if cfg is None: - continue + for entity, cfg in self._iter_entity_configs( + q, + error_message="ECBSDMXDataSource requires q.entities registry keys", + ): df = self._call(cfg, q) period_col = ( "time_period" @@ -111,20 +99,9 @@ def fetch(self, q: Query) -> pd.DataFrame: "date": date, "entity_id": str(entity), "value": pd.to_numeric(row[value_col], errors="coerce"), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/eia.py b/alphaforge/data/public_web/eia.py index 178e6d3..0578f13 100644 --- a/alphaforge/data/public_web/eia.py +++ b/alphaforge/data/public_web/eia.py @@ -8,15 +8,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .registry_loader import load_registry_entries, map_registry -from .utils import apply_query_filters, project_columns +from .registry_api import RegistryApiSourceBase +from .schema_helpers import single_value_schema -class EIADataSource(DataSource): +class EIADataSource(RegistryApiSourceBase): name = "eia" TABLE = "eia_series" @@ -30,23 +28,18 @@ def __init__( registry_entries: list[dict] | None = None, registry_path: str | Path | None = None, ) -> None: + super().__init__(http_client=http_client, cache_dir=cache_dir) self._api_key = api_key or os.getenv("EIA_API_KEY") - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) self._base_url = base_url.rstrip("/") - entries = load_registry_entries( - "eia_series.yaml", entries=registry_entries, registry_path=registry_path + self._init_registry( + "eia_series.yaml", + registry_entries=registry_entries, + registry_path=registry_path, ) - self._registry = map_registry(entries) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - ) + self.TABLE: single_value_schema(self.TABLE) } def _call(self, config: dict, q: Query) -> dict: @@ -79,19 +72,17 @@ def _parse_date(value: str) -> pd.Timestamp | None: return pd.to_datetime(txt, errors="coerce", utc=True) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) if not self._api_key: raise ValueError("EIA API key required via EIA_API_KEY or constructor arg") - entities = list(q.entities or []) - if not entities: - raise ValueError("EIADataSource requires q.entities registry keys") + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for entity in entities: - cfg = self._registry.get(str(entity)) - if cfg is None: - continue + for entity, cfg in self._iter_entity_configs( + q, + error_message="EIADataSource requires q.entities registry keys", + ): payload = self._call(cfg, q) for row in payload.get("response", {}).get("data", []): date = self._parse_date(str(row.get("period", ""))) @@ -106,20 +97,9 @@ def fetch(self, q: Query) -> pd.DataFrame: "date": date, "entity_id": str(entity), "value": pd.to_numeric(value, errors="coerce"), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/eurex_refdata_contracts.py b/alphaforge/data/public_web/eurex_refdata_contracts.py index 6e7b331..2a91d7f 100644 --- a/alphaforge/data/public_web/eurex_refdata_contracts.py +++ b/alphaforge/data/public_web/eurex_refdata_contracts.py @@ -6,20 +6,15 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .base import PublicWebSourceBase from .http import CachedHttpClient -from .utils import ( - apply_query_filters, - ensure_date_utc, - first_existing, - make_entity_id, - project_columns, -) +from .schema_helpers import table_schema +from .tabular import resolved_text_series +from .utils import ensure_date_utc, first_existing, make_entity_id -class EurexRefdataContractsSource(DataSource): +class EurexRefdataContractsSource(PublicWebSourceBase): name: str = "eurex_refdata_contracts" TABLE = "eurex.refdata.contracts" @@ -31,12 +26,12 @@ def __init__( cache_dir: str | Path | None = None, ) -> None: self._api_url = api_url - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=[ "symbol", "product_name", @@ -55,16 +50,13 @@ def schemas(self) -> dict[str, TableSchema]: "tick_size", "isin", ], - entity_column="entity_id", - time_column="date", native_freq="D", time_semantics="point", ) } def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) payload = self._http.get_bytes( url=self._api_url, @@ -76,27 +68,26 @@ def fetch(self, q: Query) -> pd.DataFrame: records = data if isinstance(data, list) else data.get("data", []) frame = pd.DataFrame(records) - schema = self.schemas()[self.TABLE] + schema = self._schema() if frame.empty: - return pd.DataFrame( - columns=[ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - ) + return self._empty_frame(schema, entity_col="entity_id") out = pd.DataFrame(index=frame.index) - out["symbol"] = frame[ - first_existing(frame, "symbol", "contract_symbol") - ].astype(str) - out["product_name"] = frame[ - first_existing(frame, "product_name", "product") - ].astype(str) - out["product_group"] = frame[ - first_existing(frame, "product_group", "group") - ].astype(str) + out["symbol"] = resolved_text_series( + frame, + ["symbol", "contract_symbol"], + default="", + ) + out["product_name"] = resolved_text_series( + frame, + ["product_name", "product"], + default="", + ) + out["product_group"] = resolved_text_series( + frame, + ["product_group", "group"], + default="", + ) out["currency"] = ( frame[first_existing(frame, "currency", "ccy")].astype(str).str.lower() ) @@ -108,7 +99,7 @@ def fetch(self, q: Query) -> pd.DataFrame: src = first_existing(frame, col) out[col] = frame[src] if src else pd.NA - snapshot = pd.Timestamp.now(tz="UTC") + snapshot = self._asof_utc(q) out["date"] = snapshot.normalize() out["asof_utc"] = snapshot @@ -123,12 +114,4 @@ def fetch(self, q: Query) -> pd.DataFrame: for symbol, expiry in zip(out["symbol"].str.lower(), out["expiry_date"]) ] - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + return self._finalize(out, q=q, schema=schema, sort_by=[]) diff --git a/alphaforge/data/public_web/eurex_stats_daily.py b/alphaforge/data/public_web/eurex_stats_daily.py index ddc2d2a..082f518 100644 --- a/alphaforge/data/public_web/eurex_stats_daily.py +++ b/alphaforge/data/public_web/eurex_stats_daily.py @@ -5,22 +5,20 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .parsing import parse_html_tables -from .utils import ( - apply_query_filters, - ensure_date_utc, - first_existing, - make_entity_id, - project_columns, - to_float, +from .schema_helpers import table_schema +from .tabular import ( + TabularDocumentSourceBase, + candidate_tables, + resolved_date_series, + resolved_numeric_series, + resolved_text_series, ) +from .utils import make_entity_id -class EurexStatsDailySource(DataSource): +class EurexStatsDailySource(TabularDocumentSourceBase): name: str = "eurex_stats_daily" TABLE = "eurex.stats.daily" URL = "https://www.eurex.com/ex-en/data/statistics/market-statistics-online" @@ -32,13 +30,13 @@ def __init__( cache_dir: str | Path | None = None, stats_url: str | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._stats_url = stats_url or self.URL - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["volume", "open_interest"], canonical_columns=[ "volume", @@ -48,89 +46,61 @@ def schemas(self) -> dict[str, TableSchema]: "contract_count", "trades", ], - entity_column="entity_id", - time_column="date", native_freq="D", time_semantics="point", ) } def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) + snapshot_date = self._snapshot_date(q) - payload = self._http.get_bytes( + tables = self._read_html_tables( url=self._stats_url, source="eurex_stats_daily", - artifact_name=Path(self._stats_url.split("?")[0]).name - or "market_stats.html", + artifact_name=Path(self._stats_url.split("?")[0]).name or "market_stats.html", ) - - tables = parse_html_tables(payload) rows: list[pd.DataFrame] = [] - for table in tables: - cols = set(table.columns) - if not ({"volume", "open_interest"} & cols): - continue - + for table in candidate_tables(tables, any_of=("volume", "open_interest")): out = pd.DataFrame(index=table.index) - date_col = first_existing(table, "date", "trading_day") - out["date"] = ( - ensure_date_utc(table[date_col]) - if date_col - else pd.Timestamp.now(tz="UTC").normalize() + out["date"] = resolved_date_series( + table, + ["date", "trading_day"], + default_date=snapshot_date, ) - - pg_col = first_existing(table, "product_group", "group") - pn_col = first_existing(table, "product_name", "product", "contract") - out["product_group"] = ( - table[pg_col].astype(str).str.lower().str.replace(" ", "_", regex=False) - if pg_col - else "unknown" + out["product_group"] = resolved_text_series( + table, + ["product_group", "group"], + default="unknown", + case="lower", + space_replacement="_", ) - out["product_name"] = ( - table[pn_col].astype(str).str.lower().str.replace(" ", "_", regex=False) - if pn_col - else "unknown" + out["product_name"] = resolved_text_series( + table, + ["product_name", "product", "contract"], + default="unknown", + case="lower", + space_replacement="_", ) - - vol_col = first_existing(table, "volume") - oi_col = first_existing(table, "open_interest", "openinterest") - out["volume"] = to_float(table[vol_col]) if vol_col else pd.NA - out["open_interest"] = to_float(table[oi_col]) if oi_col else pd.NA - - trades_col = first_existing(table, "trades") - contract_count_col = first_existing(table, "contract_count", "contracts") - out["trades"] = to_float(table[trades_col]) if trades_col else pd.NA - out["contract_count"] = ( - to_float(table[contract_count_col]) if contract_count_col else pd.NA + out["volume"] = resolved_numeric_series(table, ["volume"]) + out["open_interest"] = resolved_numeric_series( + table, + ["open_interest", "openinterest"], + ) + out["trades"] = resolved_numeric_series(table, ["trades"]) + out["contract_count"] = resolved_numeric_series( + table, + ["contract_count", "contracts"], ) out["entity_id"] = [ make_entity_id("eurex", group, name) for group, name in zip(out["product_group"], out["product_name"]) ] - out["asof_utc"] = pd.Timestamp.now(tz="UTC") + out["asof_utc"] = asof_utc rows.append(out) - schema = self.schemas()[self.TABLE] - if not rows: - return pd.DataFrame( - columns=[ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - ) - - out = pd.concat(rows, ignore_index=True) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + out = pd.concat(rows, ignore_index=True) if rows else self._empty_frame(schema) + return self._finalize(out, q=q, schema=schema, sort_by=[]) diff --git a/alphaforge/data/public_web/eurostat.py b/alphaforge/data/public_web/eurostat.py index f0c7ed5..510c487 100644 --- a/alphaforge/data/public_web/eurostat.py +++ b/alphaforge/data/public_web/eurostat.py @@ -8,15 +8,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .registry_loader import load_registry_entries, map_registry -from .utils import apply_query_filters, project_columns +from .registry_api import RegistryApiSourceBase +from .schema_helpers import single_value_schema -class EurostatDataSource(DataSource): +class EurostatDataSource(RegistryApiSourceBase): name = "eurostat" TABLE = "eurostat_series" @@ -29,25 +27,17 @@ def __init__( registry_entries: list[dict] | None = None, registry_path: str | Path | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._base_url = base_url.rstrip("/") - entries = load_registry_entries( + self._init_registry( "eurostat_series.yaml", - entries=registry_entries, + registry_entries=registry_entries, registry_path=registry_path, ) - self._registry = map_registry(entries) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - native_freq="M", - ) + self.TABLE: single_value_schema(self.TABLE, native_freq="M") } @staticmethod @@ -101,17 +91,15 @@ def _call(self, cfg: dict) -> pd.DataFrame: return frame def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - entities = list(q.entities or []) - if not entities: - raise ValueError("EurostatDataSource requires q.entities registry keys") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for entity in entities: - cfg = self._registry.get(str(entity)) - if cfg is None: - continue + for entity, cfg in self._iter_entity_configs( + q, + error_message="EurostatDataSource requires q.entities registry keys", + ): df = self._call(cfg) period_col = ( "period" @@ -134,20 +122,9 @@ def fetch(self, q: Query) -> pd.DataFrame: "date": date, "entity_id": str(entity), "value": pd.to_numeric(row[value_col], errors="coerce"), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/ezoic_adrevenue_daily.py b/alphaforge/data/public_web/ezoic_adrevenue_daily.py index f35079f..849a4c2 100644 --- a/alphaforge/data/public_web/ezoic_adrevenue_daily.py +++ b/alphaforge/data/public_web/ezoic_adrevenue_daily.py @@ -8,20 +8,14 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .utils import ( - apply_query_filters, - ensure_date_utc, - make_entity_id, - project_columns, - to_float, -) +from .schema_helpers import table_schema +from .tabular import TabularDocumentSourceBase, artifact_name_from_url +from .utils import ensure_date_utc, make_entity_id, to_float -class EzoicAdRevenueDailySource(DataSource): +class EzoicAdRevenueDailySource(TabularDocumentSourceBase): name: str = "ezoic_adrevenue_daily" TABLE = "ezoic.adrevenue.daily" URL = "https://adrevenueindex.ezoic.com/" @@ -36,16 +30,14 @@ def __init__( ) -> None: self._page_url = page_url or self.URL self._data_url = data_url - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["value"], canonical_columns=["value", "region", "category"], - entity_column="entity_id", - time_column="date", native_freq="D", time_semantics="point", ) @@ -85,8 +77,7 @@ def _load_records(self) -> list[dict[str, Any]]: payload = self._http.get_bytes( url=self._data_url, source="ezoic_adrevenue_daily", - artifact_name=Path(self._data_url.split("?")[0]).name - or "adrevenue.json", + artifact_name=artifact_name_from_url(self._data_url, "adrevenue.json"), ) parsed = json.loads(payload.decode()) if isinstance(parsed, list): @@ -99,26 +90,18 @@ def _load_records(self) -> list[dict[str, Any]]: payload = self._http.get_bytes( url=self._page_url, source="ezoic_adrevenue_daily", - artifact_name=Path(self._page_url.split("?")[0]).name or "adrevenue.html", + artifact_name=artifact_name_from_url(self._page_url, "adrevenue.html"), ) html = payload.decode(errors="ignore") return self._extract_records_from_html(html) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) records = self._load_records() - schema = self.schemas()[self.TABLE] + schema = self._schema() if not records: - return pd.DataFrame( - columns=[ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - ) + return self._empty_frame(schema) frame = pd.DataFrame(records) date_col = "date" if "date" in frame.columns else "day" @@ -135,18 +118,10 @@ def fetch(self, q: Query) -> pd.DataFrame: out["category"] = ( frame[category_col].astype(str).str.lower() if category_col else "all" ) - out["asof_utc"] = pd.Timestamp.now(tz="UTC") + out["asof_utc"] = self._asof_utc(q) out["entity_id"] = [ make_entity_id("macro", "index", "adrevenue", region, "value", "ezoic") for region in out["region"] ] - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + return self._finalize(out, q=q, schema=schema, sort_by=[]) diff --git a/alphaforge/data/public_web/finalize.py b/alphaforge/data/public_web/finalize.py new file mode 100644 index 0000000..d0c6d4a --- /dev/null +++ b/alphaforge/data/public_web/finalize.py @@ -0,0 +1,106 @@ +from __future__ import annotations + +from collections.abc import Iterable, Sequence + +import pandas as pd + +from alphaforge.data.query import Query +from alphaforge.data.schema import TableSchema + +from .utils import apply_query_filters, project_columns + + +def _ordered_unique(columns: Iterable[str]) -> list[str]: + out: list[str] = [] + seen: set[str] = set() + for column in columns: + if column not in seen: + seen.add(column) + out.append(column) + return out + + +def schema_frame_columns( + schema: TableSchema, + *, + time_col: str | None = None, + entity_col: str | None = None, +) -> list[str]: + return _ordered_unique( + [ + time_col or schema.time_column, + entity_col or schema.entity_column, + "asof_utc", + *schema.required_columns, + *schema.canonical_columns, + ] + ) + + +def empty_frame_for_schema( + schema: TableSchema, + *, + time_col: str | None = None, + entity_col: str | None = None, +) -> pd.DataFrame: + columns = schema_frame_columns(schema, time_col=time_col, entity_col=entity_col) + return pd.DataFrame(columns=columns) + + +def frame_from_records( + records: Sequence[dict] | Iterable[dict], + *, + schema: TableSchema, + time_col: str | None = None, + entity_col: str | None = None, +) -> pd.DataFrame: + rows = list(records) + if not rows: + return empty_frame_for_schema(schema, time_col=time_col, entity_col=entity_col) + return pd.DataFrame.from_records(rows) + + +def finalize_public_frame( + df: pd.DataFrame, + *, + q: Query, + schema: TableSchema, + time_col: str | None = None, + entity_col: str | None = None, + sort_by: Sequence[str] | None = None, +) -> pd.DataFrame: + resolved_time_col = time_col or schema.time_column + resolved_entity_col = entity_col or schema.entity_column + + if df.empty: + return empty_frame_for_schema( + schema, + time_col=resolved_time_col, + entity_col=resolved_entity_col, + ) + + out = apply_query_filters( + df, + q=q, + time_col=resolved_time_col, + entity_col=resolved_entity_col, + ) + out = project_columns( + out, + required_columns=schema.required_columns, + requested_columns=q.columns, + time_col=resolved_time_col, + entity_col=resolved_entity_col, + ) + sort_columns = [ + column + for column in ( + list(sort_by) + if sort_by is not None + else [resolved_entity_col, resolved_time_col] + ) + if column in out.columns + ] + if sort_columns: + out = out.sort_values(sort_columns) + return out.reset_index(drop=True) diff --git a/alphaforge/data/public_web/frb_term_structure.py b/alphaforge/data/public_web/frb_term_structure.py index 3a0b366..72bf586 100644 --- a/alphaforge/data/public_web/frb_term_structure.py +++ b/alphaforge/data/public_web/frb_term_structure.py @@ -9,20 +9,18 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .base import PublicWebSourceBase from .http import CachedHttpClient +from .schema_helpers import table_schema from .utils import ( - apply_query_filters, bucket_tenor, ensure_date_utc, make_entity_id, - project_columns, ) -class FRBTermStructureBenchmarkSource(DataSource): +class FRBTermStructureBenchmarkSource(PublicWebSourceBase): """Federal Reserve Board Kim-Wright three-factor benchmark series.""" name: str = "frb_term_structure" @@ -40,17 +38,15 @@ def __init__( cache_dir: str | Path | None = None, csv_url: str | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._csv_url = csv_url or self.CSV_URL - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["value"], canonical_columns=["value", "mnemonic", "category", "maturity_years"], - entity_column="entity_id", - time_column="date", native_freq="B", time_semantics="point", ) @@ -82,17 +78,7 @@ def _load_raw(self) -> pd.DataFrame: def _to_long(self, raw: pd.DataFrame) -> pd.DataFrame: if raw.empty: - return pd.DataFrame( - columns=[ - "date", - "mnemonic", - "category", - "maturity_years", - "value", - "entity_id", - "asof_utc", - ] - ) + return self._empty_frame(self._schema()) date_column = raw.columns[0] long = raw.rename(columns={date_column: "date"}).melt( @@ -123,18 +109,10 @@ def _to_long(self, raw: pd.DataFrame) -> pd.DataFrame: return long.sort_values(["date", "mnemonic"]).reset_index(drop=True) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) long = self._to_long(self._load_raw()) - long = apply_query_filters(long, q=q, time_col="date", entity_col="entity_id") - return project_columns( - long, - required_columns=["value"], - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) + return self._finalize(long, q=q, schema=self._schema(), sort_by=["date", "mnemonic"]) def fetch_wide(self, q: Query, *, category: str | None = None) -> pd.DataFrame: df = self.fetch( diff --git a/alphaforge/data/public_web/ibge_sidra.py b/alphaforge/data/public_web/ibge_sidra.py index 7d83e79..b986fa8 100644 --- a/alphaforge/data/public_web/ibge_sidra.py +++ b/alphaforge/data/public_web/ibge_sidra.py @@ -7,15 +7,13 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .registry_loader import load_registry_entries, map_registry -from .utils import apply_query_filters, project_columns +from .registry_api import RegistryApiSourceBase +from .schema_helpers import single_value_schema -class IBGESidraDataSource(DataSource): +class IBGESidraDataSource(RegistryApiSourceBase): name = "ibge_sidra" TABLE = "ibge_sidra_series" @@ -28,25 +26,17 @@ def __init__( registry_entries: list[dict] | None = None, registry_path: str | Path | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._base_url = base_url.rstrip("/") - entries = load_registry_entries( + self._init_registry( "ibge_sidra_series.yaml", - entries=registry_entries, + registry_entries=registry_entries, registry_path=registry_path, ) - self._registry = map_registry(entries) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, - required_columns=["value"], - canonical_columns=["value"], - entity_column="entity_id", - time_column="date", - native_freq="M", - ) + self.TABLE: single_value_schema(self.TABLE, native_freq="M") } def _call(self, cfg: dict) -> list[dict]: @@ -79,17 +69,15 @@ def _parse_period(value: str) -> pd.Timestamp | None: return pd.to_datetime(txt, errors="coerce", utc=True) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - entities = list(q.entities or []) - if not entities: - raise ValueError("IBGESidraDataSource requires q.entities registry keys") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) rows = [] - for entity in entities: - cfg = self._registry.get(str(entity)) - if cfg is None: - continue + for entity, cfg in self._iter_entity_configs( + q, + error_message="IBGESidraDataSource requires q.entities registry keys", + ): data = self._call(cfg) for row in data: period = row.get("D3C") or row.get("Mês (Código)") or row.get("V") @@ -104,20 +92,9 @@ def fetch(self, q: Query) -> pd.DataFrame: "value": pd.to_numeric( str(value).replace(",", "."), errors="coerce" ), - "asof_utc": q.asof or pd.Timestamp.now(tz="UTC"), + "asof_utc": asof_utc, } ) - out = pd.DataFrame(rows) - if out.empty: - return pd.DataFrame(columns=["date", "entity_id", "asof_utc", "value"]) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - schema = self.schemas()[self.TABLE] - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", - ) - return out.sort_values(["entity_id", "date"]).reset_index(drop=True) + out = self._frame_from_records(rows, schema=schema) + return self._finalize(out, q=q, schema=schema) diff --git a/alphaforge/data/public_web/lch_cdsclear_daily.py b/alphaforge/data/public_web/lch_cdsclear_daily.py index ed0276d..d3008ab 100644 --- a/alphaforge/data/public_web/lch_cdsclear_daily.py +++ b/alphaforge/data/public_web/lch_cdsclear_daily.py @@ -5,22 +5,20 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource from .http import CachedHttpClient -from .parsing import parse_html_tables -from .utils import ( - apply_query_filters, - ensure_date_utc, - first_existing, - make_entity_id, - project_columns, - to_float, +from .schema_helpers import table_schema +from .tabular import ( + TabularDocumentSourceBase, + candidate_tables, + resolved_date_series, + resolved_numeric_series, + resolved_text_series, ) +from .utils import make_entity_id -class LCHCDSClearDailySource(DataSource): +class LCHCDSClearDailySource(TabularDocumentSourceBase): name: str = "lch_cdsclear_daily" TABLE = "lch.cdsclear.daily" URL = "https://www.lseg.com/en/post-trade/clearing/lch-services/cdsclear/volumes" @@ -33,92 +31,67 @@ def __init__( cache_dir: str | Path | None = None, ) -> None: self._volumes_url = volumes_url or self.URL - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["value"], canonical_columns=["value", "metric", "segment"], - entity_column="entity_id", - time_column="date", native_freq="D", time_semantics="point", ) } def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) + schema = self._schema() + asof_utc = self._asof_utc(q) + snapshot_date = self._snapshot_date(q) - payload = self._http.get_bytes( + tables = self._read_html_tables( url=self._volumes_url, source="lch_cdsclear_daily", artifact_name=Path(self._volumes_url.split("?")[0]).name or "volumes.html", ) - tables = parse_html_tables(payload) rows: list[pd.DataFrame] = [] - for table in tables: - if not ({"value", "volume", "notional", "trades"} & set(table.columns)): - continue - + for table in candidate_tables( + tables, + any_of=("value", "volume", "notional", "trades"), + ): out = pd.DataFrame(index=table.index) - date_col = first_existing(table, "date", "trading_day") - out["date"] = ( - ensure_date_utc(table[date_col]) - if date_col - else pd.Timestamp.now(tz="UTC").normalize() + out["date"] = resolved_date_series( + table, + ["date", "trading_day"], + default_date=snapshot_date, ) - - metric_col = first_existing(table, "metric") - segment_col = first_existing(table, "segment", "family", "index_family") - out["metric"] = ( - table[metric_col] - .astype(str) - .str.lower() - .str.replace(" ", "_", regex=False) - if metric_col - else "volume" + out["metric"] = resolved_text_series( + table, + ["metric"], + default="volume", + case="lower", + space_replacement="_", ) - out["segment"] = ( - table[segment_col] - .astype(str) - .str.lower() - .str.replace(" ", "_", regex=False) - if segment_col - else "all" + out["segment"] = resolved_text_series( + table, + ["segment", "family", "index_family"], + default="all", + case="lower", + space_replacement="_", + ) + out["value"] = resolved_numeric_series( + table, + ["value", "volume", "notional", "trades"], ) - - value_col = first_existing(table, "value", "volume", "notional", "trades") - out["value"] = to_float(table[value_col]) if value_col else pd.NA out["entity_id"] = [ make_entity_id("credit", "cds", seg, metric, "lch") for seg, metric in zip(out["segment"], out["metric"]) ] - out["asof_utc"] = pd.Timestamp.now(tz="UTC") + out["asof_utc"] = asof_utc rows.append(out) - schema = self.schemas()[self.TABLE] - if not rows: - return pd.DataFrame( - columns=[ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - ) - - out = pd.concat(rows, ignore_index=True) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + out = pd.concat(rows, ignore_index=True) if rows else self._empty_frame(schema) + return self._finalize(out, q=q, schema=schema, sort_by=[]) diff --git a/alphaforge/data/public_web/mof_jgb.py b/alphaforge/data/public_web/mof_jgb.py index ea99b94..91ab869 100644 --- a/alphaforge/data/public_web/mof_jgb.py +++ b/alphaforge/data/public_web/mof_jgb.py @@ -25,19 +25,17 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .base import PublicWebSourceBase from .http import CachedHttpClient +from .schema_helpers import table_schema from .utils import ( - apply_query_filters, ensure_date_utc, make_entity_id, - project_columns, ) -class MOFJGBYieldCurveSource(DataSource): +class MOFJGBYieldCurveSource(PublicWebSourceBase): """Daily JGB constant-maturity par yields from the MOF website.""" name: str = "mof_jgb_yields" @@ -58,24 +56,22 @@ def __init__( csv_url: str | None = None, landing_url: str | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._csv_url = csv_url or self.CURRENT_CSV_URL self._landing_url = landing_url or self.LANDING_URL # ---- schema -------------------------------------------------------------- - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["yield_pct"], canonical_columns=[ "yield_pct", "tenor", "maturity_years", ], - entity_column="entity_id", - time_column="date", native_freq="B", time_semantics="point", ) @@ -263,18 +259,9 @@ def _to_long(self, wide: pd.DataFrame) -> pd.DataFrame: # ---- fetch --------------------------------------------------------------- def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") - - schema = self.schemas()[self.TABLE] - empty = pd.DataFrame( - columns=[ - schema.time_column, - schema.entity_column, - "asof_utc", - *schema.required_columns, - ] - ) + self._require_table(q) + + schema = self._schema() frames: list[pd.DataFrame] = [] for url in self._discover_csv_urls(): @@ -293,18 +280,10 @@ def fetch(self, q: Query) -> pd.DataFrame: continue if not frames: - return empty + return self._empty_frame(schema) out = pd.concat(frames, ignore_index=True) - out = apply_query_filters(out, q=q, time_col="date", entity_col="entity_id") - out = project_columns( - out, - required_columns=schema.required_columns, - requested_columns=q.columns, - time_col=schema.time_column, - entity_col=schema.entity_column, - ) - return out.reset_index(drop=True) + return self._finalize(out, q=q, schema=schema, sort_by=[]) # ---- convenience --------------------------------------------------------- diff --git a/alphaforge/data/public_web/philadelphia_spf.py b/alphaforge/data/public_web/philadelphia_spf.py index 630fd39..a02db56 100644 --- a/alphaforge/data/public_web/philadelphia_spf.py +++ b/alphaforge/data/public_web/philadelphia_spf.py @@ -9,20 +9,18 @@ import pandas as pd from alphaforge.data.query import Query -from alphaforge.data.schema import TableSchema -from alphaforge.data.source import DataSource +from .base import PublicWebSourceBase from .http import CachedHttpClient +from .schema_helpers import table_schema from .utils import ( - apply_query_filters, ensure_date_utc, make_entity_id, - project_columns, snake_case, ) -class PhiladelphiaSPFMeanLevelSource(DataSource): +class PhiladelphiaSPFMeanLevelSource(PublicWebSourceBase): """Historical mean SPF forecasts from the Philadelphia Fed.""" name: str = "philadelphia_spf" @@ -56,14 +54,14 @@ def __init__( workbook_url: str | None = None, release_url: str | None = None, ) -> None: - self._http = http_client or CachedHttpClient(cache_dir=cache_dir) + super().__init__(http_client=http_client, cache_dir=cache_dir) self._workbook_url = workbook_url or self.WORKBOOK_URL self._release_url = release_url or self.RELEASE_URL - def schemas(self) -> dict[str, TableSchema]: + def schemas(self): return { - self.TABLE: TableSchema( - name=self.TABLE, + self.TABLE: table_schema( + self.TABLE, required_columns=["value"], canonical_columns=[ "value", @@ -72,8 +70,6 @@ def schemas(self) -> dict[str, TableSchema]: "survey_period", "release_date", ], - entity_column="entity_id", - time_column="date", native_freq="Q", time_semantics="point", release_time_column="release_date", @@ -200,18 +196,7 @@ def _to_long(self) -> pd.DataFrame: ] frames = [frame for frame in frames if not frame.empty] if not frames: - return pd.DataFrame( - columns=[ - "date", - "release_date", - "survey_period", - "sheet_name", - "series_name", - "value", - "entity_id", - "asof_utc", - ] - ) + return self._empty_frame(self._schema()) long = pd.concat(frames, ignore_index=True) if not release_calendar.empty: @@ -231,15 +216,12 @@ def _to_long(self) -> pd.DataFrame: return long.sort_values(["date", "sheet_name", "series_name"]).reset_index(drop=True) def fetch(self, q: Query) -> pd.DataFrame: - if q.table != self.TABLE: - raise ValueError(f"Unknown table: {q.table}") + self._require_table(q) long = self._to_long() - long = apply_query_filters(long, q=q, time_col="date", entity_col="entity_id") - return project_columns( + return self._finalize( long, - required_columns=["value"], - requested_columns=q.columns, - time_col="date", - entity_col="entity_id", + q=q, + schema=self._schema(), + sort_by=["date", "sheet_name", "series_name"], ) diff --git a/alphaforge/data/public_web/registry.py b/alphaforge/data/public_web/registry.py index 6df7d36..787eb41 100644 --- a/alphaforge/data/public_web/registry.py +++ b/alphaforge/data/public_web/registry.py @@ -1,5 +1,7 @@ from __future__ import annotations +from collections.abc import Callable + from alphaforge.data.source import DataSource from .anp_fuel_prices import ANPFuelPricesDataSource @@ -7,7 +9,7 @@ from .bcb_sgs import BCBSGSDataSource from .bea import BEADataSource from .bls import BLSDataSource -from .cftc_cot import CFTCCoTSource +from .cftc_cot import CFTCCoTSource, CFTCDisaggregatedCoTSource from .cftc_swaps_weekly import CFTCWeeklySwapsSource from .cme_productslate_reference import CMEProductSlateSource from .destatis_genesis import DestatisGenesisDataSource @@ -25,32 +27,44 @@ from .mof_jgb import MOFJGBYieldCurveSource from .philadelphia_spf import PhiladelphiaSPFMeanLevelSource +DEFAULT_EUREX_REFDATA_API_URL = "https://www.eurex.com/api/refdata/contracts" + + +def _default_eurex_refdata_source() -> DataSource: + return EurexRefdataContractsSource(api_url=DEFAULT_EUREX_REFDATA_API_URL) + + +DEFAULT_PUBLIC_WEB_SOURCE_FACTORIES: tuple[Callable[[], DataSource], ...] = ( + BLSDataSource, + BEADataSource, + EIADataSource, + EurostatDataSource, + ECBSDMXDataSource, + DestatisGenesisDataSource, + ECWeeklyOilBulletinDataSource, + IBGESidraDataSource, + BCBSGSDataSource, + ANPFuelPricesDataSource, + B3HistoricalQuotesDataSource, + DTCCPPDSource, + CMEProductSlateSource, + CFTCWeeklySwapsSource, + CFTCCoTSource, + CFTCDisaggregatedCoTSource, + EurexStatsDailySource, + LCHCDSClearDailySource, + EzoicAdRevenueDailySource, + PhiladelphiaSPFMeanLevelSource, + FRBTermStructureBenchmarkSource, + _default_eurex_refdata_source, + MOFJGBYieldCurveSource, +) + def default_public_web_sources() -> dict[str, DataSource]: - sources = [ - BLSDataSource(), - BEADataSource(), - EIADataSource(), - EurostatDataSource(), - ECBSDMXDataSource(), - DestatisGenesisDataSource(), - ECWeeklyOilBulletinDataSource(), - IBGESidraDataSource(), - BCBSGSDataSource(), - ANPFuelPricesDataSource(), - B3HistoricalQuotesDataSource(), - DTCCPPDSource(), - CMEProductSlateSource(), - CFTCWeeklySwapsSource(), - CFTCCoTSource(), - EurexStatsDailySource(), - LCHCDSClearDailySource(), - EzoicAdRevenueDailySource(), - PhiladelphiaSPFMeanLevelSource(), - FRBTermStructureBenchmarkSource(), - EurexRefdataContractsSource( - api_url="https://www.eurex.com/api/refdata/contracts" - ), - MOFJGBYieldCurveSource(), - ] - return {source.name: source for source in sources} + """Construct the default public-web source registry.""" + + return { + source.name: source + for source in (factory() for factory in DEFAULT_PUBLIC_WEB_SOURCE_FACTORIES) + } diff --git a/alphaforge/data/public_web/registry_api.py b/alphaforge/data/public_web/registry_api.py new file mode 100644 index 0000000..b3b42fc --- /dev/null +++ b/alphaforge/data/public_web/registry_api.py @@ -0,0 +1,37 @@ +from __future__ import annotations + +from pathlib import Path +from typing import Any + +from alphaforge.data.query import Query + +from .base import PublicWebSourceBase +from .registry_loader import load_registry_entries, map_registry + + +class RegistryApiSourceBase(PublicWebSourceBase): + def _init_registry( + self, + file_name: str, + *, + registry_entries: list[dict[str, Any]] | None = None, + registry_path: str | Path | None = None, + ) -> None: + entries = load_registry_entries( + file_name, + entries=registry_entries, + registry_path=registry_path, + ) + self._registry = map_registry(entries) + + def _iter_entity_configs( + self, + q: Query, + *, + error_message: str, + ): + for entity in self._require_entities(q, error_message=error_message): + config = self._registry.get(str(entity)) + if config is None: + continue + yield str(entity), config diff --git a/alphaforge/data/public_web/schema_helpers.py b/alphaforge/data/public_web/schema_helpers.py new file mode 100644 index 0000000..2edf422 --- /dev/null +++ b/alphaforge/data/public_web/schema_helpers.py @@ -0,0 +1,110 @@ +from __future__ import annotations + +from collections.abc import Sequence + +from alphaforge.data.schema import TableSchema + + +def _normalize_columns(columns: Sequence[str] | None) -> list[str]: + return [str(column) for column in (columns or [])] + + +def table_schema( + name: str, + *, + required_columns: Sequence[str], + canonical_columns: Sequence[str] | None = None, + entity_column: str = "entity_id", + time_column: str = "date", + native_freq: str | None = None, + time_semantics: str | None = None, + expected_cadence_days: int | None = None, + event_time_column: str | None = None, + release_time_column: str | None = None, + revision_id_column: str | None = None, +) -> TableSchema: + canonical = _normalize_columns(canonical_columns) or _normalize_columns( + required_columns + ) + return TableSchema( + name=name, + required_columns=_normalize_columns(required_columns), + canonical_columns=canonical, + entity_column=entity_column, + time_column=time_column, + native_freq=native_freq, + time_semantics=time_semantics, + expected_cadence_days=expected_cadence_days, + event_time_column=event_time_column, + release_time_column=release_time_column, + revision_id_column=revision_id_column, + ) + + +def single_value_schema( + name: str, + *, + value_column: str = "value", + entity_column: str = "entity_id", + time_column: str = "date", + native_freq: str | None = None, + time_semantics: str | None = None, + expected_cadence_days: int | None = None, +) -> TableSchema: + return table_schema( + name, + required_columns=[value_column], + canonical_columns=[value_column], + entity_column=entity_column, + time_column=time_column, + native_freq=native_freq, + time_semantics=time_semantics, + expected_cadence_days=expected_cadence_days, + ) + + +def daily_panel_schema( + name: str, + *, + required_columns: Sequence[str], + canonical_columns: Sequence[str] | None = None, + entity_column: str = "entity_id", + time_column: str = "date", +) -> TableSchema: + return table_schema( + name, + required_columns=required_columns, + canonical_columns=canonical_columns, + entity_column=entity_column, + time_column=time_column, + native_freq="D", + expected_cadence_days=1, + ) + + +def event_table_schema( + name: str, + *, + required_columns: Sequence[str], + canonical_columns: Sequence[str] | None = None, + entity_column: str = "entity_id", + time_column: str = "ts_utc", + native_freq: str | None = None, + time_semantics: str | None = None, + expected_cadence_days: int | None = None, + release_time_column: str | None = "asof_utc", + revision_id_column: str | None = None, +) -> TableSchema: + return table_schema( + name, + required_columns=required_columns, + canonical_columns=canonical_columns, + entity_column=entity_column, + time_column=time_column, + native_freq=native_freq, + time_semantics=time_semantics, + expected_cadence_days=expected_cadence_days, + event_time_column=time_column, + release_time_column=release_time_column, + revision_id_column=revision_id_column, + ) diff --git a/alphaforge/data/public_web/tabular.py b/alphaforge/data/public_web/tabular.py new file mode 100644 index 0000000..6681b89 --- /dev/null +++ b/alphaforge/data/public_web/tabular.py @@ -0,0 +1,135 @@ +from __future__ import annotations + +from collections.abc import Iterable, Sequence +from pathlib import Path + +import pandas as pd + +from alphaforge.data.query import Query + +from .base import PublicWebSourceBase +from .parsing import parse_csv_bytes, parse_html_tables, parse_xlsx_bytes +from .utils import ensure_date_utc, first_existing, to_float + + +def artifact_name_from_url(url: str, fallback: str) -> str: + path = Path(str(url).split("?", 1)[0]) + return path.name or fallback + + +def candidate_tables( + tables: Iterable[pd.DataFrame], + *, + any_of: Iterable[str] | None = None, + all_of: Iterable[str] | None = None, +) -> list[pd.DataFrame]: + any_columns = {str(column) for column in (any_of or [])} + all_columns = {str(column) for column in (all_of or [])} + + selected: list[pd.DataFrame] = [] + for table in tables: + columns = set(table.columns) + if any_columns and not (columns & any_columns): + continue + if all_columns and not all_columns.issubset(columns): + continue + selected.append(table) + return selected + + +def resolved_date_series( + frame: pd.DataFrame, + aliases: Sequence[str], + *, + default_date: pd.Timestamp, +) -> pd.Series: + column = first_existing(frame, *aliases) + if column is None: + return pd.Series(default_date, index=frame.index) + return ensure_date_utc(frame[column]) + + +def resolved_numeric_series( + frame: pd.DataFrame, + aliases: Sequence[str], +) -> pd.Series: + column = first_existing(frame, *aliases) + if column is None: + return pd.Series(pd.NA, index=frame.index) + return to_float(frame[column]) + + +def resolved_text_series( + frame: pd.DataFrame, + aliases: Sequence[str], + *, + default: str, + case: str | None = None, + space_replacement: str | None = None, +) -> pd.Series: + column = first_existing(frame, *aliases) + if column is None: + series = pd.Series(default, index=frame.index) + else: + raw = frame[column].where(frame[column].notna(), default) + series = raw.astype(str).str.strip() + + if case == "lower": + series = series.str.lower() + elif case == "upper": + series = series.str.upper() + + if space_replacement is not None: + series = series.str.replace(" ", space_replacement, regex=False) + + return series + + +class TabularDocumentSourceBase(PublicWebSourceBase): + def _artifact_name_from_url(self, url: str, fallback: str) -> str: + return artifact_name_from_url(url, fallback) + + def _read_html_tables( + self, + *, + url: str, + source: str, + artifact_name: str, + ) -> list[pd.DataFrame]: + payload = self._http.get_bytes( + url=url, + source=source, + artifact_name=artifact_name, + ) + return parse_html_tables(payload) + + def _read_xlsx_frame( + self, + *, + url: str, + source: str, + artifact_name: str, + ) -> pd.DataFrame: + payload = self._http.get_bytes( + url=url, + source=source, + artifact_name=artifact_name, + ) + return parse_xlsx_bytes(payload) + + def _read_csv_frame( + self, + *, + url: str, + source: str, + artifact_name: str, + ) -> pd.DataFrame: + payload = self._http.get_bytes( + url=url, + source=source, + artifact_name=artifact_name, + ) + return parse_csv_bytes(payload) + + def _snapshot_date(self, q: Query) -> pd.Timestamp: + return self._asof_utc(q).normalize() diff --git a/alphaforge/data/source.py b/alphaforge/data/source.py index 2ad6941..958e0a0 100644 --- a/alphaforge/data/source.py +++ b/alphaforge/data/source.py @@ -1,3 +1,13 @@ +"""Legacy DataSource protocol for compatibility and raw-loader workflows. + +New external loading code should prefer ``SourceAdapter`` plus +``DataContext.fetch(...)`` / ``fetch_many(...)`` / ``prefetch(...)``. + +``DataSource`` remains useful where Alphaforge still exposes raw long-frame +loaders directly or where existing panel-oriented integrations have not been +migrated yet. +""" + from typing import Protocol import pandas as pd @@ -7,6 +17,8 @@ class DataSource(Protocol): + """Compatibility/raw-loader protocol, not the canonical fetch contract.""" + name: str def schemas(self) -> dict[str, TableSchema]: ... diff --git a/alphaforge/data/sources/__init__.py b/alphaforge/data/sources/__init__.py index 0b22b8c..d3c9ce2 100644 --- a/alphaforge/data/sources/__init__.py +++ b/alphaforge/data/sources/__init__.py @@ -7,7 +7,7 @@ from __future__ import annotations import logging -from typing import TYPE_CHECKING +from typing import TYPE_CHECKING, Any if TYPE_CHECKING: from alphaforge.data.adapter import SourceAdapter @@ -32,16 +32,16 @@ def discover_adapters() -> dict[str, type[SourceAdapter]]: if sys.version_info >= (3, 12): from importlib.metadata import entry_points - eps = entry_points(group="alphaforge.source_adapters") + eps: list[Any] = list(entry_points(group="alphaforge.source_adapters")) else: # Python 3.10–3.11 compat from importlib.metadata import entry_points as _ep all_eps = _ep() if isinstance(all_eps, dict): - eps = all_eps.get("alphaforge.source_adapters", []) + eps = list(all_eps.get("alphaforge.source_adapters", [])) else: - eps = all_eps.select(group="alphaforge.source_adapters") + eps = list(all_eps.select(group="alphaforge.source_adapters")) result: dict[str, type[SourceAdapter]] = {} for ep in eps: diff --git a/alphaforge/data/sources/cftc.py b/alphaforge/data/sources/cftc.py index 2890e75..c2f0bcf 100644 --- a/alphaforge/data/sources/cftc.py +++ b/alphaforge/data/sources/cftc.py @@ -1,15 +1,15 @@ """CFTCAdapter — Bulk SourceAdapter for CFTC CoT data. -Fetches from CFTC (via CFTCCoTSource), transforms with cot_to_pit_observations, -and caches the full bulk result. Subsequent queries for individual series -are served from cache. +Fetches from CFTC CoT public-web sources, transforms with +``cot_to_pit_observations``, and caches the full bulk result. Subsequent +queries for individual series are served from cache. """ from __future__ import annotations import logging from datetime import date, datetime, timedelta, timezone -from typing import Callable, Optional +from typing import Callable, Mapping, Optional import duckdb import pandas as pd @@ -22,16 +22,18 @@ logger = logging.getLogger(__name__) +_DATASET_SOURCE_NAMES = { + "cot.tff": "cftc_cot", + "cot.disagg": "cftc_cot_disagg", +} + class CFTCAdapter(SourceAdapterBase): """Cache-aware bulk adapter for CFTC Commitments of Traders data. - Parameters - ---------- - raw_fetcher : callable(start, end) -> DataFrame - Function that fetches raw CoT data (e.g. CFTCCoTSource.fetch wrapper). - cache_conn : duckdb.DuckDBPyConnection | None - DuckDB connection for caching. + Each dataset is fetched in bulk through a dataset-specific callable. + The default backward-compatible constructor still accepts a single + ``raw_fetcher`` for ``cot.tff``. """ source_name = "cftc" @@ -39,14 +41,50 @@ class CFTCAdapter(SourceAdapterBase): def __init__( self, - raw_fetcher: Callable[[Optional[pd.Timestamp], Optional[pd.Timestamp]], pd.DataFrame], + raw_fetcher: Callable[[Optional[pd.Timestamp], Optional[pd.Timestamp]], pd.DataFrame] + | None = None, + *, + raw_fetchers: Mapping[ + str, Callable[[Optional[pd.Timestamp], Optional[pd.Timestamp]], pd.DataFrame] + ] + | None = None, cache_conn: Optional[duckdb.DuckDBPyConnection] = None, ) -> None: - self._raw_fetcher = raw_fetcher + fetchers = dict(raw_fetchers or {}) + if raw_fetcher is not None: + fetchers.setdefault("cot.tff", raw_fetcher) + if not fetchers: + raise ValueError("CFTCAdapter requires at least one dataset raw_fetcher") + + self._raw_fetchers = fetchers + self._raw_fetcher = fetchers.get("cot.tff") + self.datasets = frozenset(fetchers) self._cache: CacheLayer | None = None if cache_conn is not None: self._cache = CacheLayer(cache_conn) + def _raw_fetch( + self, + dataset: str, + start: Optional[pd.Timestamp], + end: Optional[pd.Timestamp], + ) -> pd.DataFrame: + if dataset == "cot.tff" and self._raw_fetcher is not None: + return self._raw_fetcher(start, end) + if dataset not in self._raw_fetchers: + raise KeyError( + f"CFTCAdapter has no raw fetcher for dataset '{dataset}'. " + f"Available: {sorted(self._raw_fetchers)}" + ) + return self._raw_fetchers[dataset](start, end) + + def _to_pit(self, dataset: str, raw_df: pd.DataFrame) -> pd.DataFrame: + return cot_to_pit_observations( + raw_df, + key_prefix=f"cftc.{dataset}.", + source_name=_DATASET_SOURCE_NAMES.get(dataset, "cftc_cot"), + ) + def fetch( self, query: Query, @@ -54,6 +92,7 @@ def fetch( max_staleness: Optional[timedelta] = None, ) -> FetchResult: """Fetch CoT data. First call triggers bulk fetch; subsequent use cache.""" + dataset = query.table entities = list(query.entities or []) # Try cache for all requested entities @@ -65,7 +104,7 @@ def fetch( for series_key in entities: result = self._cache.lookup( series_key=series_key, - dataset="cot.tff", + dataset=dataset, source=self.source_name, is_pit=True, max_staleness=max_staleness, @@ -79,6 +118,7 @@ def fetch( if all_cached and cached_frames: combined = pd.concat(cached_frames, ignore_index=True) + combined["source"] = _DATASET_SOURCE_NAMES.get(dataset, "cftc_cot") return FetchResult( data=combined, source=self.source_name, @@ -88,14 +128,14 @@ def fetch( ) # Cache miss — bulk fetch + transform - raw_df = self._raw_fetcher(query.start, query.end) - pit_df = cot_to_pit_observations(raw_df) + raw_df = self._raw_fetch(dataset, query.start, query.end) + pit_df = self._to_pit(dataset, raw_df) # Cache the full result if self._cache is not None and not pit_df.empty: self._cache.store( pit_df, - dataset="cot.tff", + dataset=dataset, source=self.source_name, is_pit=True, ) @@ -121,17 +161,17 @@ def prefetch( start = pd.Timestamp(asof_range[0]) if asof_range else None end = pd.Timestamp(asof_range[1]) if asof_range else None - raw_df = self._raw_fetcher(start, end) - pit_df = cot_to_pit_observations(raw_df) + raw_df = self._raw_fetch(dataset, start, end) + pit_df = self._to_pit(dataset, raw_df) if self._cache is not None and not pit_df.empty: self._cache.store( pit_df, - dataset="cot.tff", + dataset=dataset, source=self.source_name, is_pit=True, ) - manifest = self._cache.get_manifest(dataset="cot.tff", source=self.source_name) + manifest = self._cache.get_manifest(dataset=dataset, source=self.source_name) if manifest is not None: return manifest diff --git a/alphaforge/data/sources/dtcc.py b/alphaforge/data/sources/dtcc.py index 65d2321..c5a1b24 100644 --- a/alphaforge/data/sources/dtcc.py +++ b/alphaforge/data/sources/dtcc.py @@ -1,61 +1,127 @@ -"""DTCCAdapter — Bulk SourceAdapter for DTCC PPD data. +"""DTCC adapters over the low-level DTCC PPD raw loader. -Fetches from DTCC (via DTCCPPDSource), transforms with -dtcc_daily_to_pit_observations, and caches the full bulk result. +`DTCCPPDSource` remains the provider-specific raw-loader implementation. +The adapter layer owns canonical `SourceAdapter` routing, cache/prefetch +behavior, and PIT transform wiring on top of that raw loader. """ from __future__ import annotations -import logging from datetime import date, datetime, timedelta, timezone -from typing import Callable, Optional +from typing import Any, Optional, cast import duckdb import pandas as pd from ..adapter import SourceAdapterBase from ..cache_layer import CacheLayer +from ..public_web.dtcc_ppd import DTCCPPDSource from ..query import Query +from ..source import DataSource from ..transforms.dtcc_pit import dtcc_daily_to_pit_observations from ..types import CacheManifest, FetchResult -logger = logging.getLogger(__name__) +_DEFAULT_RAW_COLUMNS = ( + "trade_count", + "notional_sum", + "price_mean", + "price_std", + "notional_median", + "trade_count_large", + "dv01_proxy_sum", +) -class DTCCAdapter(SourceAdapterBase): - """Cache-aware bulk adapter for DTCC PPD data. +def _build_dtcc_source( + *, + source: DataSource | None, + source_kwargs: dict[str, Any], + default_asset_codes: tuple[str, ...] | None = None, +) -> DataSource: + if source is not None and source_kwargs: + raise ValueError( + "Pass either a preconfigured source or DTCCPPDSource keyword " + "arguments, not both." + ) + + if source is not None: + return source + + resolved_kwargs = dict(source_kwargs) + if default_asset_codes is not None: + resolved_kwargs.setdefault("asset_codes", default_asset_codes) + source_factory = cast(Any, DTCCPPDSource) + return source_factory(**resolved_kwargs) + - Parameters - ---------- - raw_fetcher : callable(start, end) -> DataFrame - Function that fetches raw DTCC data. - cache_conn : duckdb.DuckDBPyConnection | None - DuckDB connection for caching. +class DTCCPPDAdapterBase(SourceAdapterBase): + """Shared cache-aware adapter plumbing over `DTCCPPDSource`. + + Subclasses define the adapter-facing dataset contract and the PIT + transform, while this base owns raw-loader fetch construction and cache + lifecycle management. """ - source_name = "dtcc" - datasets = frozenset({"dtcc.ppd"}) + source_name: str + datasets: frozenset[str] + raw_table = DTCCPPDSource.DAILY_TABLE + raw_columns = _DEFAULT_RAW_COLUMNS def __init__( self, - raw_fetcher: Callable[[Optional[pd.Timestamp], Optional[pd.Timestamp]], pd.DataFrame], + *, + source: DataSource, cache_conn: Optional[duckdb.DuckDBPyConnection] = None, ) -> None: - self._raw_fetcher = raw_fetcher + self._source = source self._cache: CacheLayer | None = None if cache_conn is not None: self._cache = CacheLayer(cache_conn) + def _require_dataset(self, dataset: str) -> None: + if dataset not in self.datasets: + raise KeyError( + f"{type(self).__name__} does not serve dataset '{dataset}'. " + f"Available: {sorted(self.datasets)}" + ) + + def _raw_table_for_dataset(self, dataset: str) -> str: + self._require_dataset(dataset) + return self.raw_table + + def _raw_columns_for_dataset(self, dataset: str) -> tuple[str, ...]: + self._require_dataset(dataset) + return self.raw_columns + + def _raw_fetch( + self, + dataset: str, + start: Optional[pd.Timestamp], + end: Optional[pd.Timestamp], + ) -> pd.DataFrame: + return self._source.fetch( + Query( + table=self._raw_table_for_dataset(dataset), + columns=list(self._raw_columns_for_dataset(dataset)), + start=start, + end=end, + ) + ) + + def _to_pit(self, dataset: str, raw_df: pd.DataFrame) -> pd.DataFrame: + raise NotImplementedError + def fetch( self, query: Query, *, max_staleness: Optional[timedelta] = None, ) -> FetchResult: - """Fetch DTCC data. Bulk fetch on miss, serve from cache on hit.""" + """Fetch DTCC data through the canonical adapter contract.""" + dataset = query.table + self._require_dataset(dataset) entities = list(query.entities or []) - # Try cache if entities and self._cache is not None: cached_frames = [] all_cached = True @@ -64,7 +130,7 @@ def fetch( for series_key in entities: result = self._cache.lookup( series_key=series_key, - dataset="dtcc.ppd", + dataset=dataset, source=self.source_name, is_pit=True, max_staleness=max_staleness, @@ -81,32 +147,29 @@ def fetch( return FetchResult( data=combined, source=self.source_name, - dataset=query.table, + dataset=dataset, is_pit=True, cached_at=cached_at, ) - # Bulk fetch + transform - raw_df = self._raw_fetcher(query.start, query.end) - pit_df = dtcc_daily_to_pit_observations(raw_df) + raw_df = self._raw_fetch(dataset, query.start, query.end) + pit_df = self._to_pit(dataset, raw_df) - # Cache if self._cache is not None and not pit_df.empty: self._cache.store( pit_df, - dataset="dtcc.ppd", + dataset=dataset, source=self.source_name, is_pit=True, ) - # Filter if entities and not pit_df.empty: pit_df = pit_df[pit_df["series_key"].isin(entities)] return FetchResult( data=pit_df, source=self.source_name, - dataset=query.table, + dataset=dataset, is_pit=True, cached_at=None, ) @@ -116,21 +179,22 @@ def prefetch( dataset: str, asof_range: tuple[date, date] | None = None, ) -> CacheManifest: - """Bulk fetch and cache all DTCC data.""" + """Bulk fetch and cache all rows for an adapter dataset.""" + self._require_dataset(dataset) start = pd.Timestamp(asof_range[0]) if asof_range else None end = pd.Timestamp(asof_range[1]) if asof_range else None - raw_df = self._raw_fetcher(start, end) - pit_df = dtcc_daily_to_pit_observations(raw_df) + raw_df = self._raw_fetch(dataset, start, end) + pit_df = self._to_pit(dataset, raw_df) if self._cache is not None and not pit_df.empty: self._cache.store( pit_df, - dataset="dtcc.ppd", + dataset=dataset, source=self.source_name, is_pit=True, ) - manifest = self._cache.get_manifest(dataset="dtcc.ppd", source=self.source_name) + manifest = self._cache.get_manifest(dataset=dataset, source=self.source_name) if manifest is not None: return manifest @@ -144,9 +208,99 @@ def prefetch( ) def list_entities(self, dataset: str) -> list[str]: - """List cached entity keys.""" + """List cached entity keys for a DTCC adapter dataset.""" + self._require_dataset(dataset) if self._cache is not None: manifest = self._cache.get_manifest(dataset=dataset, source=self.source_name) if manifest is not None: return manifest.entity_keys return [] + + +class DTCCAdapter(DTCCPPDAdapterBase): + """Built-in canonical adapter for the generic DTCC PPD dataset. + + Callers no longer need to inject a raw fetch function. By default this + adapter constructs and owns a `DTCCPPDSource` instance internally. + Tests and advanced callers may still inject a preconfigured raw source. + """ + + source_name = "dtcc" + datasets = frozenset({"dtcc.ppd"}) + + def __init__( + self, + *, + source: DataSource | None = None, + cache_conn: Optional[duckdb.DuckDBPyConnection] = None, + **source_kwargs: Any, + ) -> None: + resolved_source = _build_dtcc_source( + source=source, + source_kwargs=source_kwargs, + ) + super().__init__(source=resolved_source, cache_conn=cache_conn) + + def _to_pit(self, dataset: str, raw_df: pd.DataFrame) -> pd.DataFrame: + self._require_dataset(dataset) + return dtcc_daily_to_pit_observations(raw_df) + + +class _FilteredDTCCPPDAdapter(DTCCPPDAdapterBase): + """Shared base for concrete DTCC product-family adapters.""" + + key_prefix: str + pit_source_name: str + entity_prefix: str + default_asset_codes: tuple[str, ...] | None = None + + def __init__( + self, + *, + source: DataSource | None = None, + cache_conn: Optional[duckdb.DuckDBPyConnection] = None, + **source_kwargs: Any, + ) -> None: + resolved_source = _build_dtcc_source( + source=source, + source_kwargs=source_kwargs, + default_asset_codes=self.default_asset_codes, + ) + super().__init__(source=resolved_source, cache_conn=cache_conn) + + def _filter_raw_df(self, raw_df: pd.DataFrame) -> pd.DataFrame: + if raw_df.empty or "entity_id" not in raw_df.columns: + return raw_df + entity_ids = raw_df["entity_id"].astype(str) + return raw_df[entity_ids.str.startswith(self.entity_prefix)].copy() + + def _to_pit(self, dataset: str, raw_df: pd.DataFrame) -> pd.DataFrame: + self._require_dataset(dataset) + filtered = self._filter_raw_df(raw_df) + return dtcc_daily_to_pit_observations( + filtered, + key_prefix=self.key_prefix, + source_name=self.pit_source_name, + ) + + +class DTCCFXAdapter(_FilteredDTCCPPDAdapter): + """Canonical DTCC adapter for FX forwards and swaps.""" + + source_name = "dtcc_fx" + datasets = frozenset({"dtcc.fx"}) + key_prefix = "dtcc.fx." + pit_source_name = "dtcc_ppd_fx" + entity_prefix = "dtccppd.fx." + default_asset_codes = ("FX",) + + +class DTCCIRSAdapter(_FilteredDTCCPPDAdapter): + """Canonical DTCC adapter for interest rate swaps.""" + + source_name = "dtcc_irs" + datasets = frozenset({"dtcc.irs"}) + key_prefix = "dtcc.irs." + pit_source_name = "dtcc_ppd_irs" + entity_prefix = "dtccppd.rates.interest_rate_swap." + default_asset_codes = ("IR",) diff --git a/alphaforge/data/transforms/cot_pit.py b/alphaforge/data/transforms/cot_pit.py index 7d79c48..c900873 100644 --- a/alphaforge/data/transforms/cot_pit.py +++ b/alphaforge/data/transforms/cot_pit.py @@ -24,15 +24,20 @@ def cot_to_pit_observations( df: pd.DataFrame, metrics: Sequence[str] | None = None, + *, + key_prefix: str = "cftc.cot.tff.", + source_name: str = "cftc_cot", ) -> pd.DataFrame: """Convert CFTC CoT fetch output to PIT observation rows. Parameters ---------- - df : DataFrame from CFTCCoTSource.fetch() with columns: + df : DataFrame from a CFTC CoT source fetch() with columns: date (publication date, Friday), entity_id, long_positions, short_positions, open_interest. metrics : which series to emit. Defaults to all five. + key_prefix : Prefix for generated ``series_key`` values. + source_name : Value for the output ``source`` column. Returns ------- @@ -66,6 +71,6 @@ def cot_to_pit_observations( obs_date_col="obs_date", asof_col="asof_utc", value_vars=chosen, - key_prefix="cftc.cot.tff.", - source_name="cftc_cot", + key_prefix=key_prefix, + source_name=source_name, ) diff --git a/alphaforge/data/transforms/dtcc_pit.py b/alphaforge/data/transforms/dtcc_pit.py index 9fe7448..937aafc 100644 --- a/alphaforge/data/transforms/dtcc_pit.py +++ b/alphaforge/data/transforms/dtcc_pit.py @@ -24,6 +24,9 @@ def dtcc_daily_to_pit_observations( df: pd.DataFrame, metrics: Sequence[str] | None = None, + *, + key_prefix: str = "dtcc.ppd.daily.", + source_name: str = "dtcc_ppd", ) -> pd.DataFrame: """Convert DTCC PPD daily fetch output to PIT observation rows. @@ -32,6 +35,8 @@ def dtcc_daily_to_pit_observations( df : DataFrame from DTCCPPDSource.fetch(table="dtcc.ppd.daily") with columns: date, entity_id, asof_utc, trade_count, notional_sum, etc. metrics : which series to emit. Defaults to all five. + key_prefix : Prefix for the emitted series keys. + source_name : Value for the PIT lineage ``source`` column. Returns ------- @@ -53,6 +58,6 @@ def dtcc_daily_to_pit_observations( obs_date_col="obs_date", asof_col="asof_utc", value_vars=chosen, - key_prefix="dtcc.ppd.daily.", - source_name="dtcc_ppd", + key_prefix=key_prefix, + source_name=source_name, ) diff --git a/alphaforge/features/__init__.py b/alphaforge/features/__init__.py index f5418da..59b18ba 100644 --- a/alphaforge/features/__init__.py +++ b/alphaforge/features/__init__.py @@ -2,6 +2,7 @@ from .calendar_flags import CalendarFlagsTemplate from .event_dates import EventDateTemplate from .frame import FeatureFrame +from .market import LagReturnsTemplate, RollingVolatilityTemplate from .template import FeatureTemplate, ParamSpec, SliceSpec __all__ = [ @@ -9,6 +10,8 @@ "EventDateTemplate", "FeatureFrame", "FeatureTemplate", + "LagReturnsTemplate", "ParamSpec", + "RollingVolatilityTemplate", "SliceSpec", ] diff --git a/alphaforge/features/dataset_builder.py b/alphaforge/features/dataset_builder.py index 9b2dc3f..61c90ca 100644 --- a/alphaforge/features/dataset_builder.py +++ b/alphaforge/features/dataset_builder.py @@ -106,6 +106,25 @@ def _materialize_template( return template.transform(ctx, params, slice, state) +def _annotate_feature_frame_request(ff: FeatureFrame, req) -> FeatureFrame: + """Stamp request-level composition metadata into the feature catalog.""" + if ff.catalog is None or ff.catalog.empty: + return ff + + catalog = ff.catalog.copy() + if req.key is not None: + catalog["request_key"] = req.key + + template_name = getattr(req.template, "name", req.template.__class__.__name__) + catalog["template_name"] = template_name + template_version = getattr(req.template, "version", None) + if template_version is not None: + catalog["template_version"] = template_version + + ff.catalog = catalog + return ff + + def _materialize_target( ctx, target: TargetRequest, @@ -228,9 +247,10 @@ def build_dataset( # 1) materialize features feature_frames: List[FeatureFrame] = [] - for req in spec.features: + for req in spec.feature_requests(): s = _apply_override(base_slice, req.slice_override) ff = _materialize_template(ctx, req.template, req.params, s) + ff = _annotate_feature_frame_request(ff, req) if req.tags: ff = ff.set_tags(req.tags, overwrite=False) diff --git a/alphaforge/features/dataset_spec.py b/alphaforge/features/dataset_spec.py index 3a9c341..ab98a91 100644 --- a/alphaforge/features/dataset_spec.py +++ b/alphaforge/features/dataset_spec.py @@ -1,7 +1,7 @@ # alphaforge/features/dataset_spec.py from __future__ import annotations -from dataclasses import dataclass, field +from dataclasses import dataclass, field, replace from typing import Any, Dict, Optional, Sequence import pandas as pd @@ -26,6 +26,10 @@ class TimeSpec: grid: str = "B" # "B" daily business day grid; later can be richer asof: Optional[pd.Timestamp] = None # optional global asof cut (PIT); can be None + def __post_init__(self) -> None: + if pd.Timestamp(self.start) > pd.Timestamp(self.end): + raise ValueError("TimeSpec.start must be <= TimeSpec.end.") + @dataclass(frozen=True) class SliceOverride: @@ -41,6 +45,28 @@ class SliceOverride: asof: Optional[pd.Timestamp] = None +def _merge_slice_overrides( + parent: Optional["SliceOverride"], + child: Optional["SliceOverride"], +) -> Optional["SliceOverride"]: + if parent is None: + return child + if child is None: + return parent + return SliceOverride( + lookback=child.lookback if child.lookback is not None else parent.lookback, + grid=child.grid if child.grid is not None else parent.grid, + asof=child.asof if child.asof is not None else parent.asof, + ) + + +def _compose_request_key(parent: Optional[str], child: Optional[str]) -> Optional[str]: + pieces = [piece for piece in (parent, child) if piece] + if not pieces: + return None + return "/".join(pieces) + + @dataclass(frozen=True) class FeatureRequest: """ @@ -56,6 +82,22 @@ class FeatureRequest: tags: Dict[str, Any] = field(default_factory=dict) +@dataclass(frozen=True) +class FeatureRequestGroup: + """Composable group of feature requests with inherited metadata.""" + + requests: Sequence[FeatureRequest | "FeatureRequestGroup"] = field( + default_factory=tuple + ) + slice_override: Optional[SliceOverride] = None + key: Optional[str] = None + tags: Dict[str, Any] = field(default_factory=dict) + + def __post_init__(self) -> None: + if not self.requests: + raise ValueError("FeatureRequestGroup.requests cannot be empty.") + + @dataclass(frozen=True) class TargetRequest: """ @@ -78,6 +120,12 @@ class JoinPolicy: how: str = "inner" # "inner" safest; "outer" allowed sort_index: bool = True + def __post_init__(self) -> None: + if self.how not in {"inner", "outer"}: + raise ValueError( + f"JoinPolicy.how must be 'inner' or 'outer', got {self.how!r}." + ) + @dataclass(frozen=True) class MissingnessPolicy: @@ -87,13 +135,73 @@ class MissingnessPolicy: final_row_policy: str = "drop_if_any_nan" # or "keep" + def __post_init__(self) -> None: + if self.final_row_policy not in {"drop_if_any_nan", "keep"}: + raise ValueError( + "MissingnessPolicy.final_row_policy must be " + f"'drop_if_any_nan' or 'keep', got {self.final_row_policy!r}." + ) + + +def _flatten_feature_requests( + features: Sequence[Any], + *, + inherited_slice_override: Optional[SliceOverride] = None, + inherited_key: Optional[str] = None, + inherited_tags: Optional[Dict[str, Any]] = None, +) -> list[FeatureRequest]: + flat: list[FeatureRequest] = [] + base_tags = dict(inherited_tags or {}) + + for item in features: + if isinstance(item, FeatureRequest): + merged_tags = dict(base_tags) + merged_tags.update(item.tags) + flat.append( + replace( + item, + slice_override=_merge_slice_overrides( + inherited_slice_override, + item.slice_override, + ), + key=_compose_request_key(inherited_key, item.key), + tags=merged_tags, + ) + ) + continue + + if isinstance(item, FeatureRequestGroup): + nested_tags = dict(base_tags) + nested_tags.update(item.tags) + flat.extend( + _flatten_feature_requests( + item.requests, + inherited_slice_override=_merge_slice_overrides( + inherited_slice_override, + item.slice_override, + ), + inherited_key=_compose_request_key(inherited_key, item.key), + inherited_tags=nested_tags, + ) + ) + continue + + raise TypeError( + "DatasetSpec.features must contain FeatureRequest or " + f"FeatureRequestGroup items, got {type(item)!r}." + ) + + return flat + @dataclass(frozen=True) class DatasetSpec: universe: UniverseSpec time: TimeSpec target: TargetRequest - features: Sequence[FeatureRequest] = field(default_factory=list) + features: Sequence[FeatureRequest | FeatureRequestGroup] = field( + default_factory=list + ) join_policy: JoinPolicy = field(default_factory=JoinPolicy) missingness: MissingnessPolicy = field(default_factory=MissingnessPolicy) @@ -101,6 +209,10 @@ class DatasetSpec: name: str = "dataset" tags: Dict[str, Any] = field(default_factory=dict) + def feature_requests(self) -> list[FeatureRequest]: + """Return the flattened feature-request list used by the builder.""" + return _flatten_feature_requests(self.features) + @dataclass class DatasetArtifact: diff --git a/alphaforge/features/market.py b/alphaforge/features/market.py new file mode 100644 index 0000000..33295b1 --- /dev/null +++ b/alphaforge/features/market.py @@ -0,0 +1,280 @@ +from __future__ import annotations + +import math +from typing import Any, Sequence + +import numpy as np +import pandas as pd + +from .frame import FeatureFrame +from .ids import group_path, make_feature_id +from .template import ParamSpec, SliceSpec + + +def _coerce_windows(value: Any, *, name: str, allow_zero: bool = False) -> list[int]: + if isinstance(value, int): + items = [int(value)] + elif isinstance(value, Sequence) and not isinstance(value, (str, bytes)): + items = [int(item) for item in value] + else: + raise TypeError(f"{name} must be an int or sequence of ints.") + + if not items: + raise ValueError(f"{name} cannot be empty.") + + minimum = 0 if allow_zero else 1 + for item in items: + if item < minimum: + qualifier = "non-negative" if allow_zero else "positive" + raise ValueError(f"{name} must contain only {qualifier} integers.") + return list(dict.fromkeys(items)) + + +def _coerce_market_frame( + ctx, + *, + dataset: str, + source: str | None, + price_col: str, + slice: SliceSpec, +) -> pd.DataFrame: + result = ctx.load( + dataset, + columns=[price_col], + start=slice.start, + end=slice.end, + entities=slice.entities, + asof=slice.asof, + grid=slice.grid, + source=source, + ) + frame = result.data.copy() + empty_index = pd.MultiIndex.from_arrays( + [pd.DatetimeIndex([], tz="UTC"), pd.Index([], dtype="object")], + names=["ts_utc", "entity_id"], + ) + if frame.empty: + return pd.DataFrame({price_col: pd.Series(dtype="float64")}, index=empty_index) + + missing = {"obs_date", "series_key", price_col} - set(frame.columns) + if missing: + raise ValueError( + f"Market template fetch for '{dataset}' is missing required columns: " + f"{sorted(missing)}" + ) + + frame = frame.loc[:, ["obs_date", "series_key", price_col]].copy() + frame["ts_utc"] = pd.to_datetime(frame["obs_date"], utc=True) + frame["entity_id"] = frame["series_key"].astype(str) + frame[price_col] = pd.to_numeric(frame[price_col], errors="coerce") + + if slice.start is not None: + frame = frame[frame["ts_utc"] >= pd.Timestamp(slice.start).tz_convert("UTC")] + if slice.end is not None: + frame = frame[frame["ts_utc"] <= pd.Timestamp(slice.end).tz_convert("UTC")] + if slice.asof is not None: + frame = frame[frame["ts_utc"] <= pd.Timestamp(slice.asof).tz_convert("UTC")] + if slice.entities is not None: + allowed = {str(entity) for entity in slice.entities} + frame = frame[frame["entity_id"].isin(allowed)] + + out = frame.set_index(["ts_utc", "entity_id"])[[price_col]].sort_index() + return out[~out.index.duplicated(keep="last")] + + +def _compute_returns(prices: pd.Series, *, return_kind: str) -> pd.Series: + if return_kind == "log": + base = np.log(prices.astype(float)) + return base.groupby(level="entity_id").diff() + if return_kind == "simple": + return prices.astype(float).groupby(level="entity_id").pct_change() + raise ValueError("return_kind must be 'log' or 'simple'.") + + +class LagReturnsTemplate: + """Lagged return features from a canonical market-price dataset.""" + + name = "lag_returns" + version = "1.0" + param_space = { + "lags": ParamSpec("categorical", default=(1, 5, 21)), + "price_col": ParamSpec("categorical", default="close"), + "dataset": ParamSpec("categorical", default="market.ohlcv"), + "source": ParamSpec("categorical", default=None), + "return_kind": ParamSpec( + "categorical", default="log", choices=["log", "simple"] + ), + } + + def requires(self, params): + return [] + + def fit(self, ctx, params, fit_slice): + return None + + def transform(self, ctx, params, slice: SliceSpec, state): + dataset = str(params.get("dataset", "market.ohlcv")) + source = params.get("source") + price_col = str(params.get("price_col", "close")) + return_kind = str(params.get("return_kind", "log")).lower() + lags = _coerce_windows(params.get("lags", (1, 5, 21)), name="lags") + + prices = _coerce_market_frame( + ctx, + dataset=dataset, + source=source, + price_col=price_col, + slice=slice, + )[price_col] + returns = _compute_returns(prices, return_kind=return_kind) + + features: dict[str, pd.Series] = {} + catalog_rows: list[dict[str, Any]] = [] + for lag in lags: + feature_id = make_feature_id( + dataset, + "*", + "market", + f"{return_kind}_return_lag", + {"lag": lag, "price_col": price_col}, + ) + features[feature_id] = returns.groupby(level="entity_id").shift(lag) + catalog_rows.append( + { + "feature_id": feature_id, + "group_path": group_path( + "market", + "lag_returns", + {"lags": tuple(lags), "price_col": price_col}, + ), + "family": "market", + "transform": "lag_return", + "source_table": dataset, + "source_name": source, + "price_col": price_col, + "return_kind": return_kind, + "lag": lag, + } + ) + + X = pd.DataFrame(features, index=prices.index).sort_index() + catalog = pd.DataFrame(catalog_rows).set_index("feature_id").sort_index() + return FeatureFrame( + X=X, + catalog=catalog, + meta={ + "template": self.name, + "version": self.version, + "dataset": dataset, + "source": source, + }, + ) + + +class RollingVolatilityTemplate: + """Rolling realized-volatility features from canonical market-price data.""" + + name = "rolling_volatility" + version = "1.0" + param_space = { + "windows": ParamSpec("categorical", default=(5, 21, 63)), + "lag": ParamSpec("int", default=1, low=0), + "price_col": ParamSpec("categorical", default="close"), + "dataset": ParamSpec("categorical", default="market.ohlcv"), + "source": ParamSpec("categorical", default=None), + "return_kind": ParamSpec( + "categorical", default="log", choices=["log", "simple"] + ), + "annualization_factor": ParamSpec("int", default=252, low=1), + "min_periods": ParamSpec("int", default=None, low=1), + } + + def requires(self, params): + return [] + + def fit(self, ctx, params, fit_slice): + return None + + def transform(self, ctx, params, slice: SliceSpec, state): + dataset = str(params.get("dataset", "market.ohlcv")) + source = params.get("source") + price_col = str(params.get("price_col", "close")) + return_kind = str(params.get("return_kind", "log")).lower() + windows = _coerce_windows(params.get("windows", (5, 21, 63)), name="windows") + lag = _coerce_windows(params.get("lag", 1), name="lag", allow_zero=True)[0] + annualization_factor = float(params.get("annualization_factor", 252)) + min_periods_param = params.get("min_periods") + + prices = _coerce_market_frame( + ctx, + dataset=dataset, + source=source, + price_col=price_col, + slice=slice, + )[price_col] + returns = _compute_returns(prices, return_kind=return_kind) + + features: dict[str, pd.Series] = {} + catalog_rows: list[dict[str, Any]] = [] + for window in windows: + min_periods = window if min_periods_param is None else int(min_periods_param) + realized_vol = returns.groupby(level="entity_id").transform( + lambda values: values.rolling( + window=window, + min_periods=min_periods, + ).std() + ) + if lag: + realized_vol = realized_vol.groupby(level="entity_id").shift(lag) + if annualization_factor != 1.0: + realized_vol = realized_vol * math.sqrt(annualization_factor) + + feature_id = make_feature_id( + dataset, + "*", + "market", + f"{return_kind}_realized_volatility", + { + "window": window, + "lag": lag, + "price_col": price_col, + "annualization_factor": annualization_factor, + }, + ) + features[feature_id] = realized_vol + catalog_rows.append( + { + "feature_id": feature_id, + "group_path": group_path( + "market", + "rolling_volatility", + { + "window": window, + "lag": lag, + "price_col": price_col, + }, + ), + "family": "market", + "transform": "rolling_volatility", + "source_table": dataset, + "source_name": source, + "price_col": price_col, + "return_kind": return_kind, + "window": window, + "lag": lag, + "annualization_factor": annualization_factor, + } + ) + + X = pd.DataFrame(features, index=prices.index).sort_index() + catalog = pd.DataFrame(catalog_rows).set_index("feature_id").sort_index() + return FeatureFrame( + X=X, + catalog=catalog, + meta={ + "template": self.name, + "version": self.version, + "dataset": dataset, + "source": source, + }, + ) diff --git a/alphaforge/features/ops.py b/alphaforge/features/ops.py index 3cd692f..d0f3c30 100644 --- a/alphaforge/features/ops.py +++ b/alphaforge/features/ops.py @@ -5,12 +5,19 @@ from ..data.context import DataContext from ..store.cache import MaterializationPolicy +from ..store.store import Store from .dag import LineageGraph from .frame import FeatureFrame from .realization import FeatureRealization from .template import FeatureTemplate, SliceSpec +def _require_store(ctx: DataContext) -> Store: + if ctx.store is None: + raise ValueError("Feature materialization requires a configured store.") + return ctx.store + + def materialize( ctx: DataContext, template: FeatureTemplate, @@ -21,6 +28,7 @@ def materialize( ) -> FeatureFrame: """Materialize a FeatureRealization with caching and (optional) stateful fit.""" rid = realization.id() + store = _require_store(ctx) if lineage is not None: lineage.add( @@ -34,7 +42,7 @@ def materialize( }, ) - got = ctx.store.get_frame(rid) + got = store.get_frame(rid) if got is not None: return got @@ -46,7 +54,7 @@ def materialize( st = None if st is not None: - art = ctx.store.put_state( + art = store.put_state( st, pickle.dumps(st, protocol=pickle.HIGHEST_PROTOCOL) ) if lineage is not None: @@ -69,7 +77,7 @@ def materialize( frame.validate() if policy.persist_mode == "always": - ctx.store.put_frame(rid, frame) + store.put_frame(rid, frame) return frame diff --git a/alphaforge/futures/config.py b/alphaforge/futures/config.py index 0185c0c..75cca64 100644 --- a/alphaforge/futures/config.py +++ b/alphaforge/futures/config.py @@ -149,6 +149,8 @@ def resolve( "Missing required futures configuration values. " f"Set them explicitly, via YAML, or via env vars: {joined}" ) + assert resolved_source_dir is not None + assert resolved_artifact_root is not None return cls( source_dir=resolved_source_dir.resolve(), diff --git a/alphaforge/futures/loader.py b/alphaforge/futures/loader.py index 30c2a3d..1a7fec9 100644 --- a/alphaforge/futures/loader.py +++ b/alphaforge/futures/loader.py @@ -2,9 +2,9 @@ from __future__ import annotations +import re from dataclasses import dataclass from pathlib import Path -import re import pandas as pd @@ -56,6 +56,13 @@ def contract_sort_key(self) -> tuple[int, int]: return self.contract_year, self.contract_month +@dataclass(frozen=True) +class _RollEvent: + root_symbol: str + effective_session_date: pd.Timestamp + adjustment_factor: float + + def _parse_hhmm(value: str) -> tuple[int, int]: hour, minute = value.split(":") return int(hour), int(minute) @@ -542,8 +549,11 @@ def _build_continuous_eod( if contract_eod.empty or roll_schedule.empty: return pd.DataFrame() + def _as_float(value: object) -> float: + return float(pd.to_numeric(value, errors="raise")) + selected_frames: list[pd.DataFrame] = [] - roll_events: list[dict[str, object]] = [] + roll_events: list[_RollEvent] = [] for root_symbol, schedule in roll_schedule.groupby("root_symbol", sort=True): schedule = schedule.sort_values("start_session_date").reset_index(drop=True) @@ -561,15 +571,15 @@ def _build_continuous_eod( first_subset_row = subset.index.min() if pd.notna(row["rolled_from_contract_id"]): subset.loc[first_subset_row, "roll_flag"] = True - new_open = float(subset.iloc[0]["open"]) - prior_close = 1.0 if prior_end_row is None else float(prior_end_row["close"]) + new_open = _as_float(subset.iloc[0]["open"]) + prior_close = 1.0 if prior_end_row is None else _as_float(prior_end_row["close"]) factor = 1.0 if prior_close == 0 else new_open / prior_close roll_events.append( - { - "root_symbol": root_symbol, - "effective_session_date": pd.Timestamp(row["start_session_date"]), - "adjustment_factor": factor, - } + _RollEvent( + root_symbol=root_symbol, + effective_session_date=pd.Timestamp(row["start_session_date"]), + adjustment_factor=factor, + ) ) prior_end_row = subset.sort_values("session_date").iloc[-1] selected_frames.append(subset) @@ -584,15 +594,15 @@ def _build_continuous_eod( for event in roll_events: mask = ( - (continuous["root_symbol"] == event["root_symbol"]) - & (continuous["session_date"] < event["effective_session_date"]) + (continuous["root_symbol"] == event.root_symbol) + & (continuous["session_date"] < event.effective_session_date) ) - continuous.loc[mask, "cumulative_adjustment_factor"] *= float(event["adjustment_factor"]) + continuous.loc[mask, "cumulative_adjustment_factor"] *= event.adjustment_factor start_mask = ( - (continuous["root_symbol"] == event["root_symbol"]) - & (continuous["session_date"] == event["effective_session_date"]) + (continuous["root_symbol"] == event.root_symbol) + & (continuous["session_date"] == event.effective_session_date) ) - continuous.loc[start_mask, "adjustment_factor"] = float(event["adjustment_factor"]) + continuous.loc[start_mask, "adjustment_factor"] = event.adjustment_factor for column in ("open", "high", "low", "close"): continuous[f"raw_{column}"] = continuous[column] diff --git a/alphaforge/pipeline/__init__.py b/alphaforge/pipeline/__init__.py index d0c7b61..8199743 100644 --- a/alphaforge/pipeline/__init__.py +++ b/alphaforge/pipeline/__init__.py @@ -9,6 +9,8 @@ SourceHealthStatus, assess_health, assess_source_health, + build_health_report, + health_report, ) from .protocols import ( Filter, @@ -40,5 +42,7 @@ "HealthStatus", "HealthTracker", "assess_health", + "build_health_report", + "health_report", "adjust_weights_for_health", ] diff --git a/alphaforge/pipeline/health.py b/alphaforge/pipeline/health.py index b4045b4..e0efe2b 100644 --- a/alphaforge/pipeline/health.py +++ b/alphaforge/pipeline/health.py @@ -2,12 +2,13 @@ from __future__ import annotations from dataclasses import dataclass -from typing import TYPE_CHECKING +from typing import TYPE_CHECKING, Mapping, Sequence import pandas as pd +from pandas.tseries.offsets import MonthEnd, QuarterEnd, YearEnd if TYPE_CHECKING: - from alphaforge.pit.release_rules import ReleaseRule + from alphaforge.time.release_rules import ReleaseRule @dataclass(frozen=True) @@ -32,6 +33,8 @@ class SourceHealthStatus: asof: pd.Timestamp age: pd.Timedelta | None age_days: float | None + overdue: pd.Timedelta | None + overdue_days: float | None status: str # "ok", "late", "stale", "dead", "empty" weight_factor: float expected_next: pd.Timestamp | None @@ -52,6 +55,8 @@ def assess_source_health( asof=asof, age=None, age_days=None, + overdue=None, + overdue_days=None, status="empty", weight_factor=0.0, expected_next=None, @@ -68,45 +73,88 @@ def assess_source_health( age = _asof - _latest age_days = age.total_seconds() / 86400.0 - expected_next = _latest + policy.expected_cadence + expected_next = _resolve_expected_next_release(_latest, _asof, policy) + raw_overdue = _asof - expected_next + overdue = raw_overdue if raw_overdue > pd.Timedelta(0) else pd.Timedelta(0) + overdue_days = overdue.total_seconds() / 86400.0 - ok_limit = policy.expected_cadence + policy.grace_period - decay_start = ( - policy.weight_decay_start - if policy.weight_decay_start is not None - else policy.stale_threshold - ) - - if age <= ok_limit: - status = "ok" - weight = 1.0 - msg = f"{source_name}: OK (age {age_days:.0f}d)" - elif age <= policy.stale_threshold: - status = "late" - weight = 1.0 - msg = f"{source_name}: late (age {age_days:.0f}d, expected every {policy.expected_cadence.days}d)" - elif age <= policy.dead_threshold: - status = "stale" - decay_age = age - decay_start - hl = policy.weight_decay_half_life - weight = max(0.0, min(1.0, 2.0 ** (-decay_age / hl))) - msg = ( - f"{source_name}: stale (age {age_days:.0f}d, weight {weight:.2f})" + if policy.release_rule is not None: + decay_start = ( + policy.weight_decay_start + if policy.weight_decay_start is not None + else policy.stale_threshold ) + if overdue <= policy.grace_period: + status = "ok" + weight = 1.0 + msg = ( + f"{source_name}: OK (next release expected " + f"{expected_next.date()}, current obs age {age_days:.0f}d)" + ) + elif overdue <= policy.stale_threshold: + status = "late" + weight = 1.0 + msg = ( + f"{source_name}: late (next release expected " + f"{expected_next.date()}, overdue {overdue.days}d)" + ) + elif overdue <= policy.dead_threshold: + status = "stale" + decay_age = overdue - decay_start + hl = policy.weight_decay_half_life + weight = max(0.0, min(1.0, 2.0 ** (-decay_age / hl))) + msg = ( + f"{source_name}: stale (next release expected " + f"{expected_next.date()}, weight {weight:.2f})" + ) + else: + status = "dead" + weight = 0.0 + msg = ( + f"{source_name}: dead (next release expected " + f"{expected_next.date()}, overdue {overdue.days}d)" + ) else: - status = "dead" - weight = 0.0 - msg = ( - f"{source_name}: dead (age {age_days:.0f}d, " - f"exceeds dead threshold {policy.dead_threshold.days}d)" + ok_limit = policy.expected_cadence + policy.grace_period + decay_start = ( + policy.weight_decay_start + if policy.weight_decay_start is not None + else policy.stale_threshold ) + if age <= ok_limit: + status = "ok" + weight = 1.0 + msg = f"{source_name}: OK (age {age_days:.0f}d)" + elif age <= policy.stale_threshold: + status = "late" + weight = 1.0 + msg = ( + f"{source_name}: late (age {age_days:.0f}d, " + f"expected every {policy.expected_cadence.days}d)" + ) + elif age <= policy.dead_threshold: + status = "stale" + decay_age = age - decay_start + hl = policy.weight_decay_half_life + weight = max(0.0, min(1.0, 2.0 ** (-decay_age / hl))) + msg = f"{source_name}: stale (age {age_days:.0f}d, weight {weight:.2f})" + else: + status = "dead" + weight = 0.0 + msg = ( + f"{source_name}: dead (age {age_days:.0f}d, " + f"exceeds dead threshold {policy.dead_threshold.days}d)" + ) + return SourceHealthStatus( source_name=source_name, latest_obs_date=latest_obs_date, asof=asof, age=age, age_days=age_days, + overdue=overdue, + overdue_days=overdue_days, status=status, weight_factor=weight, expected_next=expected_next, @@ -114,6 +162,81 @@ def assess_source_health( ) +def _resolve_expected_next_release( + latest_obs_date: pd.Timestamp, + asof: pd.Timestamp, + policy: SourceHealthPolicy, +) -> pd.Timestamp: + if policy.release_rule is None: + return latest_obs_date + policy.expected_cadence + + next_obs_date = _advance_observation_date(latest_obs_date, policy.expected_cadence) + expected_release = pd.Timestamp( + policy.release_rule.expected_release_date(next_obs_date.date()) + ) + if asof.tzinfo is not None and expected_release.tzinfo is None: + return expected_release.tz_localize("UTC") + if asof.tzinfo is None and expected_release.tzinfo is not None: + return expected_release.tz_localize(None) + return expected_release + + +def _advance_observation_date( + latest_obs_date: pd.Timestamp, + expected_cadence: pd.Timedelta, +) -> pd.Timestamp: + days = expected_cadence / pd.Timedelta(days=1) + base = latest_obs_date.normalize() + + if 364 <= days <= 366 and base.is_year_end: + return base + YearEnd(1) + if 89 <= days <= 92 and base.is_quarter_end: + return base + QuarterEnd(1) + if 28 <= days <= 31 and base.is_month_end: + return base + MonthEnd(1) + return base + expected_cadence + + +def build_health_report( + statuses: Mapping[str, SourceHealthStatus] | Sequence[SourceHealthStatus], +) -> pd.DataFrame: + """Convert source health statuses into a deterministic report frame.""" + values = list(statuses.values()) if isinstance(statuses, Mapping) else list(statuses) + rows = [ + { + "source_name": status.source_name, + "latest_obs_date": status.latest_obs_date, + "asof": status.asof, + "expected_next": status.expected_next, + "age": status.age, + "age_days": status.age_days, + "overdue": status.overdue, + "overdue_days": status.overdue_days, + "status": status.status, + "weight_factor": status.weight_factor, + "message": status.message, + } + for status in values + ] + if not rows: + return pd.DataFrame( + columns=[ + "source_name", + "latest_obs_date", + "asof", + "expected_next", + "age", + "age_days", + "overdue", + "overdue_days", + "status", + "weight_factor", + "message", + ] + ) + return pd.DataFrame(rows).sort_values("source_name").reset_index(drop=True) + + # --------------------------------------------------------------------------- # Default policies for known sources # --------------------------------------------------------------------------- @@ -142,3 +265,4 @@ def assess_source_health( HealthPolicy = SourceHealthPolicy HealthStatus = SourceHealthStatus assess_health = assess_source_health +health_report = build_health_report diff --git a/alphaforge/pipeline/tracker.py b/alphaforge/pipeline/tracker.py index 7efd0d1..99b4fa8 100644 --- a/alphaforge/pipeline/tracker.py +++ b/alphaforge/pipeline/tracker.py @@ -7,7 +7,12 @@ from alphaforge.pit.accessor import PITAccessor -from .health import SourceHealthPolicy, SourceHealthStatus, assess_source_health +from .health import ( + SourceHealthPolicy, + SourceHealthStatus, + assess_source_health, + build_health_report, +) _STATUS_CODE = {"ok": 0, "late": 1, "stale": 2, "dead": 3, "empty": 4} @@ -51,6 +56,8 @@ def assess(self, source_name: str, asof: pd.Timestamp) -> SourceHealthStatus: asof=asof, age=None, age_days=None, + overdue=None, + overdue_days=None, status="empty", weight_factor=0.0, expected_next=None, @@ -93,10 +100,24 @@ def record(self, status: SourceHealthStatus) -> None: "value": status.age_days, "source": "health", }) + if status.overdue_days is not None: + rows.append( + { + "series_key": f"{base}.overdue_days", + "obs_date": obs_date, + "asof_utc": asof, + "value": status.overdue_days, + "source": "health", + } + ) df = pd.DataFrame(rows) self.pit.upsert_pit_observations(df, strict="coerce") + def report(self, asof: pd.Timestamp) -> pd.DataFrame: + """Return a dataframe report for all configured source policies.""" + return build_health_report(self.assess_all(asof)) + def history( self, source_name: str, diff --git a/alphaforge/pit/__init__.py b/alphaforge/pit/__init__.py index 6774a44..96e2c2f 100644 --- a/alphaforge/pit/__init__.py +++ b/alphaforge/pit/__init__.py @@ -54,6 +54,12 @@ PITPipelineStep, coerce_pipeline_spec, ) +from .queries import ( + RefRevisionQuery, + RefSnapshotQuery, + coerce_ref_revision_query, + coerce_ref_snapshot_query, +) from .ref_entity import make_ref_entity_id, parse_ref_entity_id from .resolvers import ( FrozenResolver, @@ -121,6 +127,10 @@ "PITPipelineSpec", "PITPipelineResult", "coerce_pipeline_spec", + "RefSnapshotQuery", + "RefRevisionQuery", + "coerce_ref_snapshot_query", + "coerce_ref_revision_query", "PITExpressionNode", "PITExpressionGraphSpec", "PITExpressionGraphResult", diff --git a/alphaforge/pit/accessor.py b/alphaforge/pit/accessor.py index 634ed87..c907b14 100644 --- a/alphaforge/pit/accessor.py +++ b/alphaforge/pit/accessor.py @@ -6,13 +6,21 @@ import uuid import warnings from dataclasses import dataclass +from pathlib import Path from typing import Any, Literal, Mapping, Sequence import duckdb import pandas as pd from pandas.tseries.offsets import MonthEnd -from alphaforge.time.ref_period import RefFreq, RefPeriod +from alphaforge.time.ref_period import ( + ObsDateAnchor, + RefFreq, + RefPeriod, + coerce_ref_period, + normalize_obs_date_anchor, + normalize_ref_freq, +) from .exceptions import ( PITCausalityError, @@ -39,6 +47,13 @@ PITPipelineSpec, coerce_pipeline_spec, ) +from .queries import ( + RefRevisionQuery, + RefSnapshotQuery, + coerce_ref_revision_query, + coerce_ref_snapshot_query, +) +from .ref_entity import make_ref_entity_id from .transforms import ( EngineMismatchPolicy, PITEngineResolution, @@ -519,6 +534,14 @@ def _evaluate_expression_series( class PITAccessor: conn: duckdb.DuckDBPyConnection + @classmethod + def open(cls, root: str | Path) -> "PITAccessor": + """Open a PIT accessor from a DuckDBParquetStore root.""" + from alphaforge.store.duckdb_parquet import DuckDBParquetStore + + store = DuckDBParquetStore(root=str(root)) + return cls(store.conn()) + def __post_init__(self) -> None: ensure_pit_table(self.conn) @@ -654,6 +677,7 @@ def get_snapshot_multi( { "series_key": pd.Series(dtype="object"), "obs_date": pd.Series(dtype="datetime64[ns, UTC]"), + "source_asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), "value": pd.Series(dtype="float64"), } ) @@ -673,11 +697,12 @@ def get_snapshot_multi( where_clause = " AND ".join(filters) query = f""" - SELECT series_key, obs_date, value + SELECT series_key, obs_date, source_asof_utc, value FROM ( SELECT series_key, obs_date, + asof_utc AS source_asof_utc, value, ROW_NUMBER() OVER ( PARTITION BY series_key, obs_date @@ -695,12 +720,14 @@ def get_snapshot_multi( { "series_key": pd.Series(dtype="object"), "obs_date": pd.Series(dtype="datetime64[ns, UTC]"), + "source_asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), "value": pd.Series(dtype="float64"), } ) out = df.copy() out["obs_date"] = to_utc_aware(out["obs_date"]) + out["source_asof_utc"] = to_utc_aware(out["source_asof_utc"]) out["value"] = pd.to_numeric(out["value"], errors="coerce") return out.reset_index(drop=True) @@ -928,16 +955,15 @@ def get_revision_path_multi(self, requests: pd.DataFrame) -> pd.DataFrame: def get_revision_timeline_ref( self, series_key: str, - ref: str | RefPeriod, + ref: object, start_asof: pd.Timestamp | None = None, end_asof: pd.Timestamp | None = None, *, freq: RefFreq | None = None, + obs_date_anchor: ObsDateAnchor | str = "end", ) -> pd.Series: - ref_period = RefPeriod.parse(ref) if isinstance(ref, str) else ref - if freq is not None and freq != ref_period.freq: - raise PITContractError("Reference period frequency does not match requested freq.") - obs_date = ref_period.end_obs_date() + ref_period = self._resolve_ref_period(ref, freq=freq, obs_date_anchor=obs_date_anchor) + obs_date = ref_period.obs_date(anchor=obs_date_anchor) return self.get_revision_timeline( series_key, obs_date, @@ -949,37 +975,110 @@ def get_snapshot_ref( self, series_key: str, asof: pd.Timestamp, - start_ref: str | RefPeriod | None = None, - end_ref: str | RefPeriod | None = None, + start_ref: object | None = None, + end_ref: object | None = None, *, freq: RefFreq | None = None, + obs_date_anchor: ObsDateAnchor | str = "end", ) -> pd.Series: - def _resolve(ref_value: str | RefPeriod | None) -> pd.Timestamp | None: + def _resolve(ref_value: object | None) -> pd.Timestamp | None: if ref_value is None: return None - if isinstance(ref_value, RefPeriod): - ref_period = ref_value - else: - ref_period = RefPeriod.parse(ref_value) - if freq is not None and ref_period.freq != freq: - raise PITContractError("Reference period frequency does not match requested freq.") - return ref_period.end_obs_date() + return self._resolve_ref_period( + ref_value, + freq=freq, + obs_date_anchor=obs_date_anchor, + ).obs_date(anchor=obs_date_anchor) start_ts = _resolve(start_ref) end_ts = _resolve(end_ref) return self.get_snapshot(series_key, asof, start=start_ts, end=end_ts) @staticmethod - def _resolve_ref_period(ref: str | RefPeriod, freq: RefFreq | None = None) -> RefPeriod: - ref_period = RefPeriod.parse(ref) if isinstance(ref, str) else ref - if freq is not None and freq != ref_period.freq: - raise PITContractError("Reference period frequency does not match requested freq.") - return ref_period + def _resolve_ref_period( + ref: object, + freq: RefFreq | None = None, + obs_date_anchor: ObsDateAnchor | str = "end", + ) -> RefPeriod: + try: + return coerce_ref_period(ref, freq=freq, obs_date_anchor=obs_date_anchor) + except ValueError as exc: + raise PITContractError(str(exc)) from exc + + def snapshot_ref( + self, + query: RefSnapshotQuery | Mapping[str, Any], + ) -> pd.Series: + query_obj = coerce_ref_snapshot_query(query) + obs_date_anchor = normalize_obs_date_anchor(query_obj.obs_date_anchor) + freq = normalize_ref_freq(query_obj.freq) + snap = self.get_snapshot_ref( + query_obj.series_key, + query_obj.asof, + start_ref=query_obj.start_ref, + end_ref=query_obj.end_ref, + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + if snap.empty: + return pd.Series( + index=pd.Index([], dtype="object", name="ref_period"), + dtype="float64", + name=query_obj.series_key, + ) + + ref_index = [ + self._resolve_ref_period( + obs_date, + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + for obs_date in snap.index + ] + if len(ref_index) != len(set(ref_index)): + raise PITContractError( + "Ref snapshot query produced duplicate reference periods. " + "Check the requested frequency and obs_date_anchor." + ) + + out = pd.Series( + snap.to_numpy(), + index=pd.Index(ref_index, name="ref_period"), + name=query_obj.series_key, + ) + out.attrs["freq"] = freq + out.attrs["obs_date_anchor"] = obs_date_anchor + return out + + def revisions_ref( + self, + query: RefRevisionQuery | Mapping[str, Any], + ) -> pd.Series: + query_obj = coerce_ref_revision_query(query) + obs_date_anchor = normalize_obs_date_anchor(query_obj.obs_date_anchor) + freq = normalize_ref_freq(query_obj.freq) + ref_period = self._resolve_ref_period( + query_obj.ref, + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + series = self.get_revision_timeline_ref( + query_obj.series_key, + ref_period, + start_asof=query_obj.start_asof, + end_asof=query_obj.end_asof, + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + series.name = make_ref_entity_id(query_obj.series_key, ref_period) + series.attrs["ref_period"] = ref_period + series.attrs["obs_date_anchor"] = obs_date_anchor + return series def list_release_stream( self, series_key: str, - ref: str | RefPeriod, + ref: object, asof: pd.Timestamp | None = None, *, freq: RefFreq | None = None, @@ -1039,7 +1138,7 @@ def list_release_stream( def resolve_release( self, series_key: str, - ref: str | RefPeriod, + ref: object, *, policy: ReleaseSelectionPolicy | Mapping[str, Any] | str = "latest", asof: pd.Timestamp | None = None, @@ -1223,63 +1322,166 @@ def _align_snapshot_index( raise PITContractError("align must be one of {'month_end', 'quarter_end'}.") return pd.DatetimeIndex(aligned).tz_localize("UTC") - def build_snapshot_panel( + @staticmethod + def _empty_snapshot_panel_long() -> pd.DataFrame: + return pd.DataFrame( + { + "series_key": pd.Series(dtype="object"), + "series_alias": pd.Series(dtype="object"), + "obs_date": pd.Series(dtype="datetime64[ns, UTC]"), + "source_obs_date": pd.Series(dtype="datetime64[ns, UTC]"), + "source_asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), + "value": pd.Series(dtype="float64"), + } + ) + + def _snapshot_bounds_from_spec( + self, + spec: SnapshotSeriesSpec, + ) -> tuple[pd.Timestamp | None, pd.Timestamp | None]: + anchor = normalize_obs_date_anchor(spec.obs_date_anchor) + freq = normalize_ref_freq(spec.freq) + + def _resolve(ref_value: object | None) -> pd.Timestamp | None: + if ref_value is None: + return None + return self._resolve_ref_period( + ref_value, + freq=freq, + obs_date_anchor=anchor, + ).obs_date(anchor=anchor) + + return _resolve(spec.start_ref), _resolve(spec.end_ref) + + def _snapshot_panel_rows_from_frame( + self, + spec: SnapshotSeriesSpec, + frame: pd.DataFrame, + *, + align: Literal["month_end", "quarter_end"], + ) -> pd.DataFrame: + if frame.empty: + return self._empty_snapshot_panel_long() + + rows = frame.copy() + rows["source_obs_date"] = to_utc_aware(rows["source_obs_date"]) + rows["source_asof_utc"] = to_utc_aware(rows["source_asof_utc"]) + rows["value"] = pd.to_numeric(rows["value"], errors="coerce") + rows["obs_date"] = self._align_snapshot_index( + pd.DatetimeIndex(rows["source_obs_date"]), + align=align, + ) + rows["series_key"] = spec.series_key + rows["series_alias"] = spec.alias or spec.series_key + rows = rows.loc[ + :, ["series_key", "series_alias", "obs_date", "source_obs_date", "source_asof_utc", "value"] + ] + rows = rows.sort_values(["series_key", "obs_date", "source_obs_date", "source_asof_utc"]) + rows = rows.drop_duplicates(subset=["series_key", "obs_date"], keep="last") + return rows.reset_index(drop=True) + + def build_snapshot_panel_long( self, series_specs: Sequence[SnapshotSeriesSpec | Mapping[str, Any]], asof: pd.Timestamp, *, align: Literal["month_end", "quarter_end"] = "month_end", - join: Literal["inner", "left", "right", "outer"] = "outer", ) -> pd.DataFrame: - if join not in {"inner", "left", "right", "outer"}: - raise PITContractError("join must be one of {'inner', 'left', 'right', 'outer'}.") - - panel: pd.DataFrame | None = None - for raw_spec in series_specs: - spec = coerce_snapshot_series_spec(raw_spec) - start_obs = ( - RefPeriod.parse(spec.start_ref).end_obs_date() - if spec.start_ref is not None - else None - ) - end_obs = ( - RefPeriod.parse(spec.end_ref).end_obs_date() if spec.end_ref is not None else None - ) + normalized_specs = [coerce_snapshot_series_spec(raw_spec) for raw_spec in series_specs] + if not normalized_specs: + return self._empty_snapshot_panel_long() + + aliases = [spec.alias or spec.series_key for spec in normalized_specs] + if len(set(aliases)) != len(aliases): + raise PITContractError("Snapshot panel aliases must be unique.") + + latest_batches: dict[ + tuple[pd.Timestamp | None, pd.Timestamp | None], + list[SnapshotSeriesSpec], + ] = {} + pieces: list[pd.DataFrame] = [] + + for spec in normalized_specs: + start_obs, end_obs = self._snapshot_bounds_from_spec(spec) + mode, _ = normalize_release_selection_policy(spec.release_policy) + if mode == "latest": + latest_batches.setdefault((start_obs, end_obs), []).append(spec) + continue - snapshot = self._snapshot_with_release_policy( + rows = self._snapshot_rows_with_release_policy( spec.series_key, asof=asof, policy=spec.release_policy, start=start_obs, end=end_obs, - ) - if snapshot.empty: - s = pd.Series(dtype="float64", name=spec.alias or spec.series_key) - else: - aligned_index = self._align_snapshot_index( - pd.DatetimeIndex(snapshot.index), + ).rename(columns={"obs_date": "source_obs_date"}) + pieces.append( + self._snapshot_panel_rows_from_frame( + spec, + rows, align=align, ) - tmp = pd.DataFrame( - { - "source_obs_date": pd.DatetimeIndex(snapshot.index), - "obs_date": aligned_index, - "value": snapshot.to_numpy(), - } + ) + + for (start_obs, end_obs), specs in latest_batches.items(): + batch = self.get_snapshot_multi( + [spec.series_key for spec in specs], + asof=asof, + start=start_obs, + end=end_obs, + ).rename(columns={"obs_date": "source_obs_date"}) + for spec in specs: + spec_rows = batch.loc[batch["series_key"] == spec.series_key].copy() + pieces.append( + self._snapshot_panel_rows_from_frame( + spec, + spec_rows, + align=align, + ) ) - tmp = tmp.sort_values(["obs_date", "source_obs_date"]) - tmp = tmp.drop_duplicates(subset=["obs_date"], keep="last") - s = pd.Series( - tmp["value"].to_numpy(), - index=pd.DatetimeIndex(tmp["obs_date"]), - name=spec.alias or spec.series_key, - ).sort_index() - - frame = s.to_frame() - panel = frame if panel is None else panel.join(frame, how=join) - - if panel is None: + + if not pieces: + return self._empty_snapshot_panel_long() + + out = pd.concat(pieces, ignore_index=True) + if out.empty: + return self._empty_snapshot_panel_long() + return out.sort_values(["obs_date", "series_alias"]).reset_index(drop=True) + + def build_snapshot_panel( + self, + series_specs: Sequence[SnapshotSeriesSpec | Mapping[str, Any]], + asof: pd.Timestamp, + *, + align: Literal["month_end", "quarter_end"] = "month_end", + join: Literal["inner", "left", "right", "outer"] = "outer", + ) -> pd.DataFrame: + if join not in {"inner", "left", "right", "outer"}: + raise PITContractError("join must be one of {'inner', 'left', 'right', 'outer'}.") + normalized_specs = [coerce_snapshot_series_spec(raw_spec) for raw_spec in series_specs] + if not normalized_specs: return pd.DataFrame() + aliases = [spec.alias or spec.series_key for spec in normalized_specs] + + long_panel = self.build_snapshot_panel_long( + normalized_specs, + asof=asof, + align=align, + ) + if long_panel.empty: + return pd.DataFrame(index=pd.DatetimeIndex([], name="obs_date"), columns=aliases) + + panel = ( + long_panel.pivot(index="obs_date", columns="series_alias", values="value") + .sort_index() + .reindex(columns=aliases) + ) + if join == "inner": + panel = panel.dropna(how="any") + elif join == "left": + panel = panel.loc[panel[aliases[0]].notna()] + elif join == "right": + panel = panel.loc[panel[aliases[-1]].notna()] panel.index = ( panel.index.tz_convert("UTC") @@ -1287,7 +1489,7 @@ def build_snapshot_panel( else panel.index.tz_localize("UTC") ) panel.index.name = "obs_date" - return panel.sort_index() + return panel def _list_candidate_asofs( self, @@ -3333,6 +3535,255 @@ def list_transforms(self, output_series_key: str | None = None) -> pd.DataFrame: df["created_utc"] = to_utc_aware(df["created_utc"]) return df + @staticmethod + def _empty_series_lineage() -> pd.DataFrame: + return pd.DataFrame( + { + "series_key": pd.Series(dtype="object"), + "obs_date": pd.Series(dtype="datetime64[ns, UTC]"), + "asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), + "value": pd.Series(dtype="float64"), + "source": pd.Series(dtype="object"), + "meta_json": pd.Series(dtype="object"), + "lineage_kind": pd.Series(dtype="object"), + "transform_id": pd.Series(dtype="object"), + "graph_id": pd.Series(dtype="object"), + "node_name": pd.Series(dtype="object"), + "op": pd.Series(dtype="object"), + "axis": pd.Series(dtype="object"), + "engine": pd.Series(dtype="object"), + "engine_requested": pd.Series(dtype="object"), + "experimental": pd.Series(dtype="bool"), + "input_series_keys": pd.Series(dtype="object"), + "source_asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), + "selected_input_series_key": pd.Series(dtype="object"), + "selected_input_asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), + "source_asof_by_series_utc": pd.Series(dtype="object"), + "max_source_asof_utc": pd.Series(dtype="datetime64[ns, UTC]"), + "causality_status": pd.Series(dtype="object"), + } + ) + + @staticmethod + def _parse_lineage_payload(meta_json: object) -> dict[str, Any]: + if meta_json is None or pd.isna(meta_json): + return {} + try: + payload = json.loads(str(meta_json)) + except (TypeError, ValueError, json.JSONDecodeError): + return {} + return payload if isinstance(payload, dict) else {} + + @staticmethod + def _parse_lineage_timestamp(value: object) -> pd.Timestamp | None: + if value is None: + return None + try: + return to_utc_aware(value) + except (TypeError, ValueError): + return None + + def _parse_lineage_timestamp_map(self, value: object) -> dict[str, pd.Timestamp]: + if not isinstance(value, Mapping): + return {} + out: dict[str, pd.Timestamp] = {} + for key, raw in value.items(): + parsed = self._parse_lineage_timestamp(raw) + if parsed is not None: + out[str(key)] = parsed + return out + + @staticmethod + def _lineage_kind(payload: Mapping[str, Any]) -> str: + if "transform_id" in payload: + return "transform" + if "graph_id" in payload: + return "expression_graph" + if payload: + return "derived" + return "raw" + + def get_series_lineage( + self, + series_key: str, + start_obs: pd.Timestamp | None = None, + end_obs: pd.Timestamp | None = None, + start_asof: pd.Timestamp | None = None, + end_asof: pd.Timestamp | None = None, + *, + limit: int | None = 500, + ) -> pd.DataFrame: + filters = ["series_key = ?"] + params: list[object] = [series_key] + if start_obs is not None: + filters.append("obs_date >= ?") + params.append(to_utc_naive(start_obs)) + if end_obs is not None: + filters.append("obs_date <= ?") + params.append(to_utc_naive(end_obs)) + if start_asof is not None: + filters.append("asof_utc >= ?") + params.append(to_utc_naive(start_asof)) + if end_asof is not None: + filters.append("asof_utc <= ?") + params.append(to_utc_naive(end_asof)) + + where_clause = " AND ".join(filters) + query = f""" + SELECT series_key, obs_date, asof_utc, value, source, meta_json + FROM {_PIT_TABLE} + WHERE {where_clause} + ORDER BY obs_date ASC, asof_utc ASC + """ + if limit is not None: + if limit <= 0: + raise PITContractError("get_series_lineage limit must be > 0 when provided.") + query += " LIMIT ?" + params.append(int(limit)) + + df = self.conn.execute(query, params).fetchdf() + if df.empty: + return self._empty_series_lineage() + + df["obs_date"] = to_utc_aware(df["obs_date"]) + df["asof_utc"] = to_utc_aware(df["asof_utc"]) + df["value"] = pd.to_numeric(df["value"], errors="coerce") + + records: list[dict[str, Any]] = [] + for row in df.itertuples(index=False): + payload = self._parse_lineage_payload(row.meta_json) + lineage_kind = self._lineage_kind(payload) + input_series_keys = payload.get("input_series_keys") + if not isinstance(input_series_keys, list): + input_key = payload.get("input_series_key") + input_series_keys = [input_key] if input_key is not None else [] + input_series_keys = tuple(str(key) for key in input_series_keys if str(key).strip()) + + source_asof = self._parse_lineage_timestamp(payload.get("source_asof_utc")) + selected_input_asof = self._parse_lineage_timestamp( + payload.get("selected_input_asof_utc") + ) + source_asof_by_series = self._parse_lineage_timestamp_map( + payload.get("source_asof_by_series_utc") + ) + max_candidates = [ + ts + for ts in [source_asof, selected_input_asof, *source_asof_by_series.values()] + if ts is not None + ] + max_source_asof = max(max_candidates) if max_candidates else None + + if lineage_kind == "raw": + causality_status = "raw" + elif max_source_asof is None: + causality_status = "unknown" + elif max_source_asof > pd.Timestamp(row.asof_utc): + causality_status = "violation" + elif bool(payload.get("experimental")): + causality_status = "experimental" + else: + causality_status = "ok" + + records.append( + { + "series_key": row.series_key, + "obs_date": pd.Timestamp(row.obs_date), + "asof_utc": pd.Timestamp(row.asof_utc), + "value": row.value, + "source": row.source, + "meta_json": row.meta_json, + "lineage_kind": lineage_kind, + "transform_id": payload.get("transform_id"), + "graph_id": payload.get("graph_id"), + "node_name": payload.get("node_name"), + "op": payload.get("op"), + "axis": payload.get("axis"), + "engine": payload.get("engine"), + "engine_requested": payload.get("engine_requested"), + "experimental": bool(payload.get("experimental", False)), + "input_series_keys": input_series_keys, + "source_asof_utc": source_asof, + "selected_input_series_key": payload.get("selected_input_series_key"), + "selected_input_asof_utc": selected_input_asof, + "source_asof_by_series_utc": source_asof_by_series, + "max_source_asof_utc": max_source_asof, + "causality_status": causality_status, + } + ) + + return pd.DataFrame(records) + + def explain_series( + self, + series_key: str, + start_obs: pd.Timestamp | None = None, + end_obs: pd.Timestamp | None = None, + start_asof: pd.Timestamp | None = None, + end_asof: pd.Timestamp | None = None, + *, + limit: int | None = 500, + ) -> dict[str, Any]: + lineage = self.get_series_lineage( + series_key, + start_obs=start_obs, + end_obs=end_obs, + start_asof=start_asof, + end_asof=end_asof, + limit=limit, + ) + if lineage.empty: + return { + "series_key": series_key, + "row_count": 0, + "derived_row_count": 0, + "lineage_kinds": [], + "input_series_keys": [], + "transform_ids": [], + "graph_ids": [], + "causality_status_counts": {}, + "causality_safe": True, + } + + derived = lineage[lineage["lineage_kind"] != "raw"] + input_series_keys = sorted( + { + key + for keys in lineage["input_series_keys"] + for key in (keys if isinstance(keys, tuple) else tuple()) + } + ) + transform_ids = sorted( + {str(value) for value in lineage["transform_id"].dropna().tolist() if str(value).strip()} + ) + graph_ids = sorted( + {str(value) for value in lineage["graph_id"].dropna().tolist() if str(value).strip()} + ) + status_counts = { + str(key): int(value) + for key, value in lineage["causality_status"].value_counts(dropna=False).items() + } + causality_safe = ( + status_counts.get("violation", 0) == 0 + and status_counts.get("unknown", 0) == 0 + and status_counts.get("experimental", 0) == 0 + ) + + return { + "series_key": series_key, + "row_count": int(len(lineage)), + "derived_row_count": int(len(derived)), + "lineage_kinds": sorted({str(value) for value in lineage["lineage_kind"].tolist()}), + "input_series_keys": input_series_keys, + "transform_ids": transform_ids, + "graph_ids": graph_ids, + "causality_status_counts": status_counts, + "causality_safe": causality_safe, + "latest_output_asof_utc": lineage["asof_utc"].max(), + "latest_source_asof_utc": lineage["max_source_asof_utc"].dropna().max() + if lineage["max_source_asof_utc"].notna().any() + else None, + } + def list_pipelines(self, pipeline_id: str | None = None) -> pd.DataFrame: query = f""" SELECT diff --git a/alphaforge/pit/adapters/alphaforge_adapter.py b/alphaforge/pit/adapters/alphaforge_adapter.py index 95b5f82..bf9431b 100644 --- a/alphaforge/pit/adapters/alphaforge_adapter.py +++ b/alphaforge/pit/adapters/alphaforge_adapter.py @@ -10,6 +10,7 @@ from alphaforge.data.context import DataContext from alphaforge.data.query import Query +from alphaforge.pit.accessor import PITAccessor from alphaforge.pit.adapters.alphaforge_layer import AlphaForgePITLayer from alphaforge.pit.adapters.base import PITAdapter from alphaforge.pit.observation import PITObservation, SeriesMetadata @@ -41,7 +42,10 @@ class AlphaForgePITAdapter(PITAdapter): """Point-in-time data adapter for AlphaForge.""" def __init__(self, ctx: DataContext) -> None: + if ctx.pit is None: + raise ValueError("AlphaForge PIT adapter requires PIT-enabled DataContext") self._ctx = ctx + self._pit: PITAccessor = ctx.pit self._layer = AlphaForgePITLayer(ctx) @property @@ -52,7 +56,7 @@ def supports_pit(self, series_id: str) -> bool: return True def list_vintages(self, query_series_key: str) -> list[date]: - conn = self._ctx.pit.conn + conn = self._pit.conn rows = conn.execute( "SELECT DISTINCT asof_utc FROM pit_observations WHERE series_key = ?", [query_series_key], @@ -131,7 +135,7 @@ def fetch_asof( "release_time_utc": pd.NaT, } ) - self._ctx.pit.upsert_pit_observations(pit_df) + self._pit.upsert_pit_observations(pit_df) snap = self._layer.snapshot(query_series_key, asof=asof_ts, start=start_ts, end=end_ts) diff --git a/alphaforge/pit/adapters/alphaforge_layer.py b/alphaforge/pit/adapters/alphaforge_layer.py index c144a85..24ba716 100644 --- a/alphaforge/pit/adapters/alphaforge_layer.py +++ b/alphaforge/pit/adapters/alphaforge_layer.py @@ -9,9 +9,10 @@ import pandas as pd from alphaforge.data.context import DataContext -from alphaforge.pit.transforms import PITTransformResult, PITTransformSpec +from alphaforge.pit.accessor import PITAccessor +from alphaforge.pit.transforms import EngineMismatchPolicy, PITTransformResult, PITTransformSpec from alphaforge.pit.utils.timestamps import coerce_utc_timestamp, normalize_utc_day -from alphaforge.time.ref_period import RefFreq, RefPeriod +from alphaforge.time.ref_period import RefFreq, coerce_ref_period class AlphaForgePITLayer: @@ -21,6 +22,7 @@ def __init__(self, ctx: DataContext) -> None: if ctx.pit is None: raise ValueError("PIT requires DuckDBParquetStore-backed DataContext") self._ctx = ctx + self._pit: PITAccessor = ctx.pit def snapshot( self, @@ -29,7 +31,7 @@ def snapshot( start: pd.Timestamp | None = None, end: pd.Timestamp | None = None, ) -> pd.Series: - return self._ctx.pit.get_snapshot(series_key, asof=asof, start=start, end=end) + return self._pit.get_snapshot(series_key, asof=asof, start=start, end=end) def snapshot_multi( self, @@ -39,9 +41,7 @@ def snapshot_multi( start: pd.Timestamp | None = None, end: pd.Timestamp | None = None, ) -> pd.DataFrame: - if self._ctx.pit is None: - raise ValueError("PIT store is not available; cannot query snapshots.") - return self._ctx.pit.get_snapshot_multi( + return self._pit.get_snapshot_multi( list(series_keys), asof=asof, start=start, @@ -52,16 +52,16 @@ def snapshot_ref( self, series_key: str, asof: pd.Timestamp, - start_ref: str | RefPeriod | None = None, - end_ref: str | RefPeriod | None = None, + start_ref: object | None = None, + end_ref: object | None = None, *, freq: RefFreq | None = None, ) -> pd.Series: - if isinstance(start_ref, str): - start_ref = RefPeriod.parse(start_ref) - if isinstance(end_ref, str): - end_ref = RefPeriod.parse(end_ref) - return self._ctx.pit.get_snapshot_ref( + if start_ref is not None: + start_ref = coerce_ref_period(start_ref, freq=freq) + if end_ref is not None: + end_ref = coerce_ref_period(end_ref, freq=freq) + return self._pit.get_snapshot_ref( series_key, asof=asof, start_ref=start_ref, @@ -76,7 +76,7 @@ def revisions( start_asof: pd.Timestamp | None = None, end_asof: pd.Timestamp | None = None, ) -> pd.Series: - return self._ctx.pit.get_revision_timeline( + return self._pit.get_revision_timeline( series_key, obs_date=obs_date, start_asof=start_asof, end_asof=end_asof ) @@ -87,30 +87,29 @@ def revision_path( start_asof: pd.Timestamp | None = None, end_asof: pd.Timestamp | None = None, ) -> pd.DataFrame: - return self._ctx.pit.get_revision_path( + return self._pit.get_revision_path( series_key, obs_date=obs_date, start_asof=start_asof, end_asof=end_asof ) def revision_path_multi(self, requests: pd.DataFrame) -> pd.DataFrame: - return self._ctx.pit.get_revision_path_multi(requests) + return self._pit.get_revision_path_multi(requests) def revisions_ref( self, series_key: str, - ref: str | RefPeriod, + ref: object, start_asof: pd.Timestamp | None = None, end_asof: pd.Timestamp | None = None, *, freq: RefFreq | None = None, ) -> pd.Series: - if isinstance(ref, str): - ref = RefPeriod.parse(ref) - return self._ctx.pit.get_revision_timeline_ref( + ref = coerce_ref_period(ref, freq=freq) + return self._pit.get_revision_timeline_ref( series_key, ref=ref, start_asof=start_asof, end_asof=end_asof, freq=freq ) def upsert(self, df: pd.DataFrame) -> None: - self._ctx.pit.upsert_pit_observations(df) + self._pit.upsert_pit_observations(df) def apply_transform( self, @@ -119,9 +118,9 @@ def apply_transform( overwrite: bool = False, persist: bool = True, allow_experimental: bool = False, - on_engine_mismatch: str = "error", + on_engine_mismatch: EngineMismatchPolicy = "error", ) -> PITTransformResult: - return self._ctx.pit.apply_transform( + return self._pit.apply_transform( spec, overwrite=overwrite, persist=persist, @@ -134,9 +133,9 @@ def explain_transform( spec: PITTransformSpec | dict[str, Any], *, allow_experimental: bool = False, - on_engine_mismatch: str = "error", + on_engine_mismatch: EngineMismatchPolicy = "error", ) -> dict[str, Any]: - return self._ctx.pit.explain_transform( + return self._pit.explain_transform( spec, allow_experimental=allow_experimental, on_engine_mismatch=on_engine_mismatch, @@ -149,9 +148,7 @@ def list_pit_observations_asof( obs_date: date, asof_date: date, ) -> pd.DataFrame: - if self._ctx.pit is None: - raise ValueError("PIT store is not available; cannot list PIT observations.") - conn = self._ctx.pit.conn + conn = self._pit.conn obs_day = normalize_utc_day(obs_date) asof_cutoff = coerce_utc_timestamp(asof_date).normalize() + pd.Timedelta(days=1) df = conn.execute( diff --git a/alphaforge/pit/adapters/fred.py b/alphaforge/pit/adapters/fred.py index 57350e8..67b32e5 100644 --- a/alphaforge/pit/adapters/fred.py +++ b/alphaforge/pit/adapters/fred.py @@ -32,12 +32,13 @@ def __init__(self, api_key: str | None = None) -> None: Args: api_key: FRED API key. If None, reads from FRED_API_KEY env var. """ - self._api_key = api_key or os.getenv("FRED_API_KEY") - if not self._api_key: + resolved_api_key = api_key or os.getenv("FRED_API_KEY") + if not resolved_api_key: raise ValueError( "FRED API key required. Set FRED_API_KEY environment variable " "or pass api_key parameter." ) + self._api_key = resolved_api_key self._session = requests.Session() self._vintage_cache: dict[str, tuple[list[date], float]] = {} self._cache_ttl = 86400 # 24 hours diff --git a/alphaforge/pit/catalog.py b/alphaforge/pit/catalog.py index f6ef5e9..8dccc9c 100644 --- a/alphaforge/pit/catalog.py +++ b/alphaforge/pit/catalog.py @@ -13,7 +13,7 @@ import yaml from alphaforge.pit.observation import SeriesMetadata -from alphaforge.pit.release_rules import ReleaseRule +from alphaforge.time.release_rules import ReleaseRule class SeriesCatalog: diff --git a/alphaforge/pit/missingness.py b/alphaforge/pit/missingness.py index 235912a..b99ba2d 100644 --- a/alphaforge/pit/missingness.py +++ b/alphaforge/pit/missingness.py @@ -1,152 +1,6 @@ -"""Missingness taxonomy and classification for nowcasting panels. +"""Compatibility shim for PIT missingness imports. -Provides a shared classifier so Phase 3 (views) and Phase 4 (imputation) -use identical logic to determine *why* a cell is NaN. +The canonical public path is now ``alphaforge.time.missingness``. """ -from __future__ import annotations - -from datetime import date -from enum import Enum -from typing import TYPE_CHECKING - -if TYPE_CHECKING: - from alphaforge.pit.release_rules import ReleaseRule - - -class MissingnessReason(str, Enum): - """Why a panel cell is NaN.""" - - STRUCTURAL = "structural" - """Frequency mismatch — e.g. quarterly obs_date in a monthly panel.""" - - FUTURE = "future" - """The observation period has not yet ended (obs_date > asof_date).""" - - RAGGED_EDGE = "ragged_edge" - """Not yet published — the expected release date is after asof_date.""" - - TRUE_MISSING = "true_missing" - """Should have been available by now but is absent — data issue.""" - - -def classify_missingness( - *, - obs_date: date, - asof_date: date, - series_frequency: str, - panel_frequency: str = "M", - release_rule: ReleaseRule | None = None, - publication_lag_months: int | None = None, - realized_release_date: date | None = None, -) -> MissingnessReason: - """Classify why a panel cell is NaN. - - Evaluation follows a strict order: - - 1. **STRUCTURAL** — frequency mismatch (e.g. quarterly series in a - monthly panel has NaN for non-quarter-end months). - 2. **FUTURE** — ``obs_date`` is after ``asof_date``. - 3. **RAGGED_EDGE vs TRUE_MISSING** — compare ``asof_date`` against the - expected (or realized) publication date: - - * If a ``realized_release_date`` is provided (from AlphaForge PIT - storage), it takes precedence over the expected schedule. - * Otherwise, ``release_rule.expected_release_date(obs_date)`` is used. - * As a last fallback, ``publication_lag_months`` gives a coarse - estimate. - * If none are available, default to ``TRUE_MISSING``. - - Parameters - ---------- - obs_date - The observation / reference-period end date of the NaN cell. - asof_date - The "as-of" date of the vintage snapshot. - series_frequency - Frequency code of the series (``"Q"``, ``"M"``, ``"W"``, etc.). - panel_frequency - Frequency code of the panel grid (default ``"M"`` for monthly). - release_rule - Optional structured release schedule from catalog metadata. - publication_lag_months - Optional simple lag heuristic (backward compat). - realized_release_date - Optional realized publication date from AlphaForge. When provided - this always takes precedence over the expected schedule. - """ - # 1. Structural: quarterly series in a monthly panel at non-quarter-end - if _is_structural(obs_date, series_frequency, panel_frequency): - return MissingnessReason.STRUCTURAL - - # 2. Future: observation period hasn't ended - if obs_date > asof_date: - return MissingnessReason.FUTURE - - # 3. Ragged edge vs true missing - expected = _resolve_expected_date( - obs_date, - release_rule=release_rule, - publication_lag_months=publication_lag_months, - realized_release_date=realized_release_date, - ) - - if expected is None: - # No release information at all — assume it should be here - return MissingnessReason.TRUE_MISSING - - if asof_date < expected: - return MissingnessReason.RAGGED_EDGE - - return MissingnessReason.TRUE_MISSING - - -# --------------------------------------------------------------------------- -# Internal helpers -# --------------------------------------------------------------------------- - -_QUARTER_END_MONTHS = {3, 6, 9, 12} - - -def _is_structural( - obs_date: date, series_frequency: str, panel_frequency: str -) -> bool: - """Return True when the NaN is a frequency-mismatch artifact.""" - sf = series_frequency.upper() - pf = panel_frequency.upper() - - if sf == "Q" and pf == "M": - # Quarterly series only has values at quarter-end months - return obs_date.month not in _QUARTER_END_MONTHS - - return False - - -def _resolve_expected_date( - obs_date: date, - *, - release_rule: ReleaseRule | None, - publication_lag_months: int | None, - realized_release_date: date | None, -) -> date | None: - """Determine the best-available expected publication date. - - Precedence: - 1. realized_release_date (from AlphaForge PIT storage) - 2. release_rule.expected_release_date(obs_date) - 3. obs_date + publication_lag_months (coarse fallback) - 4. None (no information) - """ - if realized_release_date is not None: - return realized_release_date - - if release_rule is not None: - return release_rule.expected_release_date(obs_date) - - if publication_lag_months is not None: - m = obs_date.month + publication_lag_months - y = obs_date.year + (m - 1) // 12 - m = (m - 1) % 12 + 1 - return date(y, m, 1) - - return None +from alphaforge.time.missingness import * # noqa: F403 diff --git a/alphaforge/pit/models.py b/alphaforge/pit/models.py index df9d74f..145a849 100644 --- a/alphaforge/pit/models.py +++ b/alphaforge/pit/models.py @@ -7,6 +7,13 @@ import pandas as pd +from alphaforge.time.ref_period import ( + ObsDateAnchor, + RefFreq, + normalize_obs_date_anchor, + normalize_ref_freq, +) + from .exceptions import PITContractError JoinMode = Literal["inner", "left", "right", "outer"] @@ -164,8 +171,10 @@ class PITExpressionGraphResult: class SnapshotSeriesSpec: series_key: str alias: str | None = None - start_ref: str | None = None - end_ref: str | None = None + start_ref: object | None = None + end_ref: object | None = None + freq: RefFreq | str | None = None + obs_date_anchor: ObsDateAnchor | str = "end" release_policy: ReleaseSelectionPolicy = "latest" @@ -299,7 +308,15 @@ def coerce_snapshot_series_spec( if not isinstance(spec, Mapping): raise PITContractError("Snapshot series spec must be SnapshotSeriesSpec or a mapping.") - allowed_keys = {"series_key", "alias", "start_ref", "end_ref", "release_policy"} + allowed_keys = { + "series_key", + "alias", + "start_ref", + "end_ref", + "freq", + "obs_date_anchor", + "release_policy", + } unknown_keys = sorted(set(spec.keys()) - allowed_keys) if unknown_keys: raise PITContractError(f"Unknown snapshot series spec keys: {unknown_keys}") @@ -309,8 +326,14 @@ def coerce_snapshot_series_spec( raise PITContractError("Snapshot series spec requires non-empty series_key.") alias = _cast_or_none(spec.get("alias")) - start_ref = _cast_or_none(spec.get("start_ref")) - end_ref = _cast_or_none(spec.get("end_ref")) + start_ref = spec.get("start_ref") + end_ref = spec.get("end_ref") + if isinstance(start_ref, str) and not start_ref.strip(): + start_ref = None + if isinstance(end_ref, str) and not end_ref.strip(): + end_ref = None + freq = normalize_ref_freq(spec.get("freq")) + obs_date_anchor = normalize_obs_date_anchor(spec.get("obs_date_anchor", "end")) release_policy = spec.get("release_policy", "latest") return SnapshotSeriesSpec( @@ -318,6 +341,8 @@ def coerce_snapshot_series_spec( alias=alias, start_ref=start_ref, end_ref=end_ref, + freq=freq, + obs_date_anchor=obs_date_anchor, release_policy=release_policy, ) diff --git a/alphaforge/pit/observation.py b/alphaforge/pit/observation.py index f6199ad..2dd9af1 100644 --- a/alphaforge/pit/observation.py +++ b/alphaforge/pit/observation.py @@ -18,7 +18,7 @@ import pandas as pd if TYPE_CHECKING: - from alphaforge.pit.release_rules import ReleaseRule + from alphaforge.time.release_rules import ReleaseRule PITMode = Literal["ALFRED_REALTIME", "DISCRETE_VINTAGES_SNAP", "NO_PIT"] diff --git a/alphaforge/pit/panel.py b/alphaforge/pit/panel.py index 4e40b2d..5c76735 100644 --- a/alphaforge/pit/panel.py +++ b/alphaforge/pit/panel.py @@ -34,15 +34,30 @@ def build_pit_panel( ------- Wide DataFrame with index=obs_date, columns=column_name. """ - series: dict[str, pd.Series] = {} - for col_name, skey in series_keys.items(): - snap = pit.get_snapshot(skey, asof, start=start, end=end) - series[col_name] = snap - - if not series: + if not series_keys: return pd.DataFrame() - df = pd.DataFrame(series) + batch = pit.get_snapshot_multi( + list(series_keys.values()), + asof, + start=start, + end=end, + ) + if batch.empty: + return pd.DataFrame(columns=list(series_keys.keys())) + + by_key = ( + batch.pivot_table( + index="obs_date", + columns="series_key", + values="value", + aggfunc="first", + ) + .sort_index() + ) + df = pd.DataFrame(index=by_key.index) + for col_name, skey in series_keys.items(): + df[col_name] = by_key[skey] if skey in by_key.columns else pd.Series(index=by_key.index) if align_freq is not None: new_idx = pd.date_range( diff --git a/alphaforge/pit/queries.py b/alphaforge/pit/queries.py new file mode 100644 index 0000000..46ac0ab --- /dev/null +++ b/alphaforge/pit/queries.py @@ -0,0 +1,202 @@ +"""Typed ref-period PIT query surfaces.""" + +from __future__ import annotations + +from dataclasses import dataclass +from datetime import date, datetime +from typing import Any, Mapping, TypeAlias + +import pandas as pd + +from alphaforge.time.ref_period import ( + ObsDateAnchor, + RefFreq, + RefPeriod, + coerce_ref_period, + normalize_obs_date_anchor, + normalize_ref_freq, +) + +from .exceptions import PITContractError + +RefPeriodLike: TypeAlias = RefPeriod | str | pd.Period | pd.Timestamp | date | datetime + + +def _coerce_series_key(value: object) -> str: + series_key = str(value).strip() + if not series_key: + raise PITContractError("series_key is required.") + return series_key + + +def _coerce_timestamp(value: object, *, field_name: str) -> pd.Timestamp: + try: + ts = pd.Timestamp(value) + except (TypeError, ValueError) as exc: + raise PITContractError(f"{field_name} must be datetime-like.") from exc + if pd.isna(ts): + raise PITContractError(f"{field_name} must be datetime-like.") + if ts.tzinfo is None: + ts = ts.tz_localize("UTC") + else: + ts = ts.tz_convert("UTC") + return ts + + +def _coerce_optional_timestamp(value: object | None, *, field_name: str) -> pd.Timestamp | None: + if value is None: + return None + return _coerce_timestamp(value, field_name=field_name) + + +def _coerce_ref( + value: object, + *, + field_name: str, + freq: RefFreq | None, + obs_date_anchor: ObsDateAnchor, +) -> RefPeriod: + try: + return coerce_ref_period(value, freq=freq, obs_date_anchor=obs_date_anchor) + except ValueError as exc: + raise PITContractError(f"{field_name}: {exc}") from exc + + +@dataclass(frozen=True) +class RefSnapshotQuery: + """Ref-period snapshot request for :meth:`alphaforge.pit.accessor.PITAccessor.snapshot_ref`.""" + + series_key: str + asof: pd.Timestamp + start_ref: RefPeriodLike | None = None + end_ref: RefPeriodLike | None = None + freq: RefFreq | str | None = None + obs_date_anchor: ObsDateAnchor | str = "end" + + +@dataclass(frozen=True) +class RefRevisionQuery: + """Ref-period revision request for :meth:`alphaforge.pit.accessor.PITAccessor.revisions_ref`.""" + + series_key: str + ref: RefPeriodLike + start_asof: pd.Timestamp | None = None + end_asof: pd.Timestamp | None = None + freq: RefFreq | str | None = None + obs_date_anchor: ObsDateAnchor | str = "end" + + +def coerce_ref_snapshot_query( + query: RefSnapshotQuery | Mapping[str, Any], +) -> RefSnapshotQuery: + """Normalize a ref-period snapshot query into a validated typed object.""" + + if isinstance(query, RefSnapshotQuery): + candidate = query + elif isinstance(query, Mapping): + try: + candidate = RefSnapshotQuery( + series_key=query["series_key"], + asof=query["asof"], + start_ref=query.get("start_ref"), + end_ref=query.get("end_ref"), + freq=query.get("freq"), + obs_date_anchor=query.get("obs_date_anchor", "end"), + ) + except KeyError as exc: + raise PITContractError(f"Ref snapshot query missing required field: {exc.args[0]}") from exc + else: + raise PITContractError("Ref snapshot query must be RefSnapshotQuery or a mapping.") + + series_key = _coerce_series_key(candidate.series_key) + asof = _coerce_timestamp(candidate.asof, field_name="asof") + obs_date_anchor = normalize_obs_date_anchor(candidate.obs_date_anchor) + freq = normalize_ref_freq(candidate.freq) + + start_ref: RefPeriod | None = None + if candidate.start_ref is not None: + start_ref = _coerce_ref( + candidate.start_ref, + field_name="start_ref", + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + if freq is None: + freq = start_ref.freq + + end_ref: RefPeriod | None = None + if candidate.end_ref is not None: + end_ref = _coerce_ref( + candidate.end_ref, + field_name="end_ref", + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + if freq is None: + freq = end_ref.freq + + if freq is None: + raise PITContractError( + "Ref snapshot query requires freq or at least one bounded reference period." + ) + + if start_ref is not None and end_ref is not None: + if start_ref.obs_date(anchor=obs_date_anchor) > end_ref.obs_date(anchor=obs_date_anchor): + raise PITContractError("start_ref must be <= end_ref.") + + return RefSnapshotQuery( + series_key=series_key, + asof=asof, + start_ref=start_ref, + end_ref=end_ref, + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + + +def coerce_ref_revision_query( + query: RefRevisionQuery | Mapping[str, Any], +) -> RefRevisionQuery: + """Normalize a ref-period revision query into a validated typed object.""" + + if isinstance(query, RefRevisionQuery): + candidate = query + elif isinstance(query, Mapping): + try: + candidate = RefRevisionQuery( + series_key=query["series_key"], + ref=query["ref"], + start_asof=query.get("start_asof"), + end_asof=query.get("end_asof"), + freq=query.get("freq"), + obs_date_anchor=query.get("obs_date_anchor", "end"), + ) + except KeyError as exc: + raise PITContractError( + f"Ref revision query missing required field: {exc.args[0]}" + ) from exc + else: + raise PITContractError("Ref revision query must be RefRevisionQuery or a mapping.") + + series_key = _coerce_series_key(candidate.series_key) + obs_date_anchor = normalize_obs_date_anchor(candidate.obs_date_anchor) + freq = normalize_ref_freq(candidate.freq) + ref = _coerce_ref( + candidate.ref, + field_name="ref", + freq=freq, + obs_date_anchor=obs_date_anchor, + ) + start_asof = _coerce_optional_timestamp(candidate.start_asof, field_name="start_asof") + end_asof = _coerce_optional_timestamp(candidate.end_asof, field_name="end_asof") + if start_asof is not None and end_asof is not None and start_asof > end_asof: + raise PITContractError("start_asof must be <= end_asof.") + + return RefRevisionQuery( + series_key=series_key, + ref=ref, + start_asof=start_asof, + end_asof=end_asof, + freq=ref.freq, + obs_date_anchor=obs_date_anchor, + ) diff --git a/alphaforge/pit/release_rules.py b/alphaforge/pit/release_rules.py index ff8e30b..ef59305 100644 --- a/alphaforge/pit/release_rules.py +++ b/alphaforge/pit/release_rules.py @@ -1,280 +1,6 @@ -"""Release schedule rules for macro series publication timing. +"""Compatibility shim for PIT release-rule imports. -Each rule class models a real-world release schedule pattern and exposes -``expected_release_date(obs_date)`` to compute the day a value is expected -to become publicly available. - -These rules are an *expectation layer* — realized PIT timestamps from -AlphaForge always take precedence when available. +The canonical public path is now ``alphaforge.time.release_rules``. """ -from __future__ import annotations - -from abc import ABC, abstractmethod -from dataclasses import dataclass -from datetime import date -from typing import Any - -import pandas as pd -from pandas.tseries.holiday import USFederalHolidayCalendar -from pandas.tseries.offsets import CustomBusinessDay - -_US_BD = CustomBusinessDay(calendar=USFederalHolidayCalendar()) - -# --------------------------------------------------------------------------- -# Registry for tagged-union YAML parsing (populated by @_register below) -# --------------------------------------------------------------------------- - -RULE_REGISTRY: dict[str, type] = {} - - -# --------------------------------------------------------------------------- -# Base -# --------------------------------------------------------------------------- - - -@dataclass(frozen=True) -class ReleaseRule(ABC): - """Base class for publication schedule rules.""" - - rule_type: str = "" # overridden by subclass class-var - - @abstractmethod - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - """Return the expected publication date for a given observation date. - - Parameters - ---------- - obs_date : date - The observation / reference-period end date. - release_number : int | None - For multi-release series (e.g. GDP), selects which release - (1 = advance, 2 = preliminary, …). Ignored by single-release - rules. - """ - - # Serialization helpers for round-trip YAML - def to_dict(self) -> dict[str, Any]: - """Serialize to a dict suitable for YAML output.""" - d: dict[str, Any] = {"type": self.rule_type} - for k, v in self.__dict__.items(): - if k == "rule_type": - continue - d[k] = v - return d - - @staticmethod - def from_dict(d: dict[str, Any]) -> ReleaseRule: - """Reconstruct a ReleaseRule from a YAML-style dict.""" - kwargs = dict(d) - type_key = kwargs.pop("type") - cls = RULE_REGISTRY.get(type_key) - if cls is None: - raise ValueError( - f"Unknown release rule type '{type_key}'. " - f"Known types: {sorted(RULE_REGISTRY)}" - ) - return cls(**kwargs) - - -def _register(cls: type[ReleaseRule]) -> type[ReleaseRule]: - """Decorator that adds a ReleaseRule subclass to the registry.""" - RULE_REGISTRY[cls.rule_type] = cls - return cls - - -# --------------------------------------------------------------------------- -# Concrete rule classes -# --------------------------------------------------------------------------- - - -def _month_start(ref: date, offset_months: int) -> date: - """Return the 1st day of the month that is *offset_months* after *ref*.""" - m = ref.month + offset_months - y = ref.year + (m - 1) // 12 - m = (m - 1) % 12 + 1 - return date(y, m, 1) - - -def _nth_business_day(anchor: date, n: int) -> date: - """Return the *n*-th US business day on or after *anchor*.""" - ts = pd.Timestamp(anchor) - # Move to the n-th business day (1-indexed) - result = ts + (n - 1) * _US_BD - # If anchor itself is not a business day, the offset already handles it - # but we need to make sure we start counting from the first valid day - first_bd = ts + 0 * _US_BD # snap to first business day on or after - result = first_bd + (n - 1) * _US_BD - return result.date() - - -@_register -@dataclass(frozen=True) -class NthBusinessDay(ReleaseRule): - """Published on the n-th US business day relative to an anchor month. - - Example: Employment Situation — 1st business day of the following month. - """ - - rule_type: str = "nth_business_day" - n: int = 1 - anchor: str = "following_month" # "following_month" | "same_month" - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - if self.anchor == "following_month": - start = _month_start(obs_date, 1) - elif self.anchor == "same_month": - start = _month_start(obs_date, 0) - else: - raise ValueError(f"Unknown anchor: {self.anchor}") - return _nth_business_day(start, self.n) - - -@_register -@dataclass(frozen=True) -class NthWeekday(ReleaseRule): - """Published on the n-th occurrence of a weekday in the anchor month. - - Example: CPI — typically released around the 10th–15th business day - of the following month, often modeled as ~2nd or 3rd week. - """ - - rule_type: str = "nth_weekday" - n: int = 1 - weekday: str = "Friday" # Monday–Sunday - anchor: str = "following_month" - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - _weekday_map = { - "Monday": 0, - "Tuesday": 1, - "Wednesday": 2, - "Thursday": 3, - "Friday": 4, - "Saturday": 5, - "Sunday": 6, - } - target_wd = _weekday_map[self.weekday] - - if self.anchor == "following_month": - first = _month_start(obs_date, 1) - elif self.anchor == "same_month": - first = _month_start(obs_date, 0) - else: - raise ValueError(f"Unknown anchor: {self.anchor}") - - # Find the first occurrence of target weekday in the month - days_ahead = (target_wd - first.weekday()) % 7 - first_occurrence = first + pd.Timedelta(days=days_ahead) - # Advance to the n-th occurrence - result = first_occurrence + pd.Timedelta(weeks=self.n - 1) - return result.date() if isinstance(result, pd.Timestamp) else result - - -@_register -@dataclass(frozen=True) -class CalendarDay(ReleaseRule): - """Published on a specific calendar day of the anchor month. - - Example: "15th of the following month." - """ - - rule_type: str = "calendar_day" - day: int = 15 - anchor: str = "following_month" - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - if self.anchor == "following_month": - start = _month_start(obs_date, 1) - elif self.anchor == "same_month": - start = _month_start(obs_date, 0) - elif self.anchor == "two_months_later": - start = _month_start(obs_date, 2) - else: - raise ValueError(f"Unknown anchor: {self.anchor}") - return start.replace(day=self.day) - - -@_register -@dataclass(frozen=True) -class FixedLagMonths(ReleaseRule): - """Published a fixed number of months after the observation month. - - Fallback rule for series where the exact schedule is not worth encoding. - """ - - rule_type: str = "fixed_lag_months" - months: int = 1 - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - return _month_start(obs_date, self.months) - - -@_register -@dataclass(frozen=True) -class QuarterlyRelease(ReleaseRule): - """GDP-style multi-release schedule. - - Models advance / preliminary / final releases with configurable lag - (in months after the quarter-end). - """ - - rule_type: str = "quarterly_release" - advance_lag_months: int = 1 - preliminary_lag_months: int = 2 - final_lag_months: int = 3 - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - rn = release_number or 1 - if rn == 1: - lag = self.advance_lag_months - elif rn == 2: - lag = self.preliminary_lag_months - else: - lag = self.final_lag_months - return _month_start(obs_date, lag) - - -@_register -@dataclass(frozen=True) -class WeeklyRelease(ReleaseRule): - """Weekly series released on a fixed weekday with a lag. - - Example: Initial Claims — released Thursday for the prior Saturday week. - """ - - rule_type: str = "weekly" - release_weekday: str = "Thursday" - lag_days: int = 5 - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - return obs_date + pd.Timedelta(days=self.lag_days) - - -@_register -@dataclass(frozen=True) -class CustomRule(ReleaseRule): - """Free-text description for exotic or not-yet-modeled schedules.""" - - rule_type: str = "custom" - description: str = "" - approximate_lag_months: int = 1 - - def expected_release_date( - self, obs_date: date, release_number: int | None = None - ) -> date: - return _month_start(obs_date, self.approximate_lag_months) +from alphaforge.time.release_rules import * # noqa: F403 diff --git a/alphaforge/pit/target.py b/alphaforge/pit/target.py index 307a505..e34633d 100644 --- a/alphaforge/pit/target.py +++ b/alphaforge/pit/target.py @@ -7,7 +7,6 @@ from __future__ import annotations -import re from collections.abc import Iterable from dataclasses import dataclass from datetime import date @@ -16,6 +15,7 @@ import pandas as pd from alphaforge.pit.adapters.base import PITAdapter +from alphaforge.time.ref_period import RefFreq, RefPeriod, coerce_ref_period if TYPE_CHECKING: from alphaforge.pit.catalog import SeriesCatalog @@ -51,27 +51,7 @@ class TargetPolicy: # --------------------------------------------------------------------------- -def _quarter_end_from_string(ref_quarter: str) -> date: - """Parse ``YYYYQn`` and return the calendar quarter-end date.""" - match = re.match(r"^(\d{4})[Qq](\d+)$", str(ref_quarter).strip()) - if not match: - raise ValueError(f"Expected quarterly reference in format YYYYQn, got {ref_quarter!r}") - year = int(match.group(1)) - quarter = int(match.group(2)) - if quarter not in {1, 2, 3, 4}: - raise ValueError(f"Invalid quarter: {quarter}. Must be 1, 2, 3, or 4") - month = quarter * 3 - # Calendar quarter-end day - if month in {3, 12}: - day = 31 - elif month in {6, 9}: - day = 30 - else: # pragma: no cover — unreachable for valid quarters - raise ValueError(f"Invalid quarter month: {month}") - return date(year, month, day) - - -def quarter_end_date(ref_quarter: str | pd.Period) -> date: +def quarter_end_date(ref_quarter: str | pd.Period | RefPeriod) -> date: """Return the quarter-end date for a reference quarter. Args: @@ -84,30 +64,34 @@ def quarter_end_date(ref_quarter: str | pd.Period) -> date: Raises: ValueError: If the reference quarter does not match ``YYYYQn``. """ - return _quarter_end_from_string(str(ref_quarter)) + try: + return coerce_ref_period(ref_quarter, freq=RefFreq.Q).end_obs_date().date() + except ValueError as exc: + raise ValueError( + f"Expected quarterly reference in format YYYYQn, got {ref_quarter!r}" + ) from exc -def quarter_start_date(ref_quarter: str | pd.Period) -> date: +def quarter_start_date(ref_quarter: str | pd.Period | RefPeriod) -> date: """Return the quarter-start date for a reference quarter.""" - quarter_end = quarter_end_date(ref_quarter) - if quarter_end.month == 3: - return date(quarter_end.year, 1, 1) - if quarter_end.month == 6: - return date(quarter_end.year, 4, 1) - if quarter_end.month == 9: - return date(quarter_end.year, 7, 1) - return date(quarter_end.year, 10, 1) + try: + return coerce_ref_period(ref_quarter, freq=RefFreq.Q).start_obs_date().date() + except ValueError as exc: + raise ValueError( + f"Expected quarterly reference in format YYYYQn, got {ref_quarter!r}" + ) from exc def quarter_obs_date( - ref_quarter: str | pd.Period, obs_date_anchor: Literal["end", "start"] = "end" + ref_quarter: str | pd.Period | RefPeriod, + obs_date_anchor: Literal["end", "start"] = "end", ) -> date: """Return the quarter observation date under the configured anchor policy.""" - if obs_date_anchor == "end": - return quarter_end_date(ref_quarter) - if obs_date_anchor == "start": - return quarter_start_date(ref_quarter) - raise ValueError(f"obs_date_anchor must be 'end' or 'start', got {obs_date_anchor}") + return ( + coerce_ref_period(ref_quarter, freq=RefFreq.Q) + .obs_date(anchor=obs_date_anchor) + .date() + ) # --------------------------------------------------------------------------- @@ -128,8 +112,10 @@ def resolve_target_obs_date_anchor( 2. ``"auto"`` uses catalog metadata ``obs_date_anchor`` for the target series. 3. If metadata is missing/unknown, fallback is ``"end"``. """ - if target_obs_date_anchor in {"start", "end"}: - return target_obs_date_anchor # type: ignore[return-value] + if target_obs_date_anchor == "start": + return "start" + if target_obs_date_anchor == "end": + return "end" if target_obs_date_anchor != "auto": raise ValueError( f"target_obs_date_anchor must be one of 'auto', 'start', 'end', got " @@ -137,8 +123,10 @@ def resolve_target_obs_date_anchor( ) if catalog is not None: meta = catalog.get(target_series_key) - if meta is not None and meta.obs_date_anchor in {"start", "end"}: - return meta.obs_date_anchor # type: ignore[return-value] + if meta is not None and meta.obs_date_anchor == "start": + return "start" + if meta is not None and meta.obs_date_anchor == "end": + return "end" return "end" @@ -204,7 +192,7 @@ def list_quarterly_target_releases_asof_multi( } ) try: - releases = adapter.list_pit_observations_asof_multi(requests) # type: ignore[attr-defined] + releases = adapter.list_pit_observations_asof_multi(requests) except NotImplementedError: releases = pd.DataFrame() out: dict[str, pd.DataFrame] = {} diff --git a/alphaforge/pit/tasks.py b/alphaforge/pit/tasks.py index c22a645..e759693 100644 --- a/alphaforge/pit/tasks.py +++ b/alphaforge/pit/tasks.py @@ -350,8 +350,8 @@ def _normalize_asof_grid( def _resolve_snapshot_window( - start_ref: str | None, - end_ref: str | None, + start_ref: object | None, + end_ref: object | None, ) -> tuple[pd.Timestamp | None, pd.Timestamp | None]: start_obs = RefPeriod.parse(start_ref).end_obs_date() if start_ref is not None else None end_obs = RefPeriod.parse(end_ref).end_obs_date() if end_ref is not None else None diff --git a/alphaforge/time/__init__.py b/alphaforge/time/__init__.py index 23ed117..a9d80d6 100644 --- a/alphaforge/time/__init__.py +++ b/alphaforge/time/__init__.py @@ -1,3 +1,40 @@ -from .ref_period import RefFreq, RefPeriod +from .missingness import MissingnessReason, classify_missingness +from .ref_period import ( + ObsDateAnchor, + RefFreq, + RefPeriod, + coerce_ref_period, + normalize_obs_date_anchor, + normalize_ref_freq, +) +from .release_rules import ( + RULE_REGISTRY, + CalendarDay, + CustomRule, + FixedLagMonths, + NthBusinessDay, + NthWeekday, + QuarterlyRelease, + ReleaseRule, + WeeklyRelease, +) -__all__ = ["RefFreq", "RefPeriod"] +__all__ = [ + "RefFreq", + "RefPeriod", + "ObsDateAnchor", + "coerce_ref_period", + "normalize_ref_freq", + "normalize_obs_date_anchor", + "RULE_REGISTRY", + "ReleaseRule", + "NthBusinessDay", + "NthWeekday", + "CalendarDay", + "FixedLagMonths", + "QuarterlyRelease", + "WeeklyRelease", + "CustomRule", + "MissingnessReason", + "classify_missingness", +] diff --git a/alphaforge/time/missingness.py b/alphaforge/time/missingness.py new file mode 100644 index 0000000..5c20e0e --- /dev/null +++ b/alphaforge/time/missingness.py @@ -0,0 +1,84 @@ +"""Missingness taxonomy and classification for temporal semantics.""" + +from __future__ import annotations + +from datetime import date +from enum import Enum + +from .release_rules import ReleaseRule + + +class MissingnessReason(str, Enum): + """Why a point-in-time panel cell is missing.""" + + STRUCTURAL = "structural" + FUTURE = "future" + RAGGED_EDGE = "ragged_edge" + TRUE_MISSING = "true_missing" + + +def classify_missingness( + *, + obs_date: date, + asof_date: date, + series_frequency: str, + panel_frequency: str = "M", + release_rule: ReleaseRule | None = None, + publication_lag_months: int | None = None, + realized_release_date: date | None = None, +) -> MissingnessReason: + """Classify why a cell is missing at a given as-of date.""" + if _is_structural(obs_date, series_frequency, panel_frequency): + return MissingnessReason.STRUCTURAL + + if obs_date > asof_date: + return MissingnessReason.FUTURE + + expected_release = _resolve_expected_date( + obs_date, + release_rule=release_rule, + publication_lag_months=publication_lag_months, + realized_release_date=realized_release_date, + ) + if expected_release is None: + return MissingnessReason.TRUE_MISSING + if asof_date < expected_release: + return MissingnessReason.RAGGED_EDGE + return MissingnessReason.TRUE_MISSING + + +_QUARTER_END_MONTHS = {3, 6, 9, 12} + + +def _is_structural( + obs_date: date, + series_frequency: str, + panel_frequency: str, +) -> bool: + series_freq = series_frequency.upper() + panel_freq = panel_frequency.upper() + if series_freq == "Q" and panel_freq == "M": + return obs_date.month not in _QUARTER_END_MONTHS + return False + + +def _resolve_expected_date( + obs_date: date, + *, + release_rule: ReleaseRule | None, + publication_lag_months: int | None, + realized_release_date: date | None, +) -> date | None: + if realized_release_date is not None: + return realized_release_date + if release_rule is not None: + return release_rule.expected_release_date(obs_date) + if publication_lag_months is not None: + month = obs_date.month + publication_lag_months + year = obs_date.year + (month - 1) // 12 + month = (month - 1) % 12 + 1 + return date(year, month, 1) + return None + + +__all__ = ["MissingnessReason", "classify_missingness"] diff --git a/alphaforge/time/ref_period.py b/alphaforge/time/ref_period.py index 21e896e..ded825e 100644 --- a/alphaforge/time/ref_period.py +++ b/alphaforge/time/ref_period.py @@ -2,11 +2,16 @@ import re from dataclasses import dataclass +from datetime import date, datetime from enum import Enum +from typing import Literal import pandas as pd from pandas.tseries.offsets import MonthEnd +ObsDateAnchor = Literal["start", "end"] +RefPeriodInput = "RefPeriod | str | pd.Period | pd.Timestamp | date | datetime" + class RefFreq(str, Enum): A = "A" @@ -14,7 +19,7 @@ class RefFreq(str, Enum): M = "M" -def _ts_utc_midnight(value: pd.Timestamp | str) -> pd.Timestamp: +def _ts_utc_midnight(value: pd.Timestamp | date | datetime | str) -> pd.Timestamp: ts = pd.Timestamp(value) if ts.tzinfo is None: ts = ts.tz_localize("UTC") @@ -23,6 +28,59 @@ def _ts_utc_midnight(value: pd.Timestamp | str) -> pd.Timestamp: return ts.floor("D") +def normalize_ref_freq(value: RefFreq | str | None) -> RefFreq | None: + if value is None: + return None + if isinstance(value, RefFreq): + return value + + text = str(value).strip().upper() + aliases = { + "A": RefFreq.A, + "Y": RefFreq.A, + "A-DEC": RefFreq.A, + "Y-DEC": RefFreq.A, + "YEAR": RefFreq.A, + "ANNUAL": RefFreq.A, + "Q": RefFreq.Q, + "Q-DEC": RefFreq.Q, + "QUARTER": RefFreq.Q, + "QUARTERLY": RefFreq.Q, + "M": RefFreq.M, + "MONTH": RefFreq.M, + "MONTHLY": RefFreq.M, + } + try: + return aliases[text] + except KeyError as exc: + raise ValueError(f"Unsupported reference frequency: {value!r}") from exc + + +def normalize_obs_date_anchor(anchor: ObsDateAnchor | str) -> ObsDateAnchor: + text = str(anchor).strip().lower() + if text not in {"start", "end"}: + raise ValueError(f"obs_date_anchor must be 'start' or 'end', got {anchor!r}") + return text # type: ignore[return-value] + + +def _coerce_explicit_datetime_input( + value: object, +) -> pd.Timestamp | None: + if isinstance(value, (pd.Timestamp, datetime, date)): + return _ts_utc_midnight(value) + if isinstance(value, str): + text = value.strip() + if not text: + raise ValueError("Reference period string is required.") + if not re.match(r"^\d{4}([-/]\d{2}){1,2}$", text): + return None + try: + return _ts_utc_midnight(text) + except (TypeError, ValueError): + return None + return None + + @dataclass(frozen=True) class RefPeriod: freq: RefFreq @@ -30,43 +88,103 @@ class RefPeriod: period: int @staticmethod - def parse(s: str) -> "RefPeriod": - text = str(s).strip() - if not text: - raise ValueError("Reference period string is required.") + def parse( + value: object, + *, + freq: RefFreq | str | None = None, + obs_date_anchor: ObsDateAnchor | str = "end", + ) -> "RefPeriod": + requested_freq = normalize_ref_freq(freq) + ref_period: RefPeriod | None - match = re.match(r"^(\d{4})[Qq]([1-4])$", text) - if match: - return RefPeriod(RefFreq.Q, int(match.group(1)), int(match.group(2))) - - match = re.match(r"^(\d{4})-(\d{2})-(\d{2})$", text) - if match: - try: - ts = pd.Timestamp(text) - except ValueError as exc: - raise ValueError(f"Invalid reference period date: {text}") from exc - if ts.day != (ts + MonthEnd(0)).day: - raise ValueError( - "Date reference periods must be month-end (YYYY-MM-DD)." + if isinstance(value, RefPeriod): + ref_period = value + elif isinstance(value, pd.Period): + ref_period = RefPeriod.from_period(value) + else: + text = str(value).strip() + if not text: + raise ValueError("Reference period string is required.") + + ref_period = _parse_text_ref_period(text) + if ( + ref_period is not None + and requested_freq is not None + and ref_period.freq != requested_freq + ): + obs_ts = _coerce_explicit_datetime_input(value) + if obs_ts is not None: + ref_period = RefPeriod.from_obs_date( + obs_ts, + freq=requested_freq, + obs_date_anchor=obs_date_anchor, + ) + if ref_period is None: + obs_ts = _coerce_explicit_datetime_input(value) + if obs_ts is None or requested_freq is None: + raise ValueError( + "Invalid reference period format. Expected YYYY, YYYYQq, YYYY-MM, " + "YYYY/MM, YYYY-MM-DD, pandas Period, or an explicit observation " + "date with freq=... ." + ) + ref_period = RefPeriod.from_obs_date( + obs_ts, + freq=requested_freq, + obs_date_anchor=obs_date_anchor, ) - return RefPeriod(RefFreq.M, ts.year, ts.month) - match = re.match(r"^(\d{4})[-/](\d{2})$", text) - if match: - year = int(match.group(1)) - month = int(match.group(2)) - if not 1 <= month <= 12: - raise ValueError(f"Invalid reference period month: {text}") - return RefPeriod(RefFreq.M, year, month) + if ref_period is None: + raise ValueError("Reference period could not be resolved.") + if requested_freq is not None and ref_period.freq != requested_freq: + raise ValueError( + "Reference period frequency does not match the requested frequency." + ) + return ref_period + + @staticmethod + def from_period(period: pd.Period) -> "RefPeriod": + freq = normalize_ref_freq(period.freqstr) + if freq == RefFreq.A: + return RefPeriod(freq=freq, year=period.year, period=1) + if freq == RefFreq.Q: + return RefPeriod(freq=freq, year=period.year, period=period.quarter) + if freq == RefFreq.M: + return RefPeriod(freq=freq, year=period.year, period=period.month) + raise ValueError(f"Unsupported reference frequency: {period.freqstr}") + + @staticmethod + def from_obs_date( + value: pd.Timestamp | date | datetime | str, + *, + freq: RefFreq | str, + obs_date_anchor: ObsDateAnchor | str = "end", + ) -> "RefPeriod": + resolved_freq = normalize_ref_freq(freq) + if resolved_freq is None: + raise ValueError("freq is required when normalizing observation dates.") + + obs_ts = _ts_utc_midnight(value) + anchor = normalize_obs_date_anchor(obs_date_anchor) - match = re.match(r"^(\d{4})$", text) - if match: - return RefPeriod(RefFreq.A, int(match.group(1)), 1) + if resolved_freq == RefFreq.A: + candidate = RefPeriod(freq=resolved_freq, year=obs_ts.year, period=1) + elif resolved_freq == RefFreq.Q: + candidate = RefPeriod( + freq=resolved_freq, + year=obs_ts.year, + period=((obs_ts.month - 1) // 3) + 1, + ) + elif resolved_freq == RefFreq.M: + candidate = RefPeriod(freq=resolved_freq, year=obs_ts.year, period=obs_ts.month) + else: + raise ValueError(f"Unsupported reference frequency: {resolved_freq}") - raise ValueError( - "Invalid reference period format. Expected YYYY, YYYYQq, YYYY-MM, " - "YYYY/MM, or YYYY-MM-DD." - ) + if candidate.obs_date(anchor=anchor) != obs_ts: + raise ValueError( + "obs_date does not match the requested reference period under the " + f"{anchor!r} anchor." + ) + return candidate def to_key(self) -> str: if self.freq == RefFreq.A: @@ -77,33 +195,82 @@ def to_key(self) -> str: return f"{self.year:04d}-{self.period:02d}" raise ValueError(f"Unsupported reference frequency: {self.freq}") + def __str__(self) -> str: + return self.to_key() + + def start_obs_date(self) -> pd.Timestamp: + if self.freq == RefFreq.A: + return pd.Timestamp(self.year, 1, 1, tz="UTC") + if self.freq == RefFreq.Q: + month = ((self.period - 1) * 3) + 1 + return pd.Timestamp(self.year, month, 1, tz="UTC") + if self.freq == RefFreq.M: + return pd.Timestamp(self.year, self.period, 1, tz="UTC") + raise ValueError(f"Unsupported reference frequency: {self.freq}") + def end_obs_date(self) -> pd.Timestamp: + return self.obs_date(anchor="end") + + def obs_date(self, anchor: ObsDateAnchor | str = "end") -> pd.Timestamp: + resolved_anchor = normalize_obs_date_anchor(anchor) + if resolved_anchor == "start": + return self.start_obs_date().floor("D") if self.freq == RefFreq.A: - ts = pd.Timestamp(self.year, 12, 31, tz="UTC") + end = pd.Timestamp(self.year, 12, 31, tz="UTC") elif self.freq == RefFreq.Q: month = self.period * 3 - ts = pd.Timestamp(self.year, month, 1, tz="UTC") + MonthEnd(0) - elif self.freq == RefFreq.M: - ts = pd.Timestamp(self.year, self.period, 1, tz="UTC") + MonthEnd(0) + end = pd.Timestamp(self.year, month, 1, tz="UTC") + MonthEnd(0) else: - raise ValueError(f"Unsupported reference frequency: {self.freq}") - return ts.floor("D") + end = self.start_obs_date() + MonthEnd(0) + return end.floor("D") @staticmethod - def from_obs_date_end(ts: pd.Timestamp, freq: RefFreq) -> "RefPeriod": - obs_ts = _ts_utc_midnight(ts) - if freq == RefFreq.A: - period = 1 - elif freq == RefFreq.Q: - period = (obs_ts.month - 1) // 3 + 1 - elif freq == RefFreq.M: - period = obs_ts.month - else: - raise ValueError(f"Unsupported reference frequency: {freq}") + def from_obs_date_end(ts: pd.Timestamp, freq: RefFreq | str) -> "RefPeriod": + return RefPeriod.from_obs_date(ts, freq=freq, obs_date_anchor="end") - candidate = RefPeriod(freq=freq, year=obs_ts.year, period=period) - if candidate.end_obs_date() != obs_ts: - raise ValueError( - "obs_date_end does not match the end of the requested reference period." - ) - return candidate + +def coerce_ref_period( + value: object, + *, + freq: RefFreq | str | None = None, + obs_date_anchor: ObsDateAnchor | str = "end", +) -> RefPeriod: + return RefPeriod.parse(value, freq=freq, obs_date_anchor=obs_date_anchor) + + +def _parse_text_ref_period(text: str) -> RefPeriod | None: + match = re.match(r"^(\d{4})[Qq]([1-4])$", text) + if match: + return RefPeriod(RefFreq.Q, int(match.group(1)), int(match.group(2))) + + match = re.match(r"^(\d{4})[-/](\d{2})$", text) + if match: + year = int(match.group(1)) + month = int(match.group(2)) + if not 1 <= month <= 12: + raise ValueError(f"Invalid reference period month: {text}") + return RefPeriod(RefFreq.M, year, month) + + match = re.match(r"^(\d{4})$", text) + if match: + return RefPeriod(RefFreq.A, int(match.group(1)), 1) + + match = re.match(r"^(\d{4})-(\d{2})-(\d{2})$", text) + if match: + ts = pd.Timestamp(text) + if ts.day != (ts + MonthEnd(0)).day: + return None + return RefPeriod(RefFreq.M, ts.year, ts.month) + + return None + + +__all__ = [ + "ObsDateAnchor", + "RefPeriodInput", + "RefFreq", + "RefPeriod", + "normalize_ref_freq", + "normalize_obs_date_anchor", + "coerce_ref_period", +] diff --git a/alphaforge/time/release_rules.py b/alphaforge/time/release_rules.py new file mode 100644 index 0000000..05c41ac --- /dev/null +++ b/alphaforge/time/release_rules.py @@ -0,0 +1,249 @@ +"""Release schedule rules for publication and availability semantics. + +These rules model expected public release timing for a reference-period +observation. They are an expectation layer: realized PIT timestamps always take +precedence when available from source data. +""" + +from __future__ import annotations + +from abc import ABC, abstractmethod +from dataclasses import dataclass +from datetime import date +from typing import Any + +import pandas as pd +from pandas.tseries.holiday import USFederalHolidayCalendar +from pandas.tseries.offsets import CustomBusinessDay + +_US_BD = CustomBusinessDay(calendar=USFederalHolidayCalendar()) + +RULE_REGISTRY: dict[str, type["ReleaseRule"]] = {} + +_WEEKDAY_MAP = { + "Monday": 0, + "Tuesday": 1, + "Wednesday": 2, + "Thursday": 3, + "Friday": 4, + "Saturday": 5, + "Sunday": 6, +} + + +@dataclass(frozen=True) +class ReleaseRule(ABC): + """Base class for publication schedule rules.""" + + rule_type: str = "" + + @abstractmethod + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + """Return the expected publication date for an observation date.""" + + def to_dict(self) -> dict[str, Any]: + """Serialize the rule to a YAML-friendly mapping.""" + payload: dict[str, Any] = {"type": self.rule_type} + for key, value in self.__dict__.items(): + if key != "rule_type": + payload[key] = value + return payload + + @staticmethod + def from_dict(payload: dict[str, Any]) -> "ReleaseRule": + """Reconstruct a rule from a YAML-style mapping.""" + kwargs = dict(payload) + type_key = kwargs.pop("type") + cls = RULE_REGISTRY.get(type_key) + if cls is None: + raise ValueError( + f"Unknown release rule type '{type_key}'. " + f"Known types: {sorted(RULE_REGISTRY)}" + ) + return cls(**kwargs) + + +def _register(cls: type[ReleaseRule]) -> type[ReleaseRule]: + RULE_REGISTRY[cls.rule_type] = cls + return cls + + +def _month_start(ref: date, offset_months: int) -> date: + """Return the first day of the month *offset_months* after *ref*.""" + month = ref.month + offset_months + year = ref.year + (month - 1) // 12 + month = (month - 1) % 12 + 1 + return date(year, month, 1) + + +def _nth_business_day(anchor: date, n: int) -> date: + """Return the n-th US business day on or after *anchor*.""" + first_business_day = pd.Timestamp(anchor) + 0 * _US_BD + return (first_business_day + (n - 1) * _US_BD).date() + + +def _resolve_weekday(weekday: str) -> int: + try: + return _WEEKDAY_MAP[weekday] + except KeyError as exc: + raise ValueError( + f"Unknown weekday: {weekday!r}. Known values: {sorted(_WEEKDAY_MAP)}" + ) from exc + + +@_register +@dataclass(frozen=True) +class NthBusinessDay(ReleaseRule): + """Published on the n-th US business day of an anchor month.""" + + rule_type: str = "nth_business_day" + n: int = 1 + anchor: str = "following_month" + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + if self.anchor == "following_month": + start = _month_start(obs_date, 1) + elif self.anchor == "same_month": + start = _month_start(obs_date, 0) + else: + raise ValueError(f"Unknown anchor: {self.anchor}") + return _nth_business_day(start, self.n) + + +@_register +@dataclass(frozen=True) +class NthWeekday(ReleaseRule): + """Published on the n-th occurrence of a weekday in the anchor month.""" + + rule_type: str = "nth_weekday" + n: int = 1 + weekday: str = "Friday" + anchor: str = "following_month" + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + target_weekday = _resolve_weekday(self.weekday) + + if self.anchor == "following_month": + first = _month_start(obs_date, 1) + elif self.anchor == "same_month": + first = _month_start(obs_date, 0) + else: + raise ValueError(f"Unknown anchor: {self.anchor}") + + days_ahead = (target_weekday - first.weekday()) % 7 + first_occurrence = first + pd.Timedelta(days=days_ahead) + result = first_occurrence + pd.Timedelta(weeks=self.n - 1) + return result.date() if isinstance(result, pd.Timestamp) else result + + +@_register +@dataclass(frozen=True) +class CalendarDay(ReleaseRule): + """Published on a specific calendar day of the anchor month.""" + + rule_type: str = "calendar_day" + day: int = 15 + anchor: str = "following_month" + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + if self.anchor == "following_month": + start = _month_start(obs_date, 1) + elif self.anchor == "same_month": + start = _month_start(obs_date, 0) + elif self.anchor == "two_months_later": + start = _month_start(obs_date, 2) + else: + raise ValueError(f"Unknown anchor: {self.anchor}") + return start.replace(day=self.day) + + +@_register +@dataclass(frozen=True) +class FixedLagMonths(ReleaseRule): + """Published a fixed number of months after the observation month.""" + + rule_type: str = "fixed_lag_months" + months: int = 1 + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + return _month_start(obs_date, self.months) + + +@_register +@dataclass(frozen=True) +class QuarterlyRelease(ReleaseRule): + """GDP-style multi-release schedule.""" + + rule_type: str = "quarterly_release" + advance_lag_months: int = 1 + preliminary_lag_months: int = 2 + final_lag_months: int = 3 + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + release_rank = release_number or 1 + if release_rank == 1: + lag_months = self.advance_lag_months + elif release_rank == 2: + lag_months = self.preliminary_lag_months + else: + lag_months = self.final_lag_months + return _month_start(obs_date, lag_months) + + +@_register +@dataclass(frozen=True) +class WeeklyRelease(ReleaseRule): + """Weekly series released on a fixed weekday after a lag.""" + + rule_type: str = "weekly" + release_weekday: str = "Thursday" + lag_days: int = 5 + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + anchor = obs_date + pd.Timedelta(days=self.lag_days) + target_weekday = _resolve_weekday(self.release_weekday) + days_ahead = (target_weekday - anchor.weekday()) % 7 + result = anchor + pd.Timedelta(days=days_ahead) + return result.date() if isinstance(result, pd.Timestamp) else result + + +@_register +@dataclass(frozen=True) +class CustomRule(ReleaseRule): + """Free-text description for schedules that are not yet modeled.""" + + rule_type: str = "custom" + description: str = "" + approximate_lag_months: int = 1 + + def expected_release_date( + self, obs_date: date, release_number: int | None = None + ) -> date: + return _month_start(obs_date, self.approximate_lag_months) + + +__all__ = [ + "RULE_REGISTRY", + "ReleaseRule", + "NthBusinessDay", + "NthWeekday", + "CalendarDay", + "FixedLagMonths", + "QuarterlyRelease", + "WeeklyRelease", + "CustomRule", +] diff --git a/benchmarks/__init__.py b/benchmarks/__init__.py new file mode 100644 index 0000000..9db33a0 --- /dev/null +++ b/benchmarks/__init__.py @@ -0,0 +1,14 @@ +"""Benchmark harnesses for core-platform regression tracking.""" + +from __future__ import annotations + +from importlib import import_module +from typing import Any + +__all__ = ["run_pit_contract_benchmarks"] + + +def __getattr__(name: str) -> Any: + if name == "run_pit_contract_benchmarks": + return import_module(".pit", __name__).run_pit_contract_benchmarks + raise AttributeError(f"module {__name__!r} has no attribute {name!r}") diff --git a/benchmarks/pit.py b/benchmarks/pit.py new file mode 100644 index 0000000..34d1551 --- /dev/null +++ b/benchmarks/pit.py @@ -0,0 +1,140 @@ +from __future__ import annotations + +import argparse +import json +import statistics +import tempfile +import time +from typing import Any + +import pandas as pd + +from alphaforge.pit.accessor import PITAccessor +from alphaforge.pit.queries import RefSnapshotQuery +from alphaforge.store.duckdb_parquet import DuckDBParquetStore + + +def _quarter_end_dates(periods: int) -> list[pd.Timestamp]: + return [ + period.to_timestamp(how="end").normalize() + for period in pd.period_range("2010Q1", periods=periods, freq="Q") + ] + + +def _sample_observations( + *, + periods: int, + series_count: int, + revisions_per_period: int, +) -> tuple[pd.DataFrame, list[str]]: + obs_dates = _quarter_end_dates(periods) + series_keys = [f"SERIES_{index + 1}" for index in range(series_count)] + + rows: list[dict[str, Any]] = [] + for series_index, series_key in enumerate(series_keys): + for period_index, obs_date in enumerate(obs_dates): + base_value = float((series_index + 1) * 100 + period_index) + for revision_index in range(revisions_per_period): + rows.append( + { + "series_key": series_key, + "obs_date": obs_date, + "asof_utc": pd.Timestamp(obs_date, tz="UTC") + + pd.Timedelta(days=15 + revision_index * 15), + "value": base_value + revision_index * 0.1, + } + ) + return pd.DataFrame(rows), series_keys + + +def _time_call(fn, *, iterations: int) -> list[float]: + fn() + samples: list[float] = [] + for _ in range(iterations): + started = time.perf_counter() + fn() + samples.append((time.perf_counter() - started) * 1000.0) + return samples + + +def run_pit_contract_benchmarks( + *, + iterations: int = 5, + periods: int = 40, + series_count: int = 3, + revisions_per_period: int = 2, +) -> dict[str, Any]: + with tempfile.TemporaryDirectory(prefix="alphaforge-bench-") as tmpdir: + store = DuckDBParquetStore(root=tmpdir) + pit = PITAccessor(store.conn()) + observations, series_keys = _sample_observations( + periods=periods, + series_count=series_count, + revisions_per_period=revisions_per_period, + ) + pit.upsert_pit_observations(observations) + + asof = observations["asof_utc"].max() + start_ref = pd.Period(observations["obs_date"].min(), freq="Q") + end_ref = pd.Period(observations["obs_date"].max(), freq="Q") + snapshot_query = RefSnapshotQuery( + series_key=series_keys[0], + asof=asof, + start_ref=start_ref, + end_ref=end_ref, + ) + panel_specs = [ + { + "series_key": series_key, + "alias": series_key.lower(), + "start_ref": start_ref, + "end_ref": end_ref, + "freq": "Q", + } + for series_key in series_keys + ] + + snapshot_samples = _time_call( + lambda: pit.snapshot_ref(snapshot_query), + iterations=iterations, + ) + panel_samples = _time_call( + lambda: pit.build_snapshot_panel_long( + panel_specs, + asof=asof, + align="quarter_end", + ), + iterations=iterations, + ) + + return { + "timestamp_utc": pd.Timestamp.now(tz="UTC").isoformat(), + "iterations": iterations, + "periods": periods, + "series_count": series_count, + "revisions_per_period": revisions_per_period, + "rows_loaded": int(len(observations)), + "snapshot_ref_median_ms": round(statistics.median(snapshot_samples), 3), + "snapshot_panel_long_median_ms": round(statistics.median(panel_samples), 3), + } + + +def main() -> None: + parser = argparse.ArgumentParser(description="Run PIT contract benchmarks.") + parser.add_argument("--iterations", type=int, default=5) + parser.add_argument("--periods", type=int, default=40) + parser.add_argument("--series-count", type=int, default=3) + parser.add_argument("--revisions-per-period", type=int, default=2) + args = parser.parse_args() + + metrics = run_pit_contract_benchmarks( + iterations=args.iterations, + periods=args.periods, + series_count=args.series_count, + revisions_per_period=args.revisions_per_period, + ) + print(json.dumps(metrics, indent=2, sort_keys=True)) + + +if __name__ == "__main__": + main() diff --git a/doc/plan/core_platform_roadmap.md b/doc/plan/core_platform_roadmap.md new file mode 100644 index 0000000..6c9a9e9 --- /dev/null +++ b/doc/plan/core_platform_roadmap.md @@ -0,0 +1,745 @@ +# Alphaforge Core Platform Roadmap + +**Status:** Draft for review + +## Goal + +Develop Alphaforge into a mathematically grounded platform for messy real-world +data, with a stable public API for: + +- point-in-time data and revision-aware queries +- source access, caching, and archival ingestion +- research dataset assembly for tabular and PIT-driven workflows + +This plan treats Alphaforge as a shared data platform, not as a full +strategy-orchestration framework. + +## Why Now + +The current downstream usage makes the next priorities clear: + +- `nowcast-data` uses Alphaforge as a semantic PIT backend and stresses + ref-period, revision, release, and as-of correctness. +- the volatility notebook workflow uses Alphaforge as a research assembly layer + and stresses `DatasetSpec`, templates, and notebook-friendly ergonomics. +- `positioning` uses Alphaforge as the shared operational data layer and + stresses public-web source robustness, PIT archival, and source health. + +The public-web refactor is now largely complete, so the next roadmap should +consolidate the whole library rather than continuing one subsystem at a time. + +## Design Thesis + +Alphaforge exists because real-world data is messy. The library should impose +order on that mess by making the underlying mathematical structure explicit. + +The central semantic object is not "a dataframe with dates". It is a value over +multiple axes: + +- entity +- measure +- observation time or reference period +- availability time (`asof`) + +For point-in-time data, we can think of the canonical relation as: + +`x(entity, measure, observation, asof) -> value` + +Everything else is derived from this: + +- a snapshot fixes an `asof` cutoff and selects the latest admissible value +- a revision history fixes an observation key and varies `asof` +- a panel aligns multiple snapshots on an explicit evaluation grid +- a derived series is valid only if its transform is causal with respect to + the chosen `asof` + +That viewpoint should guide the public API, the internal constructs, and the +time-series operations. + +## Semantic Laws + +These are the invariants the implementation should preserve wherever possible. + +### 1. Snapshot admissibility + +A snapshot at cutoff `a` may only depend on observations whose availability is +`<= a`. + +### 2. Revision identity + +A revision history holds the observation key fixed and varies only the +availability axis. It is not a different time series; it is a different view of +the same semantic observation. + +### 3. Explicit normalization + +Any mapping between calendar dates and reference periods must be explicit and +frequency-aware. Alphaforge should not rely on silent date coercion where +ref-period semantics matter. + +### 4. Causal transform safety + +A PIT transform is valid only if, for every cutoff `a`, it can be evaluated from +inputs admissible at `a`. Non-causal transforms should be labeled accordingly. + +### 5. Monotone information set + +Moving `asof` forward may change values through new releases or revisions, but +it should never shrink the admissible information set. + +## Mathematical Design Principles + +### 1. Make semantic axes explicit + +Never conflate: + +- observation date +- reference period +- release date +- availability / `asof` +- evaluation grid +- trading calendar session label + +If two concepts have different semantics, they should have different types, +fields, or APIs. + +### 2. Prefer typed semantic objects over loose strings + +Use explicit objects where semantics matter: + +- ref periods +- release rules +- availability policies +- time grids +- missingness classification +- query intent for snapshots, revisions, and aligned panels + +Strings and ad hoc dicts may remain at loader boundaries, but not as the main +semantic API. + +### 3. Treat operations as algebra over typed data + +Time-series operations should compose over clearly defined inputs: + +- snapshot -> aligned series +- aligned series -> causal transform +- set of aligned series -> panel +- panel + policies -> dataset + +Operations should declare their alignment and causality assumptions instead of +smuggling them through incidental pandas behavior. + +### 4. Preserve causality by construction + +For PIT workflows, every transform should either: + +- be causal and safe under `asof`, or +- explicitly declare why it is not causal and where it is intended for + evaluation-only use + +Leakage should be a first-class failure mode, not an afterthought. + +### 5. Separate semantic core from source plumbing + +Source adapters, HTTP clients, caching, and archival concerns are essential, but +they should not define the semantics of PIT operations. The semantic core should +remain usable even when the underlying source families evolve. + +### 6. Keep one canonical public path per layer + +Alphaforge can keep compatibility shims, but it should not keep multiple equally +"official" abstractions indefinitely. + +## Strategic Boundaries + +Alphaforge should own: + +- PIT semantics and transforms +- unified data access and caching contracts +- public-web and local archival source infrastructure +- dataset assembly and feature-template ergonomics +- source health, missingness, lineage, and leakage diagnostics + +Alphaforge should not grow into: + +- a full backtesting framework +- strategy-specific orchestration +- domain-specific modeling logic for every downstream project + +## Platform Pillars + +These are the enduring product pillars for Alphaforge. They are all important. +The implementation order later in this document reflects dependency structure, +not a belief that only the first few pillars matter. + +### 1. Mathematical temporal semantics + +Alphaforge should make time semantics explicit, typed, and defensible: + +- ref periods +- release rules +- availability / `asof` +- calendars and evaluation grids +- missingness and causality + +This is the formal substrate of the library. + +### 2. PIT as the flagship capability + +Revision-aware data is the strongest differentiator in the codebase. Alphaforge +should provide the best surface here: + +- snapshots +- revisions +- aligned PIT panels +- causal transforms +- lineage and explainability + +### 3. Unified source access and ingestion + +The source layer should feel like one system rather than several overlapping +ones: + +- one canonical fetch contract +- one routing story +- one compatibility story +- one archival ingestion story + +### 4. Operational observability and data quality + +Because the library deals with messy real-world data, operational discipline is +part of the product: + +- source health +- release-aware staleness +- missingness classification +- leakage diagnostics +- lineage and provenance + +### 5. Dataset algebra and research UX + +Alphaforge should make research assembly elegant enough that notebooks and +experiments do not need to work around it constantly: + +- `DatasetSpec` +- typed templates +- join and missingness policies +- notebook-friendly workflows +- reusable recipes for common feature families + +### 6. Stability, compatibility, and performance + +Downstream repos should be able to depend on Alphaforge without pinning to +implementation accidents: + +- contract tests +- benchmarks +- migration guides +- deprecation policy +- release discipline + +### 7. Public API ergonomics + +The amount of code a downstream user needs to write to load something is a core +product signal. Alphaforge should not require repo-local helper layers for +routine tasks that ought to be first-class. + +Common tasks should have short, task-shaped entry points for: + +- loading a source table or panel +- loading a PIT snapshot +- loading a PIT revision history +- wiring a default local context +- building a small dataset from a concise spec + +## Target Public API Shape + +## Public API North Star + +The public API should optimize for the common path, not merely expose the +internal architecture cleanly. + +The roadmap should explicitly improve: + +- time-to-first-frame: how many lines it takes to load a standard source table +- time-to-first-snapshot: how many lines it takes to load a PIT snapshot +- setup burden: how many objects a user must manually instantiate +- wrapper pressure: whether downstream repos still need bespoke loading facades + +### Layer 1: Temporal Semantics + +Build out `alphaforge.time` and adjacent PIT semantics around explicit objects: + +- `RefPeriod` and ref-frequency-aware utilities +- `ReleaseRule` +- calendar-aware evaluation grids +- availability and missingness semantics + +This layer is the mathematical vocabulary the rest of the library depends on. + +### Layer 2: PIT Flagship API + +Make the PIT layer the clearest public surface for revision-aware data: + +- first-class snapshot queries +- first-class revision queries +- ref-period-aware snapshot and revision APIs +- batch snapshot / panel builders +- lineage and explainability for derived PIT series + +Keep `PITAccessor` as the operational engine during migration, but move the +public surface toward typed query intent rather than ad hoc method expansion. + +### Layer 3: Unified Data Access + +Make `SourceAdapter` the canonical external data access contract. + +The intended direction is: + +- `SourceAdapter` = public, cache-aware source contract +- `DataContext.fetch(...)` = canonical routing entry point +- `DataSource` = legacy/raw-loader compatibility surface and internal bridge + +This lets Alphaforge keep the current loader ecosystem without making the old +and new access models equally normative forever. + +### Layer 4: Dataset Algebra + +Make `DatasetSpec` the canonical research assembly API. + +This layer should support: + +- reusable feature and target requests +- typed template composition +- explicit join and missingness policies +- notebook-friendly single-entity workflows +- reusable recipes for common research patterns + +### Layer 5: Operations and Observability + +Keep source health, missingness, lineage, and leakage diagnostics as first-class +operational features rather than optional utilities. + +## Downstream-Driven Success Criteria + +The roadmap is successful when: + +- `nowcast-data` no longer needs bespoke semantic wrappers for core ref-period, + revision, and release-aware PIT behavior. +- `positioning` can build PIT panels and release-aware source health flows + without repo-specific panel loops or duplicated health semantics. +- the volatility notebook workflow can express common feature-engineering paths + with fewer manual joins and fewer bespoke notebook helpers. +- common downstream loading tasks require materially less boilerplate than they + do today. +- public-web and archival workflows expose a clearer operational story for + ingestion, health, and provenance rather than just a collection of loaders. +- Alphaforge publishes a clear migration path for compatibility layers and keeps + regressions out with contract tests. + +## Ordered Implementation Roadmap + +The phases below are ordered primarily by dependency and architectural leverage. +They should not be read as "do nothing on later pillars until earlier phases are +fully complete". In practice, the roadmap should run as one foundation track and +several lighter parallel tracks. + +## Execution Model + +### Foundation Track + +This is the dependency-critical path: + +- temporal semantic core +- PIT public API consolidation +- data access contract unification + +### Parallel Product Tracks + +These should advance continuously in smaller slices while the foundation track +lands: + +- dataset algebra and research UX +- operational data platform and observability +- public API ergonomics for common loading tasks + +### Continuous Discipline + +These should start early and continue throughout the roadmap: + +- compatibility suites +- benchmarks +- migration docs +- deprecation and release policy + +### Phase 1: Temporal Semantic Core + +**Objective:** make time and release semantics explicit enough to support the +rest of the roadmap. + +**Scope:** + +- `alphaforge/time/` +- PIT release and missingness semantics +- typed ref-period and availability constructs + +**Deliverables:** + +- promote release rules into Alphaforge core +- add a unified missingness taxonomy and classifier +- standardize ref-period utilities and normalization rules +- document the semantic distinctions between observation time, ref period, + release date, and `asof` + +**Why first:** this is the foundation for both nowcast-style PIT correctness and +health / staleness generalization. + +### Phase 2: PIT Public API Consolidation + +**Objective:** make PIT the flagship user-facing surface. + +**Scope:** + +- `alphaforge/pit/` +- `alphaforge/time/ref_period.py` +- PIT docs and examples + +**Deliverables:** + +- first-class public APIs for snapshot, revision, and ref-period queries +- batch PIT panel builders for wide and long panel assembly +- lineage / explain APIs for derived PIT series +- contract tests covering causal transforms and as-of behavior + +**Why second:** `nowcast-data` is already proving the importance of this layer, +and both `nowcast-data` and `positioning` are paying a performance tax through +manual snapshot loops. + +### Phase 3: Data Access Contract Unification + +**Objective:** reduce ambiguity between `DataSource`, `SourceAdapter`, +`PITDataSource`, and downstream compatibility wrappers. + +**Scope:** + +- `alphaforge/data/context.py` +- `alphaforge/data/adapter.py` +- `alphaforge/data/source.py` +- adapter compatibility shims + +**Deliverables:** + +- publish one canonical fetch path based on `SourceAdapter` +- narrow the role of `DataSource` to raw-loader / compatibility use +- document migration guidance for legacy source registration +- add contract tests for routing, dataset resolution, and cache-aware fetches + +**Why third:** the unified access surface should be clarified before further +growth in dataset assembly or adapter families. + +### Phase 4: Dataset Algebra and Research UX + +**Objective:** make the research assembly layer feel deliberate rather than +partly-built. + +**Scope:** + +- `alphaforge/features/` +- dataset-spec docs and examples +- feature template catalog + +**Deliverables:** + +- tighten the `DatasetSpec` public contract +- expand built-in templates for calendar, event, rolling, and common market + features +- add simpler single-entity and notebook-friendly patterns +- document recipe-style workflows that mirror the volatility notebooks + +**Why fourth:** the volatility workflow is a lighter user of Alphaforge, but it +shows where the library should feel easier and more reusable. + +### Phase 5: Operational Data Platform + +**Objective:** complete the move from "loaders plus helpers" to a coherent +operational data layer. + +**Scope:** + +- source health +- archival helpers +- source metadata and diagnostics +- public-web operational workflows + +**Deliverables:** + +- integrate release-aware health policies where cadence alone is too weak +- provide better PIT archival helpers for recurring source ingestion +- standardize source metadata, health, and diagnostics surfaces +- consolidate docs for source authoring and operational usage + +**Why fifth:** the public-web refactor is already done enough that the next +work should focus on operations and observability instead of more loader-family +cleanup. + +### Phase 6: Stability, Compatibility, and Performance + +**Objective:** make Alphaforge safe to depend on across downstream repos. + +**Scope:** + +- benchmarks +- contract tests +- migration and deprecation policy +- release discipline + +**Deliverables:** + +- benchmark PIT snapshot and panel workloads +- add compatibility suites modeled on current downstream usage +- publish deprecation windows for legacy surfaces +- ensure docs cover migration between the legacy and canonical APIs + +**Why sixth:** performance and stability need to be enforced once the canonical +surfaces are clear enough to lock down. + +## Proposed Implementation Slices + +These slices define the intended implementation order. They are small enough to +become individual Linear tickets later. + +1. Temporal semantics: adopt release rules and missingness taxonomy in + Alphaforge core. +2. Temporal semantics: standardize ref-period normalization and typed helpers. +3. PIT API: add first-class ref-period snapshot and revision queries. +4. PIT performance: add batch snapshot and panel-building primitives. +5. PIT explainability: expose lineage and causal-transform diagnostics. +6. Data access: publish `SourceAdapter` as the canonical fetch contract. +7. Data access: narrow `DataSource` to a compatibility and raw-loader role. +8. Dataset API: tighten `DatasetSpec` semantics and template composition rules. +9. Public API ergonomics: reduce common loading boilerplate and setup friction. +10. Research UX: add notebook-ready templates and recipe documentation. +11. Operations: integrate release-aware health policies and archival helpers. +12. Compatibility: add downstream-inspired contract tests and benchmarks. +13. Documentation: publish migration guides and architecture references. + +## First 90 Days + +The first delivery window should stay foundation-first, but it should still +touch all major pillars in visible ways. + +1. Formalize the semantic vocabulary for ref periods, release rules, + availability, and missingness. +2. Turn the PIT layer into a clearer public API for ref-period snapshots, + revisions, and aligned panel retrieval. +3. Add contract tests and benchmarks modeled on the current `nowcast-data` and + `positioning` call patterns. +4. Publish the canonical direction for `SourceAdapter` versus `DataSource` + before any larger adapter expansion. +5. Land at least one concrete research-UX slice in `DatasetSpec` or templates + so the volatility workflow benefits early rather than waiting for a later + phase. +6. Land at least one concrete operational slice for release-aware health or PIT + archival helpers so the platform story is not purely semantic. +7. Land at least one concrete loading-ergonomics slice that reduces setup code + for a common downstream task. + +## 90-Day Workstreams + +The first 90 days should run as parallel workstreams, not as a single-file +queue. + +### Workstream A: Semantic foundation + +- release rules +- missingness taxonomy +- ref-period normalization + +### Workstream B: PIT flagship API + +- ref-period snapshots +- revision queries +- aligned panel builders + +### Workstream C: Research UX + +- one or two high-value `DatasetSpec` / template improvements +- one notebook-shaped end-to-end example + +### Workstream D: Loading ergonomics + +- define the shortest supported path for common loading tasks +- remove one major source of manual context or adapter boilerplate + +### Workstream E: Operations + +- release-aware health policy integration +- one reusable archival helper improvement + +### Workstream F: Stability + +- downstream-inspired contract tests +- PIT performance benchmarks +- migration notes for canonical versus legacy surfaces + +## Sequencing Constraints + +- Finish the temporal semantic core before expanding PIT APIs that depend on it. +- Land PIT panel primitives before migrating downstream panel builders toward + shared helpers. +- Clarify the canonical data access contract before broadening the adapter + surface further. +- Improve loading ergonomics continuously rather than waiting for all + architectural cleanup to finish. +- Treat dataset UX work as a second-order concern after PIT and access-contract + consolidation. +- Use the public-web refactor as an input to Phase 5, not as the main roadmap. +- Do not remove legacy surfaces until contract tests and migration docs exist. + +## Validation Plan + +This is a planning-only document, so TDD does not apply yet. Implementation +tickets derived from this plan should validate against: + +- targeted subsystem tests +- broader regression coverage for touched layers +- benchmark snapshots for PIT-heavy paths +- docs updates when public behavior changes + +The highest-value contract coverage should model the three existing downstream +usage patterns: + +- nowcast-style PIT semantics +- volatility notebook dataset assembly +- positioning-style operational source + health workflows + +## Linear Ticket Mirror + +These tables mirror the current Linear tickets for the core platform roadmap and +define the intended implementation order for subsequent agents. + +Linear routing for this plan: + +- Team: `ALP` (`alphaforge`) +- Project: `alphaforge` +- Umbrella issue: `ALP-9` +- Migration log: `doc/plan/migration_note.md` + +Rules for coding agents: + +- Linear is the source of truth for ticket state. +- This plan file is the repo-local execution mirror for subsequent agents. +- `doc/plan/migration_note.md` is the running migration log for public-surface + and downstream-impacting changes. +- Implement the earliest ticket in the ordered queue below whose status is not + `Done` and whose hard prerequisites are already satisfied. +- Skip tickets whose table row is `Done`. +- Epic rows are tracking rows. Do not pick them up before their earlier child + rows unless the epic has no remaining open child slices. +- Before coding, print the current ticket number and its plain-English goal on + screen. +- Follow the implementation workflow in `AGENTS.md`: + - review upstream tickets and notes first + - implement with TDD + - update docs after tests pass + - leave the structured closeout note in Linear + - mark the Linear ticket `Done` + - only then update the corresponding row in this file +- Update `doc/plan/migration_note.md` for any ticket that changes the preferred + public API, compatibility story, loading path, PIT semantics, or downstream + migration actions. +- Update this table only after the ticket is closed in Linear. + +Status mirror last synced: `2026-04-05` + +### Ordered Epic Queue + +| Ticket | Status | +| --- | --- | +| `ALP-9` Core platform: implement mathematically grounded public API and migration path | Done | + +### Ordered Core Platform Queue + +| Ticket | Status | +| --- | --- | +| `ALP-10` Temporal semantics: add release rules and missingness core | Done | +| `ALP-11` Temporal semantics: standardize typed ref-period utilities and normalization | Done | +| `ALP-12` PIT API: add first-class ref-period snapshot and revision queries | Done | +| `ALP-14` PIT performance: add batch snapshot and panel-building primitives | Done | +| `ALP-13` PIT explainability: expose lineage and causal diagnostics | Done | +| `ALP-15` Data access: canonicalize `SourceAdapter` fetch path | Done | +| `ALP-16` Data access: reduce `DataSource` to compatibility and raw-loader role | Done | +| `ALP-17` Dataset API: tighten `DatasetSpec` semantics and template composition | Done | +| `ALP-18` Public API ergonomics: reduce common loading boilerplate and setup friction | Done | +| `ALP-19` Research UX: add notebook-ready templates and recipe documentation | Done | +| `ALP-20` Operations: integrate release-aware health and archival helpers | Done | +| `ALP-21` Compatibility: add downstream contract suite and performance benchmarks | Done | +| `ALP-22` Docs: publish architecture and migration guides and maintain migration note | Done | + +### Post-Migration Cleanup Queue + +These tickets should not be picked up until their migration trigger conditions +are satisfied. The detailed cleanup backlog lives in +`doc/plan/post_migration_plan.md`. + +| Ticket | Status | +| --- | --- | +| `ALP-23` Platform cleanup: remove temporal-semantics compatibility shims after migration | Backlog | +| `ALP-24` Platform cleanup: remove SourceAdapterPITCompat bridge after PIT adapter migration | Backlog | +| `ALP-25` Platform cleanup: remove legacy DataContext source access after adapter migration | Backlog | +| `ALP-26` PIT cleanup: remove boolean strict compatibility for PIT ingestion | Backlog | +| `ALP-27` PIT API: remove legacy get_*_ref compatibility helpers after ref-query migration | Backlog | + +### Cross-Ticket Sequencing Constraints + +Use the queue order above, but also respect these concrete handoff rules: + +- Treat `ALP-9` as a tracking umbrella. Start implementation with `ALP-10`. +- Do not start `ALP-11` until `ALP-10` is landed. +- Do not start `ALP-12` or `ALP-14` until `ALP-11` is landed. +- Treat `ALP-12` and `ALP-14` as the first PIT-facing implementation queue; + use `ALP-12` before `ALP-14` so the public query surface lands before the + shared batch helpers. +- Use `ALP-13` after the first PIT public-surface slice is in place; it is an + explainability pass, not the initial PIT contract. +- Use `ALP-15` before `ALP-16` and `ALP-18` so the canonical access path is + clear before compatibility reduction and boilerplate reduction work. +- Use `ALP-17` before `ALP-19` so the dataset contract lands before the + notebook-ready template and recipe pass. +- `ALP-20` depends on the semantic foundation from `ALP-10` but otherwise + should advance alongside the main queue once that dependency is satisfied. +- `ALP-21` should begin after the first meaningful public-surface slices are in + place; it is a continuous discipline ticket, but it should not block the + roadmap from starting. +- `ALP-22` starts immediately and remains active throughout the umbrella. Do + not treat it as a substitute for implementation tickets. + +## Agent Pickup Directive + +If this roadmap is handed to a coding agent for implementation planning: + +1. start at the top of the epic queue +2. within the active queue, pick the first non-`Done` ticket whose hard + prerequisites are satisfied +3. implement only that ticket's scoped outcome +4. update `doc/plan/migration_note.md` whenever the ticket changes downstream + behavior or migration expectations +5. close the ticket in Linear first +6. then update the mirrored status row in this plan before moving on + +## Non-Goals + +- turning Alphaforge into a full backtesting or strategy package +- hiding all source-specific behavior behind a generic DSL +- forcing deep inheritance on unrelated data-source families +- introducing abstract machinery without at least one clear downstream user + +## Immediate Next Step + +The first implementation umbrella should cover Phases 1 and 2 together: + +- formalize the semantic vocabulary for time, release, and availability +- then use that vocabulary to make PIT the clearest public API in the library + +That is the highest-leverage path because it addresses the heaviest downstream +pressure while setting the mathematical foundation for the rest of the platform. diff --git a/doc/plan/migration_note.md b/doc/plan/migration_note.md new file mode 100644 index 0000000..10618f1 --- /dev/null +++ b/doc/plan/migration_note.md @@ -0,0 +1,1006 @@ +# Core Platform Migration Note + +**Status:** Core platform epic complete; downstream migration active + +## Purpose + +This file is the running migration log for the core platform roadmap tracked +under `ALP-9`. + +Use it to record downstream-impacting changes as implementation lands, with a +focus on: + +- public API changes +- compatibility boundary changes +- downstream repo follow-ups +- required migration actions +- temporary shims and planned removals + +This note complements the roadmap in +`doc/plan/core_platform_roadmap.md`. The roadmap explains what the program is +trying to build; this file records what changed and what downstream users need +to know. + +Post-migration cleanup items that should only happen after downstream repos are +fully moved live in `doc/plan/post_migration_plan.md`. + +## Downstream Repos To Watch + +- `nowcast-data` +- `positioning` +- `steveya.github.io/posts/volatility-forecasts-*` + +## Update Protocol + +Add an entry whenever a ticket changes: + +- the preferred public API +- the canonical loading path +- PIT query semantics +- dataset-spec or template behavior +- source health / archival semantics +- compatibility shims or deprecation plans + +Each entry should include: + +- date +- ticket +- summary of the change +- impacted public surface +- downstream repos affected +- migration action required, if any +- temporary compatibility path, if any +- follow-up notes + +## Current Migration Fronts + +These are the main areas where downstream code is expected to change as the +roadmap lands: + +- ref-period and release-aware PIT semantics +- the canonical fetch path around `SourceAdapter` +- the reduced public role of `DataSource` +- lower-boilerplate loading and local context setup +- dataset-spec and template ergonomics +- release-aware health and archival helpers + +## Current Canonical Paths + +Downstream agents should treat these as the current preferred public surfaces. +The historical entries below explain how they got here; this section is the +fast path for current migration work. + +### Temporal semantics + +- import release rules and missingness from `alphaforge.time` +- use `RefPeriod`, `coerce_ref_period(...)`, `normalize_ref_freq(...)`, and + `normalize_obs_date_anchor(...)` for explicit ref-period handling + +### PIT queries and panels + +- use `RefSnapshotQuery` and `RefRevisionQuery` +- use `PITAccessor.snapshot_ref(...)` and `PITAccessor.revisions_ref(...)` +- use `PITAccessor.build_snapshot_panel_long(...)` or + `build_snapshot_panel(...)` for aligned PIT panels +- use `PITAccessor.get_series_lineage(...)` and + `PITAccessor.explain_series(...)` for explainability + +### Source access and loading + +- use `SourceAdapter` as the canonical loading abstraction +- bootstrap with `DataContext.from_adapters(...)` +- load through `ctx.fetch(...)`, `ctx.fetch_many(...)`, `ctx.prefetch(...)`, + and `ctx.load(...)` + +### Dataset assembly and research UX + +- use `DatasetSpec` with explicit `FeatureRequestGroup` composition where + feature families belong together +- use built-in market templates from `alphaforge.features`, especially: + - `LagReturnsTemplate` + - `RollingVolatilityTemplate` + +### Operations and source monitoring + +- use `SourceHealthPolicy` with release-aware rules where needed +- use `build_health_report(...)` or `SourceHealthTracker.report(...)` for + structured health output +- use `discover_archive_fetches(...)` and `iter_yearly_archive_fetches(...)` + for deterministic archive planning + +## Still-Temporary Compatibility Surfaces + +These surfaces are still supported during downstream migration, but they should +not be treated as equal long-term public directions: + +- `alphaforge.pit.release_rules` +- `alphaforge.pit.missingness` +- `PITAccessor.get_snapshot_ref(...)` +- `PITAccessor.get_revision_timeline_ref(...)` +- `alphaforge.pit.adapters.source_adapter_compat.SourceAdapterPITCompat` +- `DataContext.sources` +- `DataContext.fetch_panel(...)` +- boolean `strict=True/False` for `PITAccessor.upsert_pit_observations(...)` +- defaulting to `DataSource` as the primary public loading abstraction + +Removal and cleanup of these bridges is tracked in +`doc/plan/post_migration_plan.md` under `ALP-23` through `ALP-27`. + +## Repo-By-Repo Migration Checklist + +Use these lists as the starting point for downstream migration work. + +### `nowcast-data` + +- move release-rule and missingness imports to `alphaforge.time` +- replace repo-local ref-period coercion with `RefPeriod` / + `coerce_ref_period(...)` +- move ref-period PIT reads onto `RefSnapshotQuery` / + `RefRevisionQuery` plus `snapshot_ref(...)` / `revisions_ref(...)` +- pass `obs_date_anchor` explicitly for period-start keyed series +- replace series-by-series PIT panel loops with shared panel builders +- prefer lineage APIs over direct `meta_json` parsing +- stop adding new dependencies on `get_snapshot_ref(...)` or + `get_revision_timeline_ref(...)` + +### `positioning` + +- move new loading code onto `SourceAdapter` plus adapter-backed `DataContext` +- stop teaching `DataSource`, `ctx.sources[...]`, or `fetch_panel(...)` as the + default public loading path +- use `ctx.fetch_many(...)` for batched loads instead of manual query loops +- adopt release-aware health reports via `build_health_report(...)` or + `tracker.report(...)` +- adopt deterministic archive planning via `discover_archive_fetches(...)` or + `iter_yearly_archive_fetches(...)` +- move temporal imports to `alphaforge.time` and prefer typed PIT helpers where + PIT semantics are involved + +### `steveya.github.io/posts/volatility-forecasts-*` + +- move notebook recipes onto `DataContext.from_adapters(...)` and `ctx.load(...)` +- use `FeatureRequestGroup` for grouped feature families +- import `LagReturnsTemplate` and `RollingVolatilityTemplate` from + `alphaforge.features` +- stop treating helper code under `examples/` as the reusable template API +- rely on `DatasetSpec` validation rather than permissive late failures for join + and missingness policy strings + +## Minimum Validation For Migration-Sensitive Changes + +When migration work changes Alphaforge itself rather than only downstream call +sites: + +- run `python -m pytest tests/contracts` +- rerun `python -m benchmarks.pit` when PIT retrieval performance might move +- update this file whenever the preferred public path, temporary bridge set, or + downstream migration action changes + +## Initial Setup Entry + +### 2026-04-04 + +**Tickets:** `ALP-9` through `ALP-22` + +**Summary:** + +Started the durable execution scaffolding for the core platform roadmap: + +- created the umbrella issue `ALP-9` +- created the full child ticket queue for roadmap implementation +- created this migration note +- updated the roadmap to treat migration guidance as a first-class workstream + +**Impacted public surface:** + +- None yet. This is planning and execution setup only. + +**Downstream repos affected:** + +- None yet. No implementation has landed. + +**Migration action required:** + +- None yet. + +**Temporary compatibility path:** + +- Existing downstream wrappers and compatibility layers remain unchanged. + +**Follow-up notes:** + +- The first migration-sensitive tickets are the temporal semantics, PIT API, + data-access, and loading-ergonomics slices. +- As soon as one of those tickets lands, record the exact downstream behavior + change here rather than only in PR or ticket commentary. + +## Migration Entries + +### 2026-04-04 + +**Ticket:** `ALP-10` + +**Summary:** + +Promoted release-rule and missingness semantics into the core +`alphaforge.time` package and made source health release-aware when a +`release_rule` is available. + +**Impacted public surface:** + +- canonical release-rule imports now live under `alphaforge.time` +- canonical missingness imports now live under `alphaforge.time` +- top-level `alphaforge` now re-exports the temporal semantic core +- `SourceHealthPolicy.release_rule` now changes health evaluation semantics + from cadence-only aging to next-expected-release timing + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` + +**Migration action required:** + +- Prefer `from alphaforge.time import ReleaseRule, MissingnessReason, + classify_missingness, ...` for new code. +- Existing health policies that set `release_rule` should expect status + transitions to be measured against the next expected release window rather + than only `latest_obs_date + expected_cadence`. + +**Temporary compatibility path:** + +- `alphaforge.pit.release_rules` and `alphaforge.pit.missingness` remain + import-compatible shims. + +**Follow-up notes:** + +- `ALP-11` should build on this by tightening typed ref-period handling inside + the same temporal-semantic layer. +- `ALP-20` can now rely on the shared release-aware health vocabulary instead + of inventing separate operational timing rules. +- Shim removal after downstream migration is explicitly tracked in `ALP-23` + and mirrored in `doc/plan/post_migration_plan.md`. + +### 2026-04-04 + +**Ticket:** `ALP-11` + +**Summary:** + +Standardized ref-period normalization around the typed `alphaforge.time` +surface so PIT and target helpers share one explicit path for canonical ref +keys, pandas `Period` inputs, and explicit observation dates plus declared +frequency/anchor semantics. + +**Impacted public surface:** + +- `RefPeriod.parse(...)` now accepts typed normalization inputs beyond plain + ref-key strings +- `alphaforge.time.ref_period` now exposes explicit normalization helpers such + as `coerce_ref_period`, `normalize_ref_freq`, and `normalize_obs_date_anchor` +- PIT ref helpers now accept pandas `Period` inputs through the shared + normalization path + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` + +**Migration action required:** + +- Prefer `RefPeriod` and `coerce_ref_period(...)` instead of repo-local quarter + parsing or timestamp-to-ref conversions. +- When normalizing an observation date into a reference period, pass the + intended frequency and anchor explicitly instead of relying on incidental date + coercion. + +**Temporary compatibility path:** + +- Existing string ref keys such as `2024Q4` remain supported. +- No new legacy shim module was introduced in this ticket. + +**Follow-up notes:** + +- `ALP-12` can now build ref-period snapshot and revision APIs on top of the + shared typed normalization surface instead of duplicating quarter/date logic. + +### 2026-04-04 + +**Ticket:** `ALP-12` + +**Summary:** + +Added a first-class ref-period PIT query surface based on typed query objects +instead of asking downstream code to reach into legacy-style accessor helpers. + +**Impacted public surface:** + +- new canonical query objects: + - `RefSnapshotQuery` + - `RefRevisionQuery` +- new canonical execution methods: + - `PITAccessor.snapshot_ref(...)` + - `PITAccessor.revisions_ref(...)` +- `snapshot_ref(...)` now returns a `Series` indexed by typed `RefPeriod` + values rather than raw observation timestamps +- explicit `obs_date_anchor` handling is now available on the canonical + ref-query surface for period-start or period-end keyed series + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` + +**Migration action required:** + +- Prefer `snapshot_ref(...)` and `revisions_ref(...)` with typed query objects + for new code. +- Prefer the RefPeriod-indexed snapshot output instead of repo-local wrappers + that re-key observation timestamps back into reference periods. +- When a series stores ref observations at period start instead of period end, + pass `obs_date_anchor="start"` explicitly in the query object. + +**Temporary compatibility path:** + +- `get_snapshot_ref(...)` and `get_revision_timeline_ref(...)` remain available + as compatibility helpers during migration. + +**Follow-up notes:** + +- Post-migration cleanup for the legacy helper names is tracked in `ALP-27` + and mirrored in `doc/plan/post_migration_plan.md`. +- `ALP-14` can build batch ref-period panel helpers on top of these typed query + semantics instead of inventing another batch-only contract. + +### 2026-04-04 + +**Ticket:** `ALP-14` + +**Summary:** + +Upgraded PIT batch retrieval and panel-building so common snapshot/panel +workloads can use a shared batch path with preserved source-vintage metadata +instead of series-by-series loops. + +**Impacted public surface:** + +- `get_snapshot_multi(...)` now returns `source_asof_utc` +- new canonical long helper: + - `PITAccessor.build_snapshot_panel_long(...)` +- `SnapshotSeriesSpec` now accepts `freq` and `obs_date_anchor` for explicit + ref-aware panel bounds +- `build_snapshot_panel(...)` now builds on the aligned long-form primitive + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` + +**Migration action required:** + +- Prefer `build_snapshot_panel(...)` or `build_snapshot_panel_long(...)` over + repo-local loops that call `get_snapshot(...)` series-by-series. +- When downstream logic needs to retain the supplying vintage per panel row, + consume `source_asof_utc` from `get_snapshot_multi(...)` or the long panel + builder instead of reconstructing it manually. +- For period-start keyed ref series used inside panels, pass `freq` and + `obs_date_anchor` explicitly in `SnapshotSeriesSpec`. + +**Temporary compatibility path:** + +- Existing single-series snapshot methods remain supported. +- No new compatibility-only shim module was introduced in this ticket. + +**Follow-up notes:** + +- `ALP-13` can build lineage and causal diagnostics on top of the preserved + long-form source metadata that now comes out of the shared panel builder. + +### 2026-04-04 + +**Ticket:** `ALP-13` + +**Summary:** + +Added public series-level lineage and causality inspection APIs so derived PIT +outputs can be explained directly from persisted storage instead of forcing +downstream code to parse `meta_json` by hand. + +**Impacted public surface:** + +- new lineage inspection API: + - `PITAccessor.get_series_lineage(...)` +- new series summary API: + - `PITAccessor.explain_series(...)` +- derived series now expose a normalized `causality_status` vocabulary in the + public surface: + - `raw` + - `ok` + - `unknown` + - `violation` + - `experimental` + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` + +**Migration action required:** + +- Prefer `get_series_lineage(...)` / `explain_series(...)` for persisted PIT + provenance inspection instead of repo-local JSON parsing against `meta_json`. +- Prefer the public `causality_status` summary over ad hoc checks of + `source_asof_utc` fields when inspecting whether a derived series is safe + under the stored `asof_utc` semantics. + +**Temporary compatibility path:** + +- Existing lower-level access to raw `meta_json` remains available through the + PIT table itself. +- No new compatibility shim or temporary bridge was introduced in this ticket. + +**Follow-up notes:** + +- `ALP-22` should carry these explainability APIs into the architecture and + migration guides so downstream repos adopt the public inspection path rather + than continuing to scrape raw lineage payloads. + +### 2026-04-04 + +**Ticket:** `ALP-15` + +**Summary:** + +Canonicalized adapter-based source loading around `DataContext.fetch(...)`, +`fetch_many(...)`, and `prefetch(...)`, and tightened batch routing so +`fetch_many(...)` now delegates through the resolved adapter batch contract +instead of looping one query at a time. + +**Impacted public surface:** + +- canonical loading path is now explicitly: + - `SourceAdapter` + - `DataContext.fetch(...)` + - `DataContext.fetch_many(...)` + - `DataContext.prefetch(...)` +- `DataContext.fetch_many(...)` now batches by resolved adapter and preserves + input order while forwarding `max_staleness` +- docs now treat `DataContext.sources` and `fetch_panel(...)` as + backward-compatibility surfaces rather than co-equal public directions + +**Downstream repos affected:** + +- `positioning` +- `nowcast-data` +- `steveya.github.io/posts/volatility-forecasts-*` + +**Migration action required:** + +- Prefer adapter registration plus `ctx.fetch(...)` / `ctx.fetch_many(...)` + for all new loading code. +- Stop teaching `ctx.sources[...]` or `ctx.fetch_panel(...)` as the default + public loading path in downstream helper layers and notebooks. +- Where downstream code batches multiple queries manually, prefer + `ctx.fetch_many(...)` so adapter-level cache and batch optimizations stay + available. + +**Temporary compatibility path:** + +- `DataContext.sources` remains available for legacy `DataSource` users. +- `DataContext.fetch_panel(...)` remains available for legacy panel-oriented + flows. + +**Follow-up notes:** + +- `ALP-16` should narrow the documented `DataSource` role even further now + that the canonical fetch route is explicit and batch-capable. +- Final removal of the legacy context access path remains tracked in `ALP-25` + and mirrored in `doc/plan/post_migration_plan.md`. + +### 2026-04-04 + +**Ticket:** `ALP-16` + +**Summary:** + +Reduced `DataSource` to an explicitly documented compatibility/raw-loader role +instead of leaving it as an implied co-equal public loading contract alongside +`SourceAdapter`. + +**Impacted public surface:** + +- `alphaforge.data.source.DataSource` is now explicitly documented as a + compatibility/raw-loader protocol +- `FREDDataSource` is now documented as a legacy loader, with + `FREDSourceAdapter` as the preferred new-code path +- `PITDataSource` and the public-web loader docs now explicitly position + `DataSource` usage as a bridge/raw-loader surface rather than the canonical + general loading direction + +**Downstream repos affected:** + +- `positioning` +- `nowcast-data` +- `steveya.github.io/posts/volatility-forecasts-*` + +**Migration action required:** + +- Stop treating `DataSource` as the default public abstraction in downstream + docs, helper layers, or new modules. +- Prefer `SourceAdapter` plus `DataContext.fetch(...)` for new loading code, + even when older raw-loader integrations remain in the same repo. +- Keep direct `DataSource` usage only where a loader family is still + intentionally raw-loader based or where an older panel/PIT integration has + not migrated yet. + +**Temporary compatibility path:** + +- `DataSource` remains supported for public-web raw loaders, `PITDataSource`, + and other legacy panel-oriented integrations. +- `DataContext.sources` and `fetch_panel(...)` remain available during the + migration window. + +**Follow-up notes:** + +- `ALP-18` should build the short happy path on top of the now-clearly + canonical adapter surface instead of papering over `DataSource` ambiguity. +- Eventual retirement of the remaining legacy context path is still tracked in + `ALP-25`. + +### 2026-04-04 + +**Ticket:** `ALP-17` + +**Summary:** + +Tightened the `DatasetSpec` contract around explicit feature-request +composition, early policy validation, and observable request metadata in the +feature catalog. + +**Impacted public surface:** + +- new typed composition surface: + - `FeatureRequestGroup` +- `DatasetSpec.feature_requests()` now exposes the flattened request list used + by the builder +- feature catalogs now include request/template metadata such as: + - `request_key` + - `template_name` + - `template_version` +- `JoinPolicy` now validates `how` eagerly +- `MissingnessPolicy` now validates `final_row_policy` eagerly + +**Downstream repos affected:** + +- `steveya.github.io/posts/volatility-forecasts-*` +- `positioning` + +**Migration action required:** + +- Prefer `FeatureRequestGroup` when a notebook or research recipe has a family + of related feature requests with shared tags, key prefixes, or slice + overrides. +- Stop relying on invalid join or missingness strings making it deep into the + builder; invalid values now fail at spec construction time. +- Where downstream code tracks feature families manually, prefer the stamped + `request_key` and merged `tags_json` catalog fields. + +**Temporary compatibility path:** + +- Existing flat `features=[FeatureRequest(...), ...]` specs remain supported. +- No compatibility shim was introduced for this ticket. + +**Follow-up notes:** + +- `ALP-19` can now build notebook-ready recipes on top of `FeatureRequestGroup` + instead of inventing another grouping abstraction. +- `ALP-22` should reflect the request-group composition model in the + architecture and migration guides. + +### 2026-04-04 + +**Ticket:** `ALP-18` + +**Summary:** + +Reduced happy-path loading boilerplate by adding adapter/bootstrap helpers for +common source-table and PIT read workflows. + +**Impacted public surface:** + +- new `DataContext.from_adapters(...)` classmethod for adapter-first bootstrap +- new `DataContext.load(...)` helper for common single-table loads without + explicit `Query(...)` construction +- new `PITAccessor.open(path)` bootstrap for local DuckDB-backed PIT stores + +**Downstream repos affected:** + +- `positioning` +- `nowcast-data` +- `steveya.github.io/posts/volatility-forecasts-*` + +**Migration action required:** + +- Prefer `DataContext.from_adapters(...)` over manual `sources={}`, + `adapters={...}`, and default-source boilerplate when wiring adapter-only + contexts. +- Prefer `ctx.load(...)` for straightforward source-table reads where a custom + `Query` object is not needed. +- Prefer `PITAccessor.open(path)` over manually instantiating + `DuckDBParquetStore` and then passing `store.conn()` into `PITAccessor`. + +**Temporary compatibility path:** + +- Existing `ctx.fetch(Query(...))` calls remain supported. +- Existing `PITAccessor(DuckDBParquetStore(...).conn())` construction remains + supported. + +**Follow-up notes:** + +- `ALP-21` should include contract tests around `from_adapters(...)`, + `load(...)`, and `PITAccessor.open(...)` so future migrations do not regress + the new short path. +- `ALP-22` should keep the happy-path examples synchronized with the ergonomic + helper surface. + +### 2026-04-05 + +**Ticket:** `ALP-19` + +**Summary:** + +Added canonical built-in market templates and recipe docs so volatility-style +research can use reusable lag-return and trailing-volatility feature families +instead of repo-local notebook helper cells. + +**Impacted public surface:** + +- new built-in market templates now live under `alphaforge.features`: + - `LagReturnsTemplate` + - `RollingVolatilityTemplate` +- the preferred market-research recipe path now uses: + - `DataContext.from_adapters(...)` + - `ctx.load(...)` + - `FeatureRequestGroup` + - built-in market templates rather than example-only helper modules +- quickstart and recipe docs now teach the adapter-backed volatility workflow + directly + +**Downstream repos affected:** + +- `steveya.github.io/posts/volatility-forecasts-*` + +**Migration action required:** + +- Prefer `from alphaforge.features import LagReturnsTemplate, + RollingVolatilityTemplate` for new research notebooks and template libraries. +- Prefer custom targets that build on `ctx.load(...)` rather than + `ctx.fetch_panel(...)` when following the canonical adapter-backed dataset + path. + +**Temporary compatibility path:** + +- Existing helper modules under `examples/` still exist for older demos, but + they are no longer the preferred reusable import path for market features. + +**Follow-up notes:** + +- `ALP-21` should add contract coverage around the notebook-style volatility + recipe so the new built-in template family remains stable during later + compatibility cleanup. +- `ALP-22` should incorporate the new research recipe guidance into the + architecture/migration closeout. + +### 2026-04-05 + +**Ticket:** `ALP-20` + +**Summary:** + +Added a more operational release-aware health surface and upgraded archive +ingestion helpers from raw URL lists to deterministic fetch-plan entries. + +**Impacted public surface:** + +- `SourceHealthStatus` now carries structured release-aware diagnostics: + - `overdue` + - `overdue_days` +- new operational report helper: + - `alphaforge.pipeline.health.build_health_report(...)` +- `SourceHealthTracker.report(...)` now returns the same dataframe-style health + report and tracker persistence records `overdue_days` +- public-web archive helpers now expose deterministic fetch planning via: + - `ArchiveFetchPlanEntry` + - `discover_archive_fetches(...)` + - `iter_yearly_archive_fetches(...)` +- archive-backed public-web loaders now use planned artifact names instead of + open-coded URL-to-filename logic + +**Downstream repos affected:** + +- `positioning` + +**Migration action required:** + +- Prefer structured `overdue_days` and `build_health_report(...)` / + `tracker.report(...)` over parsing release-aware delay information out of + free-form status messages. +- Prefer `discover_archive_fetches(...)` or `iter_yearly_archive_fetches(...)` + when building recurring archive ingestion flows so artifact names and year + metadata stay deterministic. + +**Temporary compatibility path:** + +- Existing low-level URL helpers (`discover_archive_links(...)`, + `filter_urls_for_years(...)`, `iter_yearly_archive_urls(...)`) remain + available for callers that intentionally want raw URL primitives. + +**Follow-up notes:** + +- `ALP-21` should lock down the release-aware report shape and archive fetch + planning semantics as part of the downstream-inspired contract suite. + +### 2026-04-05 + +**Ticket:** `ALP-21` + +**Summary:** + +Added an explicit downstream contract suite plus a PIT benchmark harness so +future public-surface changes can be validated against the workflows that +already depend on Alphaforge. + +**Impacted public surface:** + +- new contract-test slice: + - `tests/contracts/test_nowcast_pit_contract.py` + - `tests/contracts/test_volatility_dataset_contract.py` + - `tests/contracts/test_operations_contract.py` + - `tests/contracts/test_pit_benchmarks.py` +- new benchmark harness: + - `python -m benchmarks.pit` + - `benchmarks.run_pit_contract_benchmarks(...)` +- the PIT contract docs now treat the downstream contract slice and benchmark + harness as migration gates for compatibility-sensitive work + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` +- `steveya.github.io/posts/volatility-forecasts-*` + +**Migration action required:** + +- Run `python -m pytest tests/contracts` when changing the canonical PIT, + adapter/data-context, dataset-spec/template, or operational source surfaces. +- Rerun `python -m benchmarks.pit` when changing PIT retrieval paths so + benchmark snapshots stay current in the migration discussion. +- Use the contract files as explicit evidence when removing or narrowing a + compatibility surface. + +**Temporary compatibility path:** + +- No new compatibility shim was introduced in this ticket. +- The contract suite exists to police previously introduced compatibility + bridges rather than to add another one. + +**Follow-up notes:** + +- `ALP-22` should fold these regression gates into the architecture and + migration guides so downstream maintainers can see one coherent story. +- Current PIT benchmark baseline captured at + `2026-04-05T03:35:01.847144+00:00` with + `python -m benchmarks.pit --iterations 5 --periods 40 --series-count 3 --revisions-per-period 2`: + `snapshot_ref(...)` median 5.724 ms, + `build_snapshot_panel_long(...)` median 11.518 ms. + +### 2026-04-05 + +**Ticket:** `ALP-22` + +**Summary:** + +Published explicit core-platform architecture and migration guides and kept the +repo-local roadmap and migration logs synchronized through the end of the epic. + +**Impacted public surface:** + +- new architecture guide for canonical layer boundaries +- new migration guide for downstream moves onto canonical public paths +- new contracts-and-benchmarks guide for the stability discipline behind + migration and deprecation work +- development docs now call out the roadmap-specific regression gates + +**Downstream repos affected:** + +- `nowcast-data` +- `positioning` +- `steveya.github.io/posts/volatility-forecasts-*` + +**Migration action required:** + +- Use the architecture guide as the canonical map of layer ownership and public + direction. +- Use the migration guide when moving code off legacy PIT helpers, legacy + `DataContext` access, and temporary compatibility imports. +- Keep future downstream-impacting tickets updating this note rather than + relying on scattered PR context. + +**Temporary compatibility path:** + +- No new compatibility shim or temporary bridge was introduced in this ticket. +- Existing migration bridges remain tracked in + `doc/plan/post_migration_plan.md`. + +**Follow-up notes:** + +- With the core platform epic landed, future work should treat these guides as + the canonical documentation entry points for architecture and migration. + +## Maintenance Queue + +### 2026-04-05 + +| Ticket | Title | Status | +| --- | --- | --- | +| `ALP-28` | `Data access: fail on ambiguous adapter routing without a default source` | Done | +| `ALP-29` | `Temporal semantics: honor WeeklyRelease weekday configuration` | Done | +| `ALP-30` | `CFTC adapter: preserve disaggregated PIT source lineage` | Done | +| `ALP-31` | `Public web: surface CFTC archive fetch failures instead of silently dropping years` | Done | +| `ALP-32` | `Canonical loading: migrate public-web outliers onto adapter-backed access` | Open | +| `ALP-33` | `DTCC adapters: add shared product-family adapter base over DTCC PPD raw loader` | Done | +| `ALP-34` | `DTCC adapters: add first product-family adapters and dataset contracts` | Done | +| `ALP-35` | `MOF JGB: add adapter-backed canonical load path for constant-maturity yields` | Open | +| `ALP-36` | `Philadelphia SPF: add adapter-backed canonical load path for mean-level surveys` | Open | +| `ALP-37` | `Research UX: add canonical wide curve and maturity-order helpers for adapter-backed loads` | Open | +| `ALP-38` | `Migration: move short_rates and notebook examples off raw-loader outlier APIs` | Blocked by `ALP-35`, `ALP-36`, `ALP-37` | + +### 2026-04-05 + +**Tickets:** `ALP-28`, `ALP-29`, `ALP-30`, `ALP-31` + +**Summary:** + +Hardened four migration-sensitive public surfaces: + +- canonical adapter routing now fails on ambiguous shared datasets unless a + default or explicit source is provided +- `WeeklyRelease` now honors its configured weekday instead of behaving like a + raw lag-only offset +- the multi-dataset CFTC adapter now preserves distinct PIT lineage for + disaggregated CoT rows +- CFTC archive-backed public-web loads now fail fast on broken requested + archives instead of silently returning partial history + +**Impacted public surface:** + +- `DataContext.from_adapters(...)`, `ctx.fetch(...)`, and `ctx.load(...)` for + datasets served by more than one adapter +- `WeeklyRelease.expected_release_date(...)` and release-aware health + semantics that consume weekly rules +- CFTC PIT row provenance for `cot.disagg` +- CFTC public-web archive fetch behavior and operational error visibility + +**Downstream repos affected:** + +- `positioning` +- `nowcast-data` + +**Migration action required:** + +- downstream callers that relied on adapter registration order for shared + datasets must now configure `default_sources` or pass `source=` +- weekly release schedules should expect the configured weekday to matter after + the lag anchor +- workflows that previously tolerated silently partial CFTC archive history + should now expect an explicit failure and handle or fix the broken archive + instead + +**Temporary compatibility path:** + +- No new compatibility shim was introduced. +- These tickets narrow ambiguous behavior on canonical paths rather than + creating another bridge. + +**Follow-up notes:** + +- Existing cached `cot.disagg` PIT rows written before `ALP-30` keep their old + lineage until repopulated. +- The next public-web maintenance pass should decide whether other + archive-backed loaders should adopt the same fail-fast behavior as `ALP-31`. + +### 2026-04-05 + +**Ticket:** `ALP-33` + +**Summary:** + +Introduced a shared DTCC adapter base above `DTCCPPDSource` so canonical DTCC +adapters can own raw-loader wiring internally instead of forcing callers to +inject ad hoc raw-fetch closures. + +**Impacted public surface:** + +- new shared DTCC adapter base: + - `DTCCPPDAdapterBase` +- `DTCCAdapter` now constructs and owns `DTCCPPDSource` internally by default +- future DTCC product-family adapters can vary dataset contracts and PIT + transforms without duplicating cache, prefetch, and raw-query plumbing + +**Downstream repos affected:** + +- `positioning` +- `steveya.github.io` + +**Migration action required:** + +- Prefer `DTCCAdapter(...)` with raw-source keyword arguments instead of + building raw-fetcher closures around `DTCCPPDSource`. +- When adding DTCC canonical datasets such as FX options or IRS, subclass + `DTCCPPDAdapterBase` rather than copying `DTCCAdapter` cache/fetch logic. +- Keep direct `DTCCPPDSource` usage only on the raw-loader compatibility side + of the API boundary. + +**Temporary compatibility path:** + +- Tests and advanced callers may still inject a preconfigured `source=` + object into `DTCCAdapter`. +- Direct `DTCCPPDSource` usage remains available through the public-web + raw-loader surface. + +**Follow-up notes:** + +- `ALP-34` should define the first concrete DTCC product-family adapters and + their dataset contracts on top of `DTCCPPDAdapterBase`. +- The remaining outlier canonicalization work for MOF, SPF, and notebook-style + wide helpers remains tracked in `ALP-35` through `ALP-38`. + +### 2026-04-05 + +**Ticket:** `ALP-34` + +**Summary:** + +Added the first concrete DTCC product-family adapters so the preferred DTCC +adapter surface is no longer one generic `dtcc.ppd` dataset. + +**Impacted public surface:** + +- new canonical DTCC family adapters: + - `DTCCFXAdapter` with dataset `dtcc.fx` + - `DTCCIRSAdapter` with dataset `dtcc.irs` +- `dtcc_daily_to_pit_observations(...)` now accepts custom `key_prefix` and + `source_name` so family adapters can stamp distinct series keys and PIT + lineage without copying the transform +- `DTCCAdapter` remains available as the broader generic wrapper, but canonical + docs now teach `dtcc.fx` and `dtcc.irs` first + +**Downstream repos affected:** + +- `positioning` +- `steveya.github.io` + +**Migration action required:** + +- Prefer `DataContext.from_adapters(DTCCFXAdapter(...), DTCCIRSAdapter(...))` + plus `ctx.load("dtcc.fx", ...)` / `ctx.load("dtcc.irs", ...)` for new DTCC + research or ingestion code. +- Stop introducing new downstream dependencies on the generic `dtcc.ppd` + dataset when the workflow is specifically FX or IRS. +- Keep direct `DTCCPPDSource` usage only where the raw-loader compatibility + surface is explicitly intended. + +**Temporary compatibility path:** + +- `DTCCAdapter` still serves the broader `dtcc.ppd` dataset. +- `DTCCPPDSource` remains available as the low-level raw-loader surface. + +**Follow-up notes:** + +- `DTCCFXAdapter` currently covers FX forwards and swaps; it is not yet an + FX-options-specific contract because current provider fixtures do not expose + that family cleanly. +- `DTCCIRSAdapter` intentionally excludes OIS and cross-currency swaps even + though they share the same low-level IR artifact family. +- Future DTCC slices can add more family adapters on top of the filtered-family + pattern without reopening the raw-loader ownership work from `ALP-33`. diff --git a/doc/plan/post_migration_plan.md b/doc/plan/post_migration_plan.md new file mode 100644 index 0000000..972f10e --- /dev/null +++ b/doc/plan/post_migration_plan.md @@ -0,0 +1,265 @@ +# Post-Migration Plan + +**Status:** Active backlog + +## Purpose + +Track cleanup work that should happen only after downstream migrations are +complete, so temporary compatibility paths do not become permanent by default. + +This file is intentionally small and explicit. Add entries here when a roadmap +slice introduces a temporary shim, deprecation bridge, or compatibility-only +surface that should be removed once migration is done. + +## Update Notes + +### 2026-04-04 + +- `ALP-11` added no new post-migration cleanup artifact. The ref-period work + extended the canonical `alphaforge.time` surface and reused existing string + compatibility without introducing a new shim or temporary bridge. +- `ALP-12` introduced canonical typed ref-query APIs and kept + `get_snapshot_ref(...)` / `get_revision_timeline_ref(...)` as temporary + compatibility helpers. Their cleanup is tracked in `ALP-27`. +- `ALP-14` added no new post-migration cleanup artifact. It widened the + canonical PIT batch/panel surface and upgraded existing snapshot metadata, + but it did not introduce a new temporary bridge or shim. +- `ALP-13` added no new post-migration cleanup artifact. It exposed public + lineage and causality inspection APIs over existing persisted metadata + without adding a temporary compatibility layer. +- `ALP-15` added no new post-migration cleanup artifact. It made the + `SourceAdapter` plus `DataContext.fetch(...)` / `fetch_many(...)` route the + canonical public path and strengthened the migration rationale for `ALP-25`, + but it did not introduce an additional shim or bridge. +- `ALP-16` added no new post-migration cleanup artifact. It narrowed + `DataSource` to an explicitly documented compatibility/raw-loader role and + further clarified why the remaining legacy context surface should eventually + be retired under `ALP-25`, but it did not introduce another temporary shim. +- `ALP-17` added no new post-migration cleanup artifact. It tightened the + canonical `DatasetSpec` surface and introduced `FeatureRequestGroup` as part + of that public contract, not as a temporary compatibility bridge. +- `ALP-18` added no new post-migration cleanup artifact. It introduced + ergonomic bootstrap helpers (`DataContext.from_adapters(...)`, + `DataContext.load(...)`, `PITAccessor.open(...)`) as part of the intended + long-term public surface rather than as temporary shims. +- `ALP-19` added no new post-migration cleanup artifact. It promoted + notebook-ready market templates into the canonical `alphaforge.features` + surface and updated docs/examples toward that path, but it did not introduce + a temporary bridge or shim that should later be removed. +- `ALP-20` added no new post-migration cleanup artifact. The new health-report + helpers and archive fetch-plan objects are intended long-term operational + surfaces, while the older raw URL helpers remain useful lower-level + primitives rather than temporary migration shims. +- `ALP-21` added no new post-migration cleanup artifact. It introduced + contract tests and a PIT benchmark harness to police existing compatibility + bridges, but it did not create a new temporary surface that should later be + removed. +- `ALP-22` added no new post-migration cleanup artifact. It published the + architecture and migration guides that explain the canonical versus legacy + split, but it did not introduce another compatibility bridge. + +## Active Post-Migration Queue + +### 2026-04-04 + +**Ticket:** `ALP-23` + +**Title:** + +`Platform cleanup: remove temporal-semantics compatibility shims after migration` + +**Trigger:** + +- `nowcast-data` +- `positioning` +- other supported downstream consumers + +must no longer import: + +- `alphaforge.pit.release_rules` +- `alphaforge.pit.missingness` + +**Cleanup to perform:** + +- remove the compatibility shim modules at: + - `alphaforge/pit/release_rules.py` + - `alphaforge/pit/missingness.py` +- remove any remaining internal imports that still rely on those legacy PIT + paths +- update migration docs and release notes to reflect shim removal +- keep `alphaforge.time` as the only supported public path for release rules + and missingness + +**Why it exists:** + +`ALP-10` deliberately kept these shims to avoid breaking downstream consumers +during migration. They add short-term migration value, but they also preserve +API ambiguity. They should be removed once the downstream migration window is +closed. + +**Validation at cleanup time:** + +- targeted import-path regression tests +- relevant downstream compatibility checks +- `ruff check .` +- relevant `pytest` slices + +### 2026-04-04 + +**Ticket:** `ALP-24` + +**Title:** + +`Platform cleanup: remove SourceAdapterPITCompat bridge after PIT adapter migration` + +**Trigger:** + +- `nowcast-data` +- any other supported PIT consumer + +must no longer require the legacy `PITAdapter` interface or the bridge at: + +- `alphaforge.pit.adapters.source_adapter_compat.SourceAdapterPITCompat` + +**Cleanup to perform:** + +- remove `alphaforge/pit/adapters/source_adapter_compat.py` +- remove API docs that present the bridge as a supported public path +- migrate any remaining internal callers to the canonical adapter-based path +- remove bridge-specific compatibility tests and fixtures + +**Why it exists:** + +The bridge is an intentional migration aid from legacy `PITAdapter` consumers +to the unified `SourceAdapter` layer. It is useful during migration, but it +keeps a parallel adapter model alive and should not remain indefinitely. + +**Validation at cleanup time:** + +- targeted PIT adapter migration checks +- relevant downstream compatibility checks +- `ruff check .` +- relevant `pytest` slices + +### 2026-04-04 + +**Ticket:** `ALP-25` + +**Title:** + +`Platform cleanup: remove legacy DataContext source access after adapter migration` + +**Trigger:** + +- public docs and supported downstream consumers + +must no longer rely on: + +- `DataContext.sources` +- `DataContext.fetch_panel(...)` +- `ctx.sources[...]` as a user-facing loading pattern + +**Cleanup to perform:** + +- remove or retire the legacy `sources` mapping from the supported public path +- remove or retire `fetch_panel(...)` as a supported legacy access path +- update docs, examples, and tests that still teach or rely on the legacy path +- keep any remaining raw-loader internals only if they are explicitly treated + as non-public + +**Why it exists:** + +The canonical data-access direction is moving toward `SourceAdapter` plus +`DataContext.fetch(...)`, but the legacy context surface still exists for +backward compatibility. Leaving both paths alive indefinitely keeps the public +API ambiguous. + +**Validation at cleanup time:** + +- targeted data-context routing tests +- relevant downstream compatibility checks +- `ruff check .` +- relevant `pytest` slices + +### 2026-04-04 + +**Ticket:** `ALP-26` + +**Title:** + +`PIT cleanup: remove boolean strict compatibility for PIT ingestion` + +**Trigger:** + +- supported callers +- docs and examples + +must no longer rely on: + +- `strict=True` +- `strict=False` + +for `PITAccessor.upsert_pit_observations(...)` + +**Cleanup to perform:** + +- remove boolean handling from PIT ingestion policy resolution +- require explicit string policies only: + - `"error"` + - `"warn"` + - `"coerce"` +- update internal callers, docs, tests, and migration notes to use explicit + strings only + +**Why it exists:** + +Boolean `strict` support is a backward-compatibility overload. The explicit +string policy is the clearer contract, and keeping both forms permanently adds +needless API ambiguity. + +**Validation at cleanup time:** + +- targeted PIT ingestion validation tests +- relevant downstream compatibility checks +- `ruff check .` +- relevant `pytest` slices + +### 2026-04-04 + +**Ticket:** `ALP-27` + +**Title:** + +`PIT API: remove legacy get_*_ref compatibility helpers after ref-query migration` + +**Trigger:** + +- `nowcast-data` +- `positioning` +- other supported PIT consumers + +must no longer rely on: + +- `PITAccessor.get_snapshot_ref(...)` +- `PITAccessor.get_revision_timeline_ref(...)` + +**Cleanup to perform:** + +- remove or retire the legacy helper names from `alphaforge/pit/accessor.py` +- update PIT docs and examples to point only at: + - `PITAccessor.snapshot_ref(...)` + - `PITAccessor.revisions_ref(...)` +- remove compatibility-only tests that keep the old helper names alive + +**Why it exists:** + +`ALP-12` introduced the canonical typed ref-query surface but intentionally +left the older helper names in place to avoid an immediate downstream break. +Those helpers should not remain a parallel public API indefinitely. + +**Validation at cleanup time:** + +- targeted ref-query regression tests +- relevant downstream compatibility checks +- `ruff check .` +- relevant `pytest` slices diff --git a/doc/plan/public_web_refactor_plan.md b/doc/plan/public_web_refactor_plan.md new file mode 100644 index 0000000..489727b --- /dev/null +++ b/doc/plan/public_web_refactor_plan.md @@ -0,0 +1,611 @@ +# `alphaforge.data.public_web` Refactor Plan + +## Goal + +Refactor `alphaforge.data.public_web` to reduce repeated source boilerplate, make new loaders cheaper to add, and centralize the patterns that are already shared in practice without forcing unrelated loaders into an artificial hierarchy. + +This plan is intentionally incremental. The module contains several loader families, but it also contains true outliers. The refactor should extract stable abstractions around the shared fetch pipeline, not attempt to make every source look identical. + +## Current State + +Across the module, most loaders repeat the same high-level flow: + +1. Build or accept a `CachedHttpClient`. +2. Validate `q.table`. +3. Download one or more artifacts. +4. Parse bytes into `DataFrame` objects. +5. Normalize dates, entity ids, and values. +6. Build an output frame with `date` / `entity_id` / `asof_utc`. +7. Apply `apply_query_filters(...)`. +8. Apply `project_columns(...)`. +9. Return either an empty schema-compatible frame or a sorted normalized frame. + +That pattern is visible in many current loaders, including: + +- `bea.py` +- `eia.py` +- `bcb_sgs.py` +- `eurostat.py` +- `destatis_genesis.py` +- `ibge_sidra.py` +- `eurex_stats_daily.py` +- `lch_cdsclear_daily.py` +- `ec_weekly_oil_bulletin.py` +- `cftc_swaps_weekly.py` +- `cftc_cot.py` +- `b3_historical_quotes.py` + +There are also clear source families: + +### 1. Registry-driven HTTP API sources + +These sources load entity metadata from YAML registries, iterate requested entities, call an API per entity/config, then map provider-specific payloads into the standard long frame. + +Candidates: + +- `bea.py` +- `eia.py` +- `ecb_sdmx.py` +- `eurostat.py` +- `ibge_sidra.py` +- `destatis_genesis.py` + +### 2. Simple single-endpoint API sources + +These do not use registries but still follow the same "call endpoint -> normalize frame -> finalize" model. + +Candidates: + +- `bcb_sgs.py` +- `bls.py` + +### 3. Tabular document sources + +These fetch HTML / CSV / XLSX / ZIP artifacts, parse tables, identify relevant columns, then normalize them. + +Candidates: + +- `eurex_stats_daily.py` +- `lch_cdsclear_daily.py` +- `ec_weekly_oil_bulletin.py` +- `ezoic_adrevenue_daily.py` +- `cme_productslate_reference.py` +- `eurex_refdata_contracts.py` +- `frb_term_structure.py` + +### 4. Archive / bulk-file sources + +These discover or generate one or more artifact URLs, download archives, parse them, and normalize the combined result. + +Candidates: + +- `cftc_cot.py` +- `cftc_swaps_weekly.py` +- `b3_historical_quotes.py` + +### 5. Complex outliers + +These should likely adopt only the shallow shared helpers, not a deep family base class. + +Candidates: + +- `dtcc_ppd.py` +- `mof_jgb.py` +- `philadelphia_spf.py` + +## Design Principles + +### Prefer composable helpers over deep inheritance + +The current module is diverse enough that a large abstract base class would become brittle. The better target is: + +- small shared helper functions +- narrow mixins / base classes for clear source families +- explicit per-source normalization logic + +### Keep one source per file + +The refactor should not collapse unrelated loaders into generic frameworks. The implementation logic for each provider should remain discoverable in its own file. + +### Standardize the outer fetch pipeline + +The biggest immediate win is not parsing logic reuse. It is making the outer source lifecycle consistent: + +- table validation +- empty-frame construction +- `asof_utc` defaults +- common filter/project/finalize behavior +- common HTTP wiring + +### Avoid a mandatory "generic DSL" + +Do not introduce a declarative framework that every source must conform to. A small helper library is lower-risk and more maintainable than a large meta-source layer. + +## Proposed Target Structure + +Introduce a small internal foundation layer under `alphaforge/data/public_web/`: + +- `base.py` +- `schema_helpers.py` +- `finalize.py` +- `tabular.py` +- `registry_api.py` +- `archive.py` + +These names are suggestions, not requirements. The key point is to separate: + +- fetch lifecycle helpers +- schema/empty-frame helpers +- parsing/document helpers +- registry-backed API iteration helpers +- archive/batch helpers + +## Proposed Abstractions + +## A. `PublicWebSourceBase` + +Add a light shared base for HTTP-backed sources. This should solve only the repeated shell around fetching. + +Responsibilities: + +- initialize / inject `CachedHttpClient` +- expose `self._now_utc()` +- validate `q.table` +- create empty frames from schema metadata +- centralize the final `apply_query_filters` + `project_columns` step + +Suggested interface: + +```python +class PublicWebSourceBase(DataSource): + name: str + + def _require_table(self, q: Query, expected: str) -> None: ... + def _empty_frame(self, schema: TableSchema) -> pd.DataFrame: ... + def _finalize( + self, + df: pd.DataFrame, + *, + q: Query, + schema: TableSchema, + time_col: str | None = None, + entity_col: str | None = None, + sort_by: list[str] | None = None, + ) -> pd.DataFrame: ... + def _now_utc(self) -> pd.Timestamp: ... +``` + +This should replace repeated source-level code, not provider-specific parsing code. + +## B. Schema helpers + +Common `TableSchema` construction is repeated across the module, especially for: + +- single-value macro series +- daily market tables +- event tables +- weekly interval-end tables + +Add small helper constructors such as: + +- `single_value_schema(...)` +- `daily_panel_schema(...)` +- `event_table_schema(...)` + +This is mostly a readability improvement and should be kept shallow. + +## C. Finalization helpers + +Create a common helper for the dominant pattern: + +```python +out = apply_query_filters(...) +out = project_columns(...) +return out.sort_values(...).reset_index(drop=True) +``` + +This should live in one place rather than being repeated in nearly every source. + +It should also standardize: + +- empty result columns +- `asof_utc` presence +- default sorting + +## D. Registry-backed API base + +Several sources share a very similar shape: + +- load YAML registry +- require `q.entities` +- look up entity config +- call remote API per entity +- parse provider payload per entity +- assemble standardized rows + +Add a narrow base or helper like: + +```python +class RegistryApiSourceBase(PublicWebSourceBase): + def _load_registry(...) + def _iter_entity_configs(...) +``` + +Likely adopters: + +- `bea.py` +- `eia.py` +- `ecb_sdmx.py` +- `eurostat.py` +- `ibge_sidra.py` +- `destatis_genesis.py` + +Important constraint: + +This base should not try to standardize provider payload shapes. It should only standardize registry loading, entity iteration, and row accumulation. + +## E. Tabular document helpers + +Multiple sources do: + +- fetch bytes +- parse HTML tables or ZIP/CSV/XLSX sheets +- detect date/entity/metric columns using `first_existing(...)` +- normalize into standard columns + +Add shared utilities for: + +- selecting candidate tables by required columns +- resolving date columns from a list of aliases +- resolving metric columns from alias sets +- turning a normalized intermediate frame into a finalized output frame + +This is especially relevant for: + +- `eurex_stats_daily.py` +- `lch_cdsclear_daily.py` +- `ec_weekly_oil_bulletin.py` +- `ezoic_adrevenue_daily.py` + +## F. Archive / batch helpers + +Add helpers for sources that work from year-partitioned or discovered artifact lists: + +- static yearly URL generation +- optional historical-batch URL inclusion +- ZIP member selection +- concatenation and de-duplication + +This would directly benefit: + +- `cftc_cot.py` +- `cftc_swaps_weekly.py` +- `b3_historical_quotes.py` + +The recent CoT refactor already demonstrates the value of this direction. + +## G. Naming and contract cleanup + +Standardize the internal conventions used by all sources: + +- `TABLE` for single-table sources +- `*_TABLE` for multi-table sources +- `name` always equal to registry/source key +- `date` for interval-end / observation date +- `ts_utc` for event timestamps only +- `entity_id` unless the schema intentionally exposes a different entity column + +Also standardize the internal helper naming: + +- `_call(...)` for remote API calls +- `_read_*` for artifact parsing +- `_normalize_*` for provider-to-canonical transformation +- `_discover_*` or `_list_*` for remote artifact enumeration + +## Recommended Phases + +## Phase 1: Extract the safe shallow helpers + +Scope: + +- add `PublicWebSourceBase` +- add finalization helpers +- add schema helpers + +Do not change provider logic yet. + +Candidate first adopters: + +- `bcb_sgs.py` +- `bea.py` +- `eia.py` + +Success condition: + +- these sources get materially smaller without semantic changes +- tests remain unchanged except for helper-specific coverage + +## Phase 2: Registry-backed API family + +Scope: + +- add `RegistryApiSourceBase` +- migrate registry-driven loaders to the new base + +Candidate order: + +1. `eia.py` +2. `bea.py` +3. `eurostat.py` +4. `ecb_sdmx.py` +5. `ibge_sidra.py` +6. `destatis_genesis.py` + +Success condition: + +- registry loading / entity iteration duplication is eliminated +- per-provider parsing remains local to each file + +## Phase 3: Tabular-document family + +Scope: + +- add table-selection and alias-resolution helpers +- migrate HTML/table-driven sources + +Candidate order: + +1. `eurex_stats_daily.py` +2. `lch_cdsclear_daily.py` +3. `ec_weekly_oil_bulletin.py` +4. `ezoic_adrevenue_daily.py` + +Success condition: + +- repeated `first_existing(...)` / `ensure_date_utc(...)` / `entity_id` construction logic is reduced +- source-specific column semantics still remain explicit + +## Phase 4: Archive / batch family + +Scope: + +- extract shared archive helpers +- migrate `cftc_swaps_weekly.py` and `b3_historical_quotes.py` +- keep `cftc_cot.py` as the reference implementation for the family + +Success condition: + +- URL generation, historical-batch support, ZIP reading, and empty-frame behavior are standardized + +## Phase 5: Outlier integration + +Scope: + +- apply only shallow helper adoption to `dtcc_ppd.py`, `mof_jgb.py`, and `philadelphia_spf.py` + +These should likely use: + +- common HTTP setup +- common finalization +- common empty-frame helpers + +They should probably not be forced into the family bases from phases 2-4. + +## Phase 6: Public module cleanup + +Scope: + +- clean import/export organization in `__init__.py` +- keep registry construction simple in `registry.py` +- update developer docs for how to add a new source + +Deliverables: + +- a short "how to add a public_web source" document +- explicit source family guidance + +## Risks + +### 1. Over-abstraction + +Risk: + +Trying to unify all sources under one base class will make the abstraction worse than the current repetition. + +Mitigation: + +- keep bases shallow +- use family-level helpers only where there are at least 3 strong adopters +- treat `dtcc_ppd.py`, `mof_jgb.py`, and `philadelphia_spf.py` as exceptions by default + +### 2. Silent output drift + +Risk: + +Refactoring the outer fetch flow can change empty-frame columns, sorting, time normalization, or `asof_utc` behavior. + +Mitigation: + +- add focused tests around output columns and sort order before migration +- migrate a few sources at a time +- preserve existing schema contracts exactly + +### 3. Test churn without value + +Risk: + +Changing too many files at once can produce large low-signal diffs. + +Mitigation: + +- phase the work +- land family-level helpers first +- migrate only a small source set per PR + +### 4. Forcing registry-backed assumptions into non-registry sources + +Risk: + +Registry-backed sources and bulk archive sources have very different query shapes. + +Mitigation: + +- separate the bases +- do not let `RegistryApiSourceBase` leak into archive or document loaders + +## Suggested Initial PR Breakdown + +### PR 1 + +- add `PublicWebSourceBase` +- add finalization helpers +- migrate `bcb_sgs.py`, `bea.py`, `eia.py` + +### PR 2 + +- add `RegistryApiSourceBase` +- migrate remaining registry-driven sources + +### PR 3 + +- add tabular helpers +- migrate `eurex_stats_daily.py`, `lch_cdsclear_daily.py`, `ec_weekly_oil_bulletin.py` + +### PR 4 + +- add archive helpers +- migrate `cftc_swaps_weekly.py`, `b3_historical_quotes.py` + +### PR 5 + +- shallow helper adoption for `dtcc_ppd.py`, `mof_jgb.py`, `philadelphia_spf.py` +- docs update + +## Linear Ticket Mirror + +These tables mirror the current Linear public-web refactor tickets and define +the intended implementation order for subsequent agents. + +Linear routing for this plan: + +- Team: `ALP` (`alphaforge`) +- Project: `alphaforge` +- Umbrella issue: `ALP-1` + +Rules for coding agents: + +- Linear is the source of truth for ticket state. +- This plan file is the repo-local execution mirror for subsequent agents. +- Implement the earliest ticket in the ordered queue below whose status is not + `Done` and whose upstream prerequisites are already satisfied. +- Skip tickets whose table row is `Done`. +- Epic rows are tracking rows. Do not pick them up before their earlier child + rows unless the epic has no remaining open child slices. +- Before coding, print the current ticket number and its plain-English goal on + screen. +- Follow the implementation workflow in `AGENTS.md`: + - review upstream tickets and notes first + - implement with TDD + - update docs after tests pass + - leave the structured closeout note in Linear + - mark the Linear ticket `Done` + - only then update the corresponding row in this file +- Update this table only after the ticket is closed in Linear. + +Status mirror last synced: `2026-04-04` + +### Ordered Epic Queue + +| Ticket | Status | +| --- | --- | +| `ALP-1` Public web: refactor shared abstractions and source-family cleanup | Done | + +### Ordered Public Web Refactor Queue + +| Ticket | Status | +| --- | --- | +| `ALP-2` Public web foundation: add shared source base, schema helpers, and finalization helpers | Done | +| `ALP-3` Registry APIs: add a shared base and migrate registry-backed public sources | Done | +| `ALP-4` Simple APIs: migrate remaining single-endpoint public sources onto shared helpers | Done | +| `ALP-5` Tabular loaders: extract document helpers and migrate table-driven public sources | Done | +| `ALP-6` Archive loaders: extract batch helpers and migrate bulk-file public sources | Done | +| `ALP-7` Public web outliers: adopt shallow shared helpers in complex sources | Done | +| `ALP-8` Public web docs: add source-authoring guidance and complete module cleanup | Done | + +### Cross-Ticket Sequencing Constraints + +Use the queue order above, but also respect these concrete handoff rules: + +- Treat `ALP-1` as a tracking epic. Start implementation with `ALP-2`. +- Do not start `ALP-3`, `ALP-4`, `ALP-5`, `ALP-6`, or `ALP-7` until `ALP-2` is + landed. +- Treat `ALP-3`, `ALP-4`, `ALP-5`, `ALP-6`, and `ALP-7` as an ordered + implementation queue even though the Linear dependency graph only records the + shared blocker on `ALP-2`. +- Use `ALP-3` before `ALP-4` so the registry-backed family is migrated before + the remaining simple API cleanup slice. +- Use `ALP-5` before `ALP-6` so the table/document helper layer lands before + the archive/batch helper layer. +- Use `ALP-7` only after the family migrations are done; it is a shallow + cleanup pass for true outliers, not a substitute for the family abstractions. +- Do not start `ALP-8` until `ALP-2`, `ALP-3`, `ALP-4`, `ALP-5`, `ALP-6`, and + `ALP-7` are landed. + +### Agent Pickup Directive + +If you hand this file to a coding agent for end-to-end implementation, the +agent should: + +1. start at the top of the epic queue +2. within the active queue, pick the first non-`Done` ticket whose + prerequisites are satisfied +3. implement only that ticket's scoped outcome +4. close the ticket in Linear first +5. then update the mirrored status row in this plan before moving on + +Do not skip ahead to later slices because a later ticket looks smaller. This +plan is intentionally ordered to land the shared foundation first, then the +clear source families, then the outlier/doc cleanup work. + +## Epic Close-Out + +The epic is complete. + +Delivered outcomes: + +- shared foundation helpers in `base.py`, `finalize.py`, and `schema_helpers.py` +- narrow family helpers for registry-backed APIs, tabular/document loaders, and + archive/batch loaders +- shallow helper adoption in the outlier sources without forcing deeper + inheritance +- public-web authoring guidance, API docs, quickstart updates, and package + export cleanup + +Final validation: + +- `/Users/steveyang/miniforge3/bin/python -m pytest tests/test_public_web_registry_exports.py tests/public_web -k 'not live_sources' tests/test_cftc_dtcc_adapter.py -q` +- `/Users/steveyang/miniforge3/bin/python -m ruff check .` +- `/Users/steveyang/miniforge3/bin/python -m mkdocs build --strict` + +## Success Criteria + +The refactor is successful if: + +- new public-web loaders can be added with materially less boilerplate +- common fetch/finalize behavior is centralized +- registry-backed and archive-backed families have explicit support +- complex outliers are not damaged by the abstraction effort +- no public schemas or table names change +- source tests remain one-to-one with modules and retain current behavior + +## Non-Goals + +- rewriting all loaders into a declarative framework +- changing public table names or entity-id contracts +- merging unrelated source files +- replacing pandas-based parsing with a different execution model +- changing `CachedHttpClient` transport semantics in the same refactor unless required by a separate bug fix + +## Immediate Recommendation + +Start with Phase 1 only. The module already has enough evidence that shallow fetch/finalize abstraction is worth it. That first slice should produce a smaller diff, validate the direction, and make the later family-specific abstractions easier to judge with real code instead of speculation. diff --git a/docs/api/data-context.md b/docs/api/data-context.md index 3edda89..711ba6e 100644 --- a/docs/api/data-context.md +++ b/docs/api/data-context.md @@ -1,3 +1,17 @@ # DataContext +`DataContext.fetch(...)`, `fetch_many(...)`, and `prefetch(...)` are the +canonical public data-loading entry points for new code. + +For the shortest happy path, build contexts with `DataContext.from_adapters(...)` +and load a table with `DataContext.load(...)`. + +If multiple adapters serve the same dataset, canonical loads must disambiguate +through `default_sources` or an explicit `source=` override. Alphaforge no +longer falls back to an arbitrary first-registered adapter for shared datasets. + +`DataContext.sources` and `fetch_panel(...)` remain available as +backward-compatibility surfaces for legacy `DataSource`-backed loaders, but +they are no longer the preferred external loading contract. + ::: alphaforge.data.context diff --git a/docs/api/dataset-spec.md b/docs/api/dataset-spec.md index 48dbee1..28a5b1e 100644 --- a/docs/api/dataset-spec.md +++ b/docs/api/dataset-spec.md @@ -1,3 +1,16 @@ # Dataset Spec +`DatasetSpec` is the canonical research-assembly contract. + +Notable composition surfaces: + +- `FeatureRequest` +- `FeatureRequestGroup` +- `JoinPolicy` +- `MissingnessPolicy` + +Built-in notebook-ready template families live alongside the spec. See +[Market Templates](market-templates.md) for the canonical lag-return and +rolling-volatility helpers. + ::: alphaforge.features.dataset_spec diff --git a/docs/api/market-templates.md b/docs/api/market-templates.md new file mode 100644 index 0000000..f5a6001 --- /dev/null +++ b/docs/api/market-templates.md @@ -0,0 +1,8 @@ +# Market Templates + +Notebook-ready built-in templates for common market-price research features. + +These templates are intended to replace repo-local helper cells for common +return and trailing-volatility feature families. + +::: alphaforge.features.market diff --git a/docs/api/pit-accessor.md b/docs/api/pit-accessor.md index 4cd5eb1..909ad8c 100644 --- a/docs/api/pit-accessor.md +++ b/docs/api/pit-accessor.md @@ -1,3 +1,6 @@ # PIT Accessor +For the shortest local bootstrap, use `PITAccessor.open(path)` and then call +the snapshot or revision helpers you need. + ::: alphaforge.pit.accessor diff --git a/docs/api/pit-missingness.md b/docs/api/pit-missingness.md index ea3438c..65f078c 100644 --- a/docs/api/pit-missingness.md +++ b/docs/api/pit-missingness.md @@ -1,3 +1,6 @@ -# PIT Missingness +# Time Missingness -::: alphaforge.pit.missingness +Canonical import path: `alphaforge.time.missingness`. +`alphaforge.pit.missingness` remains as a compatibility shim. + +::: alphaforge.time.missingness diff --git a/docs/api/pit-queries.md b/docs/api/pit-queries.md new file mode 100644 index 0000000..6923db5 --- /dev/null +++ b/docs/api/pit-queries.md @@ -0,0 +1,3 @@ +# PIT Ref Queries + +::: alphaforge.pit.queries diff --git a/docs/api/pit-release-rules.md b/docs/api/pit-release-rules.md index df1b9a3..3a0c467 100644 --- a/docs/api/pit-release-rules.md +++ b/docs/api/pit-release-rules.md @@ -1,3 +1,6 @@ -# PIT Release Rules +# Time Release Rules -::: alphaforge.pit.release_rules +Canonical import path: `alphaforge.time.release_rules`. +`alphaforge.pit.release_rules` remains as a compatibility shim. + +::: alphaforge.time.release_rules diff --git a/docs/api/public-web.md b/docs/api/public-web.md index c9da940..26bf9ea 100644 --- a/docs/api/public-web.md +++ b/docs/api/public-web.md @@ -1,7 +1,95 @@ -# Public Web Sources +# Public Web API -Public web source modules are documented in development branches that include -`alphaforge.data.public_web`. +`alphaforge.data.public_web` contains source loaders for public datasets that do +not justify a dedicated adapter package but still need stable schemas, +deterministic normalization, and point-in-time-safe `asof_utc` handling. -The current `main` branch API reference only covers modules that are part of the -published package build. +These public-web loaders are part of the legacy/raw-loader `DataSource` +surface. They remain supported because the public-web pack has not been fully +adapterized, but they are not the canonical general loading contract for new +code outside this loader family. + +## Public entry points + +- `alphaforge.data.public_web.default_public_web_sources()` builds the default + raw-loader registry used by compatibility-oriented `DataContext` bootstrap + code. +- `alphaforge.data.public_web` re-exports the concrete public loader classes. +- `alphaforge.default_public_web_sources()` mirrors the same registry at the + package root for convenience. + +## Default source catalog + +### Registry-backed APIs + +| Source name | Class | Table(s) | +| --- | --- | --- | +| `bcb_sgs` | `BCBSGSDataSource` | `bcb_sgs_series` | +| `bea` | `BEADataSource` | `bea_series` | +| `bls` | `BLSDataSource` | `bls_series` | +| `destatis_genesis` | `DestatisGenesisDataSource` | `destatis_series` | +| `ecb_sdmx` | `ECBSDMXDataSource` | `ecb_sdmx_series` | +| `eia` | `EIADataSource` | `eia_series` | +| `eurostat` | `EurostatDataSource` | `eurostat_series` | +| `ibge_sidra` | `IBGESidraDataSource` | `ibge_sidra_series` | + +### Tabular and document loaders + +| Source name | Class | Table(s) | +| --- | --- | --- | +| `cme_productslate` | `CMEProductSlateSource` | `cme.productslate.reference` | +| `ec_weekly_oil_bulletin` | `ECWeeklyOilBulletinDataSource` | `ec_oil_bulletin_weekly` | +| `eurex_refdata_contracts` | `EurexRefdataContractsSource` | `eurex.refdata.contracts` | +| `eurex_stats_daily` | `EurexStatsDailySource` | `eurex.stats.daily` | +| `ezoic_adrevenue_daily` | `EzoicAdRevenueDailySource` | `ezoic.adrevenue.daily` | +| `frb_term_structure` | `FRBTermStructureBenchmarkSource` | `frb.term_structure` | +| `lch_cdsclear_daily` | `LCHCDSClearDailySource` | `lch.cdsclear.daily` | + +### Archive and batch loaders + +| Source name | Class | Table(s) | +| --- | --- | --- | +| `b3_historical_quotes` | `B3HistoricalQuotesDataSource` | `b3_equity_quotes_daily` | +| `cftc_cot` | `CFTCCoTSource` | `cftc.cot.tff` | +| `cftc_cot_disagg` | `CFTCDisaggregatedCoTSource` | `cftc.cot.disagg` | +| `cftc_swaps_weekly` | `CFTCWeeklySwapsSource` | `cftc.swaps.weekly` | + +Archive-backed CFTC loaders now raise an explicit failure when a requested +archive cannot be downloaded or parsed, rather than silently skipping the bad +year and returning partial history. + +### Outliers and provider-specific loaders + +| Source name | Class | Table(s) | +| --- | --- | --- | +| `anp_fuel_prices` | `ANPFuelPricesDataSource` | `anp_fuel_prices_weekly` | +| `dtcc_ppd` | `DTCCPPDSource` | `dtcc.ppd.events`, `dtcc.ppd.daily` | +| `mof_jgb_yields` | `MOFJGBYieldCurveSource` | `mof.jgb.yields` | +| `philadelphia_spf` | `PhiladelphiaSPFMeanLevelSource` | `philadelphia.spf.mean_level` | + +`DTCCPPDSource` remains the low-level raw-loader surface for DTCC. Canonical +adapter-backed access for the first DTCC families now lives in +`alphaforge.data.sources.dtcc` through `DTCCFXAdapter` (`dtcc.fx`) and +`DTCCIRSAdapter` (`dtcc.irs`). + +## Shared helper layers + +These helpers reduce boilerplate inside `alphaforge.data.public_web`, but they +are still internal implementation utilities rather than a broad compatibility +promise for third-party extensions. + +| Helper module | Purpose | +| --- | --- | +| `base.py` | shared HTTP client setup, empty-frame construction, and finalization | +| `finalize.py` | schema-aware projection, sorting, entity/date filtering, and `asof_utc` handling | +| `schema_helpers.py` | concise `TableSchema` builders for daily panels, single-value series, and event tables | +| `registry_api.py` | shallow base for registry-driven entity loaders | +| `tabular.py` | helpers for HTML/XLSX/CSV document parsing and column resolution | +| `archive.py` | helpers for archive discovery, deterministic fetch planning, yearly archive planning, and ZIP member reads | + +## Authoring guidance + +For new loaders or refactors, use the contributor workflow in the +[public-web source authoring guide](../guides/public-web-source-authoring.md). +That guide covers helper-family selection, registry and export wiring, +defensive parsing, and the required test and doc updates. diff --git a/docs/api/source-adapters.md b/docs/api/source-adapters.md index 270ae33..7b2ca37 100644 --- a/docs/api/source-adapters.md +++ b/docs/api/source-adapters.md @@ -1,6 +1,16 @@ # Source Adapters -Unified data layer providing a single query interface across all data sources. +Canonical unified data-loading surface for Alphaforge. + +New code should register `SourceAdapter` implementations in +`DataContext.adapters` and load data through: + +- `DataContext.fetch(...)` +- `DataContext.fetch_many(...)` +- `DataContext.prefetch(...)` + +Legacy `DataSource` and `fetch_panel(...)` usage remains supported only as a +compatibility boundary while older loaders migrate. ## Protocol & Base @@ -34,8 +44,54 @@ Unified data layer providing a single query interface across all data sources. ### DTCC (Swap Derivatives) +The preferred DTCC adapter-backed datasets now split along the first concrete +product families: + +- `DTCCFXAdapter` serves `dtcc.fx` for FX forwards and swaps +- `DTCCIRSAdapter` serves `dtcc.irs` for interest rate swaps + +These family adapters own `DTCCPPDSource` construction internally, stamp +product-family-specific PIT lineage, and keep dataset routing distinct in +`DataContext`. + +```python +from alphaforge.data.context import DataContext +from alphaforge.data.sources.dtcc import DTCCFXAdapter, DTCCIRSAdapter + +ctx = DataContext.from_adapters( + DTCCFXAdapter(list_provider=..., artifact_provider=...), + DTCCIRSAdapter(list_provider=..., artifact_provider=...), +) + +fx = ctx.load( + "dtcc.fx", + columns=["value"], + entities=["dtcc.fx.dtccppd.fx.fx_forward.usd.1m.trade_count"], +) +irs = ctx.load( + "dtcc.irs", + columns=["value"], + entities=["dtcc.irs.dtccppd.rates.interest_rate_swap.usd.5y.trade_count"], +) +``` + +`DTCCAdapter` remains available as the broader generic `dtcc.ppd` wrapper, and +future DTCC product-family adapters should build on `DTCCPPDAdapterBase` so +raw-loader fetch wiring, cache behavior, and PIT transform plumbing stay +consistent. + +::: alphaforge.data.sources.dtcc.DTCCFXAdapter + +::: alphaforge.data.sources.dtcc.DTCCIRSAdapter + ::: alphaforge.data.sources.dtcc.DTCCAdapter +::: alphaforge.data.sources.dtcc.DTCCPPDAdapterBase + ## PIT Compatibility Bridge +`SourceAdapterPITCompat` is a temporary migration bridge. It remains documented +because supported downstream PIT integrations still rely on it, but it is not +the long-term canonical adapter model. + ::: alphaforge.pit.adapters.source_adapter_compat.SourceAdapterPITCompat diff --git a/docs/api/time-ref-period.md b/docs/api/time-ref-period.md new file mode 100644 index 0000000..059bab0 --- /dev/null +++ b/docs/api/time-ref-period.md @@ -0,0 +1,5 @@ +# Time Ref Periods + +Canonical import path: `alphaforge.time.ref_period`. + +::: alphaforge.time.ref_period diff --git a/docs/getting-started/quickstart-dataset.md b/docs/getting-started/quickstart-dataset.md index 1d5ec05..1f2429b 100644 --- a/docs/getting-started/quickstart-dataset.md +++ b/docs/getting-started/quickstart-dataset.md @@ -1,17 +1,22 @@ # Quickstart: Build a Dataset -This example uses a dummy source with two feature families and one target. +This example uses the canonical adapter path plus built-in notebook-ready +market templates. ```python import numpy as np import pandas as pd +from alphaforge.data.adapter import SourceAdapterBase from alphaforge.data.context import DataContext from alphaforge.data.query import Query +from alphaforge.data.types import FetchResult +from alphaforge.features import LagReturnsTemplate, RollingVolatilityTemplate from alphaforge.features.dataset_builder import build_dataset from alphaforge.features.dataset_spec import ( DatasetSpec, FeatureRequest, + FeatureRequestGroup, JoinPolicy, MissingnessPolicy, TargetRequest, @@ -19,49 +24,37 @@ from alphaforge.features.dataset_spec import ( UniverseSpec, ) from alphaforge.features.target_template import TargetFrame -from alphaforge.store.local_parquet import LocalParquetStore +from alphaforge.features.template import SliceSpec from alphaforge.time.calendar import TradingCalendar -from examples.dummy_source import DummySource -from examples.features_lag_returns import LagReturnsTemplate -from examples.features_macro_carry import MacroCarryTemplate -cal = TradingCalendar("XNYS", tz="UTC") -dates = cal.sessions("2020-01-01", "2020-03-31") -entities = ["AAA", "BBB"] - -rng = np.random.default_rng(123) -rows = [] -for entity in entities: - px = 100 + np.cumsum(rng.normal(0, 1, size=len(dates))) - for date, price in zip(dates, px): - rows.append({"date": date, "entity_id": entity, "close": float(price)}) - -ohlcv = pd.DataFrame(rows) -macro = pd.DataFrame( - [ - {"date": pd.Timestamp("2020-01-31"), "entity_id": "CPI", "value": 1.0}, - {"date": pd.Timestamp("2020-02-29"), "entity_id": "CPI", "value": 2.0}, - {"date": pd.Timestamp("2020-03-31"), "entity_id": "CPI", "value": 3.0}, - ] -) - -ctx = DataContext( - sources={"dummy": DummySource(ohlcv_long=ohlcv, macro_long=macro)}, - calendars={"XNYS": cal}, - store=LocalParquetStore("./alphaforge_demo_store"), -) +class InMemoryMarketAdapter(SourceAdapterBase): + source_name = "market" + datasets = frozenset({"market.ohlcv"}) + + def __init__(self, frame: pd.DataFrame) -> None: + self._frame = frame.copy() + + def fetch(self, query: Query, *, max_staleness=None) -> FetchResult: + frame = self._frame[self._frame["series_key"].isin(query.entities)].copy() + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.start is not None: + frame = frame[obs >= query.start] + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.end is not None: + frame = frame[obs <= query.end] + + keep = ["series_key", "obs_date"] + list(query.columns) + return FetchResult( + data=frame[keep].reset_index(drop=True), + source=self.source_name, + dataset=query.table, + is_pit=False, + cached_at=None, + ) -features = [ - FeatureRequest( - template=LagReturnsTemplate(), - params={"lags": 5, "source": "dummy", "table": "market.ohlcv", "price_col": "close"}, - ), - FeatureRequest( - template=MacroCarryTemplate(), - params={"source": "dummy", "table": "macro.series", "value_col": "value", "method": "ffill"}, - ), -] + def list_entities(self, dataset: str) -> list[str]: + return sorted(self._frame["series_key"].unique()) class NextDaySqLogRetTarget: @@ -72,38 +65,109 @@ class NextDaySqLogRetTarget: def fit(self, ctx, params, fit_slice): return None - def transform(self, ctx, params, slice, state): - panel = ctx.fetch_panel( - "dummy", - Query( - table="market.ohlcv", - columns=["close"], - start=slice.start, - end=slice.end, - entities=slice.entities, - asof=slice.asof, - grid=slice.grid, - ), + def transform(self, ctx, params, slice: SliceSpec, state): + result = ctx.load( + "market.ohlcv", + columns=["close"], + start=slice.start, + end=slice.end, + entities=slice.entities, + asof=slice.asof, + grid=slice.grid, + source="market", ) - px = panel.df["close"].astype(float) - logret = np.log(px).groupby(level="entity_id").diff() + frame = result.data.copy() + calendar = ctx.calendars["XNYS"] + frame["ts_utc"] = [ + calendar.session_close_utc(ts) + for ts in pd.to_datetime(frame["obs_date"], utc=True) + ] + prices = ( + frame.set_index(["ts_utc", "series_key"])["close"] + .rename_axis(index=["ts_utc", "entity_id"]) + .sort_index() + .astype(float) + ) + logret = np.log(prices).groupby(level="entity_id").diff() y = (logret.groupby(level="entity_id").shift(-1) ** 2).rename("y") return TargetFrame(y=y, meta={"definition": "(logret_{t+1})^2"}) +cal = TradingCalendar("XNYS", tz="UTC") +dates = cal.sessions("2024-01-02", "2024-02-16") +entities = ["AAA", "BBB"] + +rng = np.random.default_rng(123) +rows = [] +for entity in entities: + prices = 100 + np.cumsum(rng.normal(0, 1, size=len(dates))) + for obs_date, close in zip(dates, prices, strict=True): + rows.append({"series_key": entity, "obs_date": obs_date, "close": float(close)}) + +ctx = DataContext.from_adapters( + InMemoryMarketAdapter(pd.DataFrame(rows)), + calendars={"XNYS": cal}, + store=None, +) + +features = [ + FeatureRequestGroup( + key="volatility", + tags={"recipe": "quickstart"}, + requests=[ + FeatureRequest( + template=LagReturnsTemplate(), + key="returns", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "lags": [1, 5, 10], + }, + ), + FeatureRequest( + template=RollingVolatilityTemplate(), + key="trailing_vol", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "windows": [5, 10], + "lag": 1, + "annualization_factor": 252, + }, + ), + ], + ) +] + spec = DatasetSpec( universe=UniverseSpec(entities=entities), - time=TimeSpec(start=pd.Timestamp("2020-01-01"), end=pd.Timestamp("2020-03-31"), calendar="XNYS", grid="B"), + time=TimeSpec( + start=pd.Timestamp("2024-01-02"), + end=pd.Timestamp("2024-02-16"), + calendar="XNYS", + grid="B", + ), features=features, - target=TargetRequest(template=NextDaySqLogRetTarget(), params={}, horizon=1, name="y"), + target=TargetRequest( + template=NextDaySqLogRetTarget(), + params={}, + horizon=1, + name="y", + ), join_policy=JoinPolicy(how="inner", sort_index=True), missingness=MissingnessPolicy(final_row_policy="drop_if_any_nan"), name="demo_dataset", ) -artifact = build_dataset(ctx, spec, persist=True) +artifact = build_dataset(ctx, spec, persist=False) print(artifact.X.shape) print(int(artifact.y.notna().sum())) ``` -See [Dataset Spec guide](../guides/dataset-spec.md) for full options. +For a full runnable version of this pattern, see +`examples/volatility_dataset_recipe.py`. + +See [Dataset Spec guide](../guides/dataset-spec.md) for the full contract and +[Research Recipes](../guides/research-recipes.md) for notebook-shaped patterns. diff --git a/docs/getting-started/quickstart-public-web.md b/docs/getting-started/quickstart-public-web.md index 32c2256..cfd2420 100644 --- a/docs/getting-started/quickstart-public-web.md +++ b/docs/getting-started/quickstart-public-web.md @@ -1,28 +1,40 @@ # Quickstart: Public Web Loaders -Alphaforge includes a public web loader pack under `alphaforge.data.public_web`. +Alphaforge ships a public loader pack under `alphaforge.data.public_web` for +public datasets that need stable schemas but not a full custom adapter stack. -## Supported table families +These loaders currently sit on the `DataSource` side of the compatibility +boundary. They are appropriate when you want raw long-form frames from the +public-web pack directly, but they are not the canonical general-purpose +loading path for new adapter-based integrations. -- `dtcc.ppd.events` -- `dtcc.ppd.daily` -- `cftc.swaps.weekly` -- `eurex.stats.daily` -- `eurex.refdata.contracts` -- `lch.cdsclear.daily` -- `cme.productslate.reference` -- `ezoic.adrevenue.daily` +## Loader families + +- Registry-backed APIs: `bcb_sgs_series`, `bea_series`, `bls_series`, + `destatis_series`, `ecb_sdmx_series`, `eia_series`, `eurostat_series`, + `ibge_sidra_series` +- Tabular and document loaders: `cme.productslate.reference`, + `ec_oil_bulletin_weekly`, `eurex.refdata.contracts`, `eurex.stats.daily`, + `ezoic.adrevenue.daily`, `frb.term_structure`, `lch.cdsclear.daily` +- Archive and batch loaders: `b3_equity_quotes_daily`, `cftc.cot.tff`, + `cftc.cot.disagg`, `cftc.swaps.weekly` +- Provider-specific outliers: `anp_fuel_prices_weekly`, `dtcc.ppd.events`, + `dtcc.ppd.daily`, `mof.jgb.yields`, `philadelphia.spf.mean_level` ## Minimal usage +This quickstart intentionally uses the raw loader objects directly: + ```python import pandas as pd +from alphaforge import default_public_web_sources from alphaforge.data.query import Query -from alphaforge.data.public_web import DTCCPPDSource, EurexStatsDailySource -dtcc = DTCCPPDSource(artifact_provider=...) -eurex = EurexStatsDailySource() +sources = default_public_web_sources() + +dtcc = sources["dtcc_ppd"] +cot = sources["cftc_cot_disagg"] dtcc_daily = dtcc.fetch( Query( @@ -33,22 +45,39 @@ dtcc_daily = dtcc.fetch( ) ) -eurex_stats = eurex.fetch( +cot_frame = cot.fetch( Query( - table="eurex.stats.daily", - columns=["volume", "open_interest"], + table="cftc.cot.disagg", + columns=["value"], + entities=["wheat_srw"], + start=pd.Timestamp("2025-01-01", tz="UTC"), ) ) ``` -## Testing loaders +If you instantiate a source directly instead of using +`default_public_web_sources()`, keep constructor-specific requirements explicit. +For example, `EurexRefdataContractsSource` requires an `api_url`. + +If you are building a broader application-level loading layer, prefer wrapping +that usage behind `SourceAdapter` plus `DataContext.fetch(...)` rather than +treating direct `DataSource.fetch(...)` calls as the primary public contract. + +## Validation workflow + +Fast local checks: ```bash -pytest -o addopts='' tests/public_web +/Users/steveyang/miniforge3/bin/python -m pytest tests/public_web -k 'not live_sources' +/Users/steveyang/miniforge3/bin/python -m ruff check . ``` Optional live tests: ```bash -ALPHAFORGE_NETWORK_TESTS=1 pytest -o addopts='' tests/public_web/test_live_sources.py +ALPHAFORGE_NETWORK_TESTS=1 /Users/steveyang/miniforge3/bin/python -m pytest \ + tests/public_web/test_live_sources.py -q ``` + +For contributor guidance on adding or refactoring loaders, see the +[public-web source authoring guide](../guides/public-web-source-authoring.md). diff --git a/docs/guides/contracts-and-benchmarks.md b/docs/guides/contracts-and-benchmarks.md new file mode 100644 index 0000000..3fea2c1 --- /dev/null +++ b/docs/guides/contracts-and-benchmarks.md @@ -0,0 +1,91 @@ +# Core Platform Contracts And Benchmarks + +Use the downstream-inspired contract suite and PIT benchmark harness as the +minimum regression gate for public-surface work in the core platform epic. + +## Contract suite + +The contract tests are intentionally modeled on the three downstream usage +patterns that shaped the roadmap: + +- `tests/contracts/test_nowcast_pit_contract.py` + covers ref-period PIT snapshots and panel-building in a nowcast-style flow +- `tests/contracts/test_volatility_dataset_contract.py` + covers adapter-backed `DataContext` loading plus notebook-style volatility + dataset assembly +- `tests/contracts/test_operations_contract.py` + covers release-aware health reporting and archive fetch planning +- `tests/contracts/test_pit_benchmarks.py` + smoke-tests the benchmark harness interface so performance tracking itself + stays callable + +Run the full slice with: + +```bash +python -m pytest tests/contracts +``` + +## When this gate is required + +Run `tests/contracts` whenever a change touches: + +- `alphaforge.time` +- `alphaforge.pit` +- `alphaforge.data.context` +- `alphaforge.features.dataset_*` +- built-in market templates +- release-aware health reporting or archive fetch planning +- any compatibility shim or migration surface that downstream repos still use + +Future deprecations should cite the specific contract file that covers the +behavior they are changing. + +## PIT benchmark harness + +The PIT benchmark harness lives in `benchmarks/pit.py` and exercises the two +highest-value regression paths from the roadmap: + +- `PITAccessor.snapshot_ref(...)` +- `PITAccessor.build_snapshot_panel_long(...)` + +Run it with: + +```bash +python -m benchmarks.pit +``` + +Optional knobs: + +- `--iterations` +- `--periods` +- `--series-count` +- `--revisions-per-period` + +Treat the benchmark output as a baseline snapshot, not as a cross-machine SLA. +Compare runs on the same developer machine or CI runner class before drawing +performance conclusions. + +## Current baseline snapshot + +Baseline captured from: + +```bash +python -m benchmarks.pit --iterations 5 --periods 40 --series-count 3 --revisions-per-period 2 +``` + +| Captured at (UTC) | Rows | `snapshot_ref(...)` median | `build_snapshot_panel_long(...)` median | +| --- | ---: | ---: | ---: | +| `2026-04-05T03:35:01.847144+00:00` | 240 | 5.724 ms | 11.518 ms | + +## Migration and deprecation workflow + +Before removing a compatibility surface or tightening a public contract: + +1. run the targeted subsystem tests +2. run `python -m pytest tests/contracts` +3. rerun `python -m benchmarks.pit` if the change touches PIT retrieval paths +4. update `doc/plan/migration_note.md` with the downstream impact +5. update `doc/plan/post_migration_plan.md` if a temporary bridge was added + +That discipline keeps migration work anchored to explicit checks instead of +repo-local assumptions. diff --git a/docs/guides/core-platform-architecture.md b/docs/guides/core-platform-architecture.md new file mode 100644 index 0000000..34e5b7a --- /dev/null +++ b/docs/guides/core-platform-architecture.md @@ -0,0 +1,98 @@ +# Core Platform Architecture + +This guide describes the canonical layer boundaries established by the core +platform roadmap. + +## Layer 1: Temporal semantics + +`alphaforge.time` is the canonical home for shared time semantics: + +- `RefPeriod` +- `ReleaseRule` +- release-rule helpers such as `FixedLagMonths` +- missingness classification + +Use this layer when code needs to talk about reference periods, release timing, +or missingness explicitly. Do not treat legacy PIT import paths as co-equal +public directions. + +## Layer 2: PIT flagship API + +`alphaforge.pit` is the flagship analytical layer built on top of the temporal +core. + +The canonical ref-period query path is: + +- `RefSnapshotQuery` +- `RefRevisionQuery` +- `PITAccessor.snapshot_ref(...)` +- `PITAccessor.revisions_ref(...)` +- `PITAccessor.build_snapshot_panel_long(...)` + +Lineage and causality diagnostics also live here, so PIT transforms, +explainability, and revision-aware panels share one semantic surface. + +## Layer 3: Source access and routing + +The canonical data-loading path is adapter-first: + +- `SourceAdapter` +- `DataContext.from_adapters(...)` +- `DataContext.fetch(...)` +- `DataContext.fetch_many(...)` +- `DataContext.load(...)` + +`DataSource` and `DataContext.sources` remain for raw-loader and migration +scenarios, but they are not the long-term public direction. + +## Layer 4: Dataset algebra and research UX + +Research assembly sits in `alphaforge.features` and `alphaforge.features.dataset_spec`. + +The intended composition model is: + +- `DatasetSpec` +- `FeatureRequestGroup` +- built-in templates such as `LagReturnsTemplate` + and `RollingVolatilityTemplate` +- `build_dataset(...)` + +This layer should express joins, missingness policy, grouped feature families, +and notebook-friendly recipes without downstream helper scaffolding. + +## Layer 5: Operations and observability + +Operational data workflows build on the same canonical surfaces: + +- release-aware source health via `SourceHealthPolicy` + and `build_health_report(...)` +- deterministic archive planning via `ArchiveFetchPlanEntry`, + `discover_archive_fetches(...)`, and `iter_yearly_archive_fetches(...)` + +These APIs are intended to keep health monitoring and archival ingestion +explicit, typed, and reusable across public-web sources. + +## Layer 6: Stability and migration discipline + +Stability is enforced with two repo-local gates: + +- `tests/contracts` +- `python -m benchmarks.pit` + +Use them alongside targeted subsystem tests whenever work changes a canonical +surface or a temporary compatibility path. + +## Canonical versus compatibility surfaces + +Use the canonical side for new work: + +| Area | Canonical surface | Compatibility-only or legacy surface | +| --- | --- | --- | +| Temporal semantics | `alphaforge.time.*` | `alphaforge.pit.release_rules`, `alphaforge.pit.missingness` | +| PIT ref queries | `snapshot_ref(...)`, `revisions_ref(...)` | `get_snapshot_ref(...)`, `get_revision_timeline_ref(...)` | +| Data loading | adapter-backed `DataContext.fetch/load/fetch_many` | `DataContext.sources`, `fetch_panel(...)`, raw `DataSource` routing | +| PIT adapter bridge | source-adapter path | `SourceAdapterPITCompat` | +| PIT ingestion strictness | `"error"`, `"warn"`, `"coerce"` | boolean `strict=True/False` | + +Those compatibility surfaces remain only to support downstream migration. +Their removal backlog is tracked in `doc/plan/post_migration_plan.md`. diff --git a/docs/guides/core-platform-migration.md b/docs/guides/core-platform-migration.md new file mode 100644 index 0000000..71444fd --- /dev/null +++ b/docs/guides/core-platform-migration.md @@ -0,0 +1,88 @@ +# Core Platform Migration Guide + +Use this guide when moving downstream code onto the canonical public surfaces +from the core platform roadmap. + +## Downstream workflows in scope + +The migration guidance is centered on the workflows that already depend on +Alphaforge: + +- nowcast-style PIT semantics +- positioning-style operational source workflows +- volatility notebook dataset assembly + +## Preferred public paths + +Move new code toward these canonical surfaces: + +| Area | Prefer this path | +| --- | --- | +| Temporal semantics | `from alphaforge.time import ...` | +| PIT ref queries | `RefSnapshotQuery`, `RefRevisionQuery`, `PITAccessor.snapshot_ref(...)`, `PITAccessor.revisions_ref(...)` | +| Adapter bootstrap | `DataContext.from_adapters(...)` | +| Common table loads | `ctx.load(...)` | +| Canonical routing | `ctx.fetch(...)`, `ctx.fetch_many(...)` | +| Research dataset assembly | `DatasetSpec`, `FeatureRequestGroup`, built-in templates | +| Source operations | `build_health_report(...)`, `discover_archive_fetches(...)`, `iter_yearly_archive_fetches(...)` | + +## Compatibility surfaces still alive during migration + +These surfaces are still available, but they should be treated as temporary +bridges rather than equal public directions: + +- `alphaforge.pit.release_rules` +- `alphaforge.pit.missingness` +- `PITAccessor.get_snapshot_ref(...)` +- `PITAccessor.get_revision_timeline_ref(...)` +- `SourceAdapterPITCompat` +- `DataContext.sources` +- `DataContext.fetch_panel(...)` +- boolean `strict=True/False` PIT ingestion + +Cleanup for those bridges is tracked in the post-migration backlog: + +- `ALP-23` +- `ALP-24` +- `ALP-25` +- `ALP-26` +- `ALP-27` + +## Recommended migration order + +1. Move imports and typed semantics first. +2. Move PIT readers onto typed query objects and canonical execution methods. +3. Move data loading onto adapter-backed `DataContext` helpers. +4. Move notebook and dataset specs onto grouped requests and built-in + templates where they fit. +5. Replace operational ad hoc logic with release-aware health reports and + archive fetch plans. +6. Only remove a legacy surface after contract coverage and downstream callers + are both clean. + +## Required checks before deprecation or removal + +Use the repo-local stability gates before tightening a contract: + +```bash +python -m pytest tests/contracts +python -m benchmarks.pit +``` + +Also rerun the targeted subsystem tests for the touched layer and update the +repo-local migration notes: + +- `doc/plan/migration_note.md` +- `doc/plan/post_migration_plan.md` when a temporary bridge is added or retired + +## Documentation expectations + +When a migration-sensitive ticket lands: + +- update the canonical docs first +- leave legacy paths documented only as temporary compatibility notes +- record the downstream impact in `doc/plan/migration_note.md` +- keep the roadmap mirror in `doc/plan/core_platform_roadmap.md` aligned with + Linear state + +That keeps downstream migrations tied to one canonical direction at a time. diff --git a/docs/guides/data-sources.md b/docs/guides/data-sources.md index 58c2b22..96d8afe 100644 --- a/docs/guides/data-sources.md +++ b/docs/guides/data-sources.md @@ -1,10 +1,21 @@ # Data Sources Guide -Alphaforge data access is mediated by `DataContext`, which wires source names to source objects. +Alphaforge data access is mediated by `DataContext`. + +For new code, the canonical public loading path is: + +- build a context with `DataContext.from_adapters(...)` +- call `DataContext.load(...)`, `fetch_many(...)`, or `prefetch(...)` + +The older `DataContext.sources` and `fetch_panel(...)` path remains available +only for backward compatibility and raw-loader flows that have not migrated to +adapters yet. ## Unified Data Layer -The unified data layer provides a single `SourceAdapter` protocol for all data sources, whether they serve PIT macro data, market OHLCV, or bulk positioning data. Key components: +The unified data layer provides a single `SourceAdapter` protocol for all data +sources, whether they serve PIT macro data, market OHLCV, or bulk positioning +data. Key components: - **`SourceAdapter`** — protocol that every adapter implements (`fetch`, `prefetch`, `list_entities`) - **`SourceAdapterBase`** — mixin with default `fetch_many` (iterates) and `prefetch` (no-op) @@ -18,32 +29,67 @@ from alphaforge.data.context import DataContext from alphaforge.data.sources.tiingo import TiingoAdapter from alphaforge.data.sources.fred import FREDSourceAdapter -ctx = DataContext( - sources={"tiingo": legacy_source}, # legacy path (backward compat) +ctx = DataContext.from_adapters( + TiingoAdapter(api_key="..."), + FREDSourceAdapter(api_key="..."), calendars={"XNYS": cal}, store=store, - adapters={ # unified path - "tiingo": TiingoAdapter(api_key="..."), - "fred": FREDSourceAdapter(api_key="..."), - }, default_sources={"market.ohlcv": "tiingo", "macro.fred": "fred"}, ) ``` +`from_adapters(...)` derives the adapter map for you and automatically sets +default sources for datasets that are served by exactly one adapter. + +If a dataset is served by more than one adapter, canonical routing now requires +either: + +- a `default_sources` entry for that dataset, or +- an explicit `source=` on the fetch/load call + +Alphaforge no longer guesses by taking the first registered adapter for an +ambiguous dataset. + ### Fetching data ```python -from alphaforge.data.query import Query - -# Unified fetch — routes to the correct adapter automatically -result = ctx.fetch(Query(table="market.ohlcv", entities=["SPY"], start=start, end=end)) +# Happy-path single-table load +result = ctx.load( + "market.ohlcv", + columns=["close", "volume"], + entities=["SPY"], + start=start, + end=end, +) result.data # DataFrame result.source # "tiingo" -# Legacy path still works -panel = ctx.fetch_panel("tiingo", query) +# Canonical batch fetch preserves input order and lets adapters optimize +results = ctx.fetch_many( + [ + Query(table="market.ohlcv", entities=["SPY"], start=start, end=end), + Query(table="macro.fred", entities=["GDP"], start=start, end=end), + ] +) ``` +`fetch_many(...)` groups queries by the resolved adapter and delegates through +the adapter batch contract, so cache-aware sources can optimize multi-query +loads without changing the caller surface. + +### Compatibility boundary + +These older patterns still work during migration, but they are not the +preferred public API for new code: + +- `ctx.sources[...]` +- `ctx.fetch_panel(...)` +- direct `DataSource` usage as the primary loading contract + +Keep using them only where a loader family has not been migrated to +`SourceAdapter` yet, or where a raw `PanelFrame` conversion is still required +internally. + ### Entry-point discovery Adapters are registered as `alphaforge.source_adapters` entry points. Third-party packages can add their own adapters by declaring an entry point in their `pyproject.toml`: @@ -85,9 +131,20 @@ Most source fetches are driven by `alphaforge.data.query.Query`, including: Some public web sources are configured through YAML registries in `alphaforge/data/registries`. +For the public-web loader pack specifically, see the +[public-web source authoring guide](public-web-source-authoring.md) for +helper-family selection, registry wiring, and validation expectations. + ## Practical recommendation -For production pipelines, keep source instantiation and registry/version pins explicit in one bootstrap module so dataset builds remain reproducible over time. +For production pipelines, keep adapter instantiation and registry/version pins +explicit in one bootstrap module so dataset builds remain reproducible over +time. + +## Operational helpers + +For source monitoring and recurring archive-backed ingestion, see +[Source Operations](source-operations.md). ## Local futures artifacts diff --git a/docs/guides/dataset-spec.md b/docs/guides/dataset-spec.md index 8ff25f4..4b83adb 100644 --- a/docs/guides/dataset-spec.md +++ b/docs/guides/dataset-spec.md @@ -2,11 +2,16 @@ `DatasetSpec` is the declarative contract for reproducible dataset builds. +For new research assembly code, treat it as the canonical dataset-building +surface. + ## Main components - `UniverseSpec`: entity universe - `TimeSpec`: start/end/calendar/grid/asof settings - `FeatureRequest`: feature template + params (+ optional slice override) +- `FeatureRequestGroup`: composable group of feature requests with inherited + tags, key prefixes, and slice overrides - `TargetRequest`: target template + params/horizon/name - `JoinPolicy`: feature join policy (`inner` or `outer`) - `MissingnessPolicy`: final row policy (`drop_if_any_nan` or `keep`) @@ -19,10 +24,134 @@ 4. Call `build_dataset(ctx, spec, persist=True)`. 5. Consume `DatasetArtifact` (`X`, `y`, `catalog`, metadata). +## Template composition + +`DatasetSpec.features` can contain either flat `FeatureRequest` objects or +nested `FeatureRequestGroup` objects. + +Composition rules are explicit: + +- group `tags` are inherited by all nested requests +- more specific tags win on key collision: + outer group -> inner group -> request +- group `slice_override` is inherited per field +- request-level `slice_override` wins per field when both are present +- group `key` prefixes nested request keys with `/` + +Example: + +```python +from alphaforge.features.dataset_spec import FeatureRequest, FeatureRequestGroup, SliceOverride + +features = [ + FeatureRequestGroup( + key="macro", + tags={"family": "macro", "recipe": "volatility"}, + slice_override=SliceOverride(lookback=pd.Timedelta(days=30)), + requests=[ + FeatureRequest( + template=CarryTemplate(), + key="carry", + tags={"series": "carry"}, + ), + FeatureRequest( + template=InflationTemplate(), + key="inflation", + tags={"series": "cpi"}, + ), + ], + ) +] +``` + +The resulting feature catalog records request-level composition metadata such +as: + +- `request_key` +- `template_name` +- `template_version` +- merged `tags_json` + ## Slice overrides Use `SliceOverride` on a per-feature/per-target basis when a request needs a different lookback, grid, or as-of value than the global spec. +When a request lives inside a `FeatureRequestGroup`, the group override is +applied first and the request override refines it. + +## Join policy + +`JoinPolicy` controls how feature families combine before final missingness +handling. + +- `inner`: keep only timestamps/entities present across all feature families +- `outer`: union feature-family rows first, then rely on the missingness policy + to decide what survives + +The builder always aligns features onto the explicit evaluation grid defined by +`TimeSpec`, so join policy governs feature-family composition, not whether the +dataset has a deterministic time/entity index. + +## Missingness policy + +- `drop_if_any_nan`: keep only final rows where every feature column and the + target are present +- `keep`: preserve the aligned dataset even when some features or target rows + are missing + +## Template behavior expectations + +Feature templates are expected to return a `FeatureFrame` whose: + +- `X` uses a `MultiIndex` of `(ts_utc, entity_id)` +- `catalog` contains one row per feature id +- output timestamps respect the requested slice semantics, especially `asof` + for PIT-sensitive templates + +The dataset builder preserves request tags, annotates request/template metadata +in the catalog, and keeps leakage detection as a best-effort warning when a +template returns timestamps beyond the requested `asof`. + +## Built-in notebook-ready templates + +Alphaforge now ships a small built-in template family for common market-price +research work: + +- `LagReturnsTemplate` +- `RollingVolatilityTemplate` + +These templates use the canonical adapter-backed loading path +(`DataContext.from_adapters(...)` plus `ctx.load(...)`) and are intended to +replace repeated notebook helper cells for lagged returns and trailing +volatility windows. + +```python +from alphaforge.features import LagReturnsTemplate, RollingVolatilityTemplate + +FeatureRequestGroup( + key="volatility", + tags={"recipe": "volatility"}, + requests=[ + FeatureRequest( + template=LagReturnsTemplate(), + key="returns", + params={"dataset": "market.ohlcv", "source": "market", "lags": [1, 5, 10]}, + ), + FeatureRequest( + template=RollingVolatilityTemplate(), + key="trailing_vol", + params={ + "dataset": "market.ohlcv", + "source": "market", + "windows": [5, 10, 21], + "lag": 1, + "annualization_factor": 252, + }, + ), + ], +) +``` + ## Output contract `build_dataset` returns a `DatasetArtifact` with: diff --git a/docs/guides/development.md b/docs/guides/development.md index d9bccaa..dccf1f3 100644 --- a/docs/guides/development.md +++ b/docs/guides/development.md @@ -15,6 +15,23 @@ mypy alphaforge pytest ``` +## Core platform regression gates + +When a change touches canonical PIT, data-context, dataset, or operational +surfaces from the roadmap, also run: + +```bash +python -m pytest tests/contracts +python -m benchmarks.pit +``` + +## Public web contributors + +For the public loader pack under `alphaforge.data.public_web`, use the +dedicated [public-web source authoring guide](public-web-source-authoring.md). +It covers helper-family selection, registry/export wiring, targeted test +expectations, and the required docs plus Linear plan updates. + ## Build package ```bash diff --git a/docs/guides/pit-api-contract.md b/docs/guides/pit-api-contract.md index 598dfcf..1868d08 100644 --- a/docs/guides/pit-api-contract.md +++ b/docs/guides/pit-api-contract.md @@ -2,6 +2,9 @@ This guide defines stable contracts for PIT ingestion, transforms, and data-source queries. +The repo-local regression gates for this contract live in +[Core Platform Contracts And Benchmarks](contracts-and-benchmarks.md). + ## Error model PIT uses typed exceptions: @@ -185,6 +188,41 @@ Error mode rejects: ## Release helper contract +## Ref query contract + +Typed ref-period PIT queries use: + +- `RefSnapshotQuery` +- `RefRevisionQuery` + +Public execution APIs: + +- `PITAccessor.snapshot_ref(query)` + - accepts `RefSnapshotQuery` or a mapping with equivalent fields + - returns a `Series` indexed by typed `RefPeriod` values + - requires explicit `freq` when both `start_ref` and `end_ref` are omitted + - supports `obs_date_anchor="start" | "end"` for series whose stored + observation dates are period-start or period-end keyed + - normalizes `freq` to canonical `RefFreq` values before execution and stores + the resolved `freq` / `obs_date_anchor` on the output `Series.attrs` +- `PITAccessor.revisions_ref(query)` + - accepts `RefRevisionQuery` or a mapping with equivalent fields + - returns a revision timeline indexed by `asof_utc` + - resolves the input ref to a canonical `RefPeriod` before execution + - names the output with the canonical ref-entity id form and stores the + resolved `RefPeriod` on `Series.attrs["ref_period"]` + +Compatibility wrappers: + +- `get_snapshot_ref(...)` +- `get_revision_timeline_ref(...)` + +remain supported during migration, but they are compatibility helpers rather +than the preferred public surface. + +The nowcast-style contract coverage for this surface lives in +`tests/contracts/test_nowcast_pit_contract.py`. + Release stream helpers for reference periods: - `list_release_stream(series_key, ref, asof=None, freq=None)` @@ -192,6 +230,50 @@ Release stream helpers for reference periods: - `resolve_release(series_key, ref, policy=..., asof=None, freq=None)` - supports policies: `"first"`, `"latest"`, `{"mode":"rank","rank":n}`, `{"mode":"horizon","horizon":...}`. +## Series explainability contract + +Persisted derived series can be inspected with: + +- `get_series_lineage(series_key, start_obs=None, end_obs=None, start_asof=None, end_asof=None, limit=...)` +- `explain_series(series_key, start_obs=None, end_obs=None, start_asof=None, end_asof=None, limit=...)` + +`get_series_lineage(...)` returns row-level provenance columns including: + +- `lineage_kind` +- `transform_id` +- `graph_id` +- `node_name` +- `input_series_keys` +- `source_asof_utc` +- `selected_input_asof_utc` +- `source_asof_by_series_utc` +- `max_source_asof_utc` +- `causality_status` + +`lineage_kind` values: + +- `raw` +- `transform` +- `expression_graph` +- `derived` for other persisted lineage payloads + +`causality_status` values: + +- `raw` +- `ok` +- `unknown` +- `violation` +- `experimental` + +`explain_series(...)` summarizes the row-level lineage into: + +- unique input series keys +- transform ids +- expression graph ids +- row counts and derived-row counts +- aggregate causality status counts +- a boolean `causality_safe` summary flag + ## Expression graph contract Expression graphs define deterministic, dependency-ordered multi-series PIT transforms. @@ -212,8 +294,29 @@ Each node applies deterministic as-of alignment using union vintages of direct i - `list_union_vintages(series_keys, start, end, mode=\"event|calendar\")` - `build_snapshot_panel(series_specs, asof, align=\"month_end|quarter_end\", join=...)` - -Snapshot panels support per-series release policies and deterministic ref alignment. +- `build_snapshot_panel_long(series_specs, asof, align=\"month_end|quarter_end\")` + +Snapshot panel semantics: + +- `get_snapshot_multi(...)` returns batch rows with: + - `series_key` + - `obs_date` + - `source_asof_utc` + - `value` +- `build_snapshot_panel_long(...)` returns aligned long rows with: + - `series_key` + - `series_alias` + - `obs_date` + - `source_obs_date` + - `source_asof_utc` + - `value` +- `build_snapshot_panel(...)` is the wide pivot over the aligned long form and + preserves explicit `join` semantics (`inner | left | right | outer`) +- `SnapshotSeriesSpec` supports per-series `release_policy` plus optional + `start_ref`, `end_ref`, `freq`, and `obs_date_anchor` for explicit ref-aware + bounds +- panel alignment is deterministic and explicit; aligned panel dates do not + discard the underlying source observation or source vintage metadata ## PIT contract versioning @@ -226,7 +329,10 @@ Migration entries for contract/validation changes are recorded in `docs/guides/p ## Data-source query contract -`PITDataSource` table semantics: +`PITDataSource` remains the legacy/raw-loader `DataSource` bridge into PIT +storage. Canonical PIT access lives on `PITAccessor`, but when a panel-style +integration still needs the `DataSource` contract, `PITDataSource` exposes +these table semantics: - `pit.snapshot` - requires `Query.asof` diff --git a/docs/guides/pit-migrations.md b/docs/guides/pit-migrations.md index 955a6d6..42f3cdc 100644 --- a/docs/guides/pit-migrations.md +++ b/docs/guides/pit-migrations.md @@ -12,6 +12,16 @@ This guide tracks PIT contract versions and required migration actions for break ## Entries +### Version 2.2.0 + +- **Date:** 2026-04-05 +- **Change Type:** non-breaking +- **Summary:** Stabilized the typed ref-period PIT query surface around `RefSnapshotQuery` / `RefRevisionQuery` and the canonical `snapshot_ref(...)` / `revisions_ref(...)` APIs, while keeping the older direct helper methods as compatibility wrappers. +- **Required Actions:** + 1. Existing `get_snapshot_ref(...)` and `get_revision_timeline_ref(...)` callers remain valid; no forced migration is required. + 2. Prefer `snapshot_ref(...)` and `revisions_ref(...)` with typed query objects or equivalent mappings for new code so `freq` and `obs_date_anchor` normalization happen in one place. + 3. Fresh environments no longer need an out-of-band `pytz` install for PIT/docs examples because timezone support is declared in package metadata. + ### Version 2.1.0 - **Date:** 2026-03-10 diff --git a/docs/guides/pit.md b/docs/guides/pit.md index a432004..f5a3937 100644 --- a/docs/guides/pit.md +++ b/docs/guides/pit.md @@ -9,6 +9,36 @@ Alphaforge PIT provides: 5. Revision/staleness helpers (`alphaforge.pit.tasks`) 6. Model-importance attribution back to dataset source fields and request tags +The semantic vocabulary for release timing and missingness is now anchored in +[`alphaforge.time`](temporal-semantics.md). PIT keeps compatibility imports for +those types, but the canonical public path is the time package. + +Ref-period semantics now route through that same layer. PIT ref helpers share +one normalization path for canonical ref keys, pandas `Period` inputs, and +explicit observation dates plus declared frequency/anchor semantics. + +## Minimal bootstrap + +The shortest supported local bootstrap is: + +```python +import pandas as pd + +from alphaforge import PITAccessor + +pit = PITAccessor.open("./alphaforge_store") +``` + +That opens the DuckDB-backed PIT store at the given root and ensures the PIT +tables exist. + +Minimal snapshot and revision-history reads then look like: + +```python +snap = pit.get_snapshot("GDP", pd.Timestamp("2025-03-01", tz="UTC")) +timeline = pit.get_revision_timeline("GDP", pd.Timestamp("2024-12-31", tz="UTC")) +``` + ## Canonical PIT schema | Column | Type | Notes | @@ -53,6 +83,67 @@ record = pit.resolve_release( ) ``` +Equivalent typed ref input also works: + +```python +stream = pit.list_release_stream("GDP", pd.Period("2024Q4", freq="Q")) +``` + +## First-class ref-period queries + +`PITAccessor` now exposes typed ref-period query objects for the common +snapshot and revision workflows: + +```python +from alphaforge import PITAccessor, RefFreq, RefRevisionQuery, RefSnapshotQuery + +snap = pit.snapshot_ref( + RefSnapshotQuery( + series_key="GDP", + asof=pd.Timestamp("2025-06-01", tz="UTC"), + start_ref="2024Q4", + end_ref="2025Q1", + ) +) + +revisions = pit.revisions_ref( + RefRevisionQuery( + series_key="GDP", + ref="2024Q4", + end_asof=pd.Timestamp("2025-03-01", tz="UTC"), + ) +) +``` + +Key semantics: + +- `snapshot_ref(...)` returns a `Series` indexed by typed `RefPeriod` values + instead of raw observation timestamps. +- `revisions_ref(...)` keeps an `asof_utc` index and names the output with the + canonical ref-entity id form, for example `GDP|2024Q4`. +- both query types accept canonical ref keys, pandas `Period` objects, or + explicit observation dates plus declared `freq` and `obs_date_anchor` + semantics. + +Example with start-anchored monthly observations: + +```python +snap = pit.snapshot_ref( + RefSnapshotQuery( + series_key="CPI", + asof=pd.Timestamp("2025-03-01", tz="UTC"), + start_ref="2025-01-01", + end_ref="2025-02-01", + freq=RefFreq.M, + obs_date_anchor="start", + ) +) +``` + +The older `get_snapshot_ref(...)` and `get_revision_timeline_ref(...)` methods +remain available during migration, but the query-object APIs above are the +preferred public path. + ## Transform API ### PIT operator capability map @@ -382,8 +473,84 @@ panel = pit.build_snapshot_panel( align=\"month_end\", join=\"outer\", ) + +panel_long = pit.build_snapshot_panel_long( + [ + {"series_key": "GDP", "alias": "gdp"}, + {"series_key": "CPI", "alias": "cpi", "release_policy": "latest"}, + ], + asof=pd.Timestamp("2025-06-30", tz="UTC"), + align="month_end", +) +``` + +`get_snapshot_multi(...)` now returns `source_asof_utc` alongside the snapshot +values, and `build_snapshot_panel_long(...)` preserves both: + +- aligned `obs_date` +- original `source_obs_date` +- `source_asof_utc` + +This makes downstream panel assembly explicitly causal instead of forcing +repo-local loops to reconstruct which vintage supplied each panel row. + +`SnapshotSeriesSpec` also accepts `freq` and `obs_date_anchor` so ref-period +bounded panel requests can express period-start keyed series without ad hoc +timestamp normalization: + +```python +panel_long = pit.build_snapshot_panel_long( + [ + { + "series_key": "CPI", + "alias": "cpi", + "start_ref": "2025-01-01", + "end_ref": "2025-02-01", + "freq": "M", + "obs_date_anchor": "start", + } + ], + asof=pd.Timestamp("2025-03-01", tz="UTC"), +) +``` + +## Lineage and causal diagnostics + +Derived PIT series can now be inspected directly from storage without manually +parsing `meta_json` payloads: + +```python +lineage = pit.get_series_lineage("GDP_MINUS_CPI") +summary = pit.explain_series("GDP_MINUS_CPI") ``` +`get_series_lineage(...)` returns row-level provenance with fields such as: + +- `lineage_kind` +- `transform_id` +- `graph_id` +- `input_series_keys` +- `source_asof_utc` +- `source_asof_by_series_utc` +- `max_source_asof_utc` +- `causality_status` + +`explain_series(...)` rolls that up into a series-level summary: + +- which transforms or expression graphs produced the series +- which input series keys were used +- whether the persisted lineage is causally safe under the stored `asof_utc` + semantics +- whether the lineage is raw, transform-derived, or expression-graph-derived + +Current `causality_status` values are: + +- `raw` +- `ok` +- `unknown` +- `violation` +- `experimental` + ## Engine contract - `engine="auto"` prefers `duckdb` for supported op+axis+params combinations. @@ -411,6 +578,11 @@ Available in `alphaforge.pit.tasks`: ## PITDataSource integration +`PITDataSource` is the legacy/raw-loader bridge for exposing PIT storage +through the `DataSource` contract. Prefer `PITAccessor` directly for canonical +PIT querying; use `PITDataSource` when an older panel-oriented integration +still expects the `DataSource` shape. + `PITDataSource` exposes two tables: - `pit.snapshot`: requires `Query.asof`, currently supports only `vintage="latest"` diff --git a/docs/guides/public-web-source-authoring.md b/docs/guides/public-web-source-authoring.md new file mode 100644 index 0000000..d4a7a61 --- /dev/null +++ b/docs/guides/public-web-source-authoring.md @@ -0,0 +1,90 @@ +# Public Web Source Authoring + +Use `alphaforge.data.public_web` for public datasets that should expose a +stable `Query`-driven interface but do not need a separate adapter package. + +The refactor goal is modest: centralize fetch/finalize boilerplate while +keeping provider-specific normalization logic explicit inside each source +module. + +## Choose the shallowest helper family + +| Loader shape | Use when | Preferred helpers | Current examples | +| --- | --- | --- | --- | +| Registry-backed API | entity definitions come from YAML registry metadata | `PublicWebSourceBase`, `RegistryApiSourceBase`, `schema_helpers.py` | `bcb_sgs.py`, `bea.py`, `bls.py`, `destatis_genesis.py`, `ecb_sdmx.py`, `eia.py`, `eurostat.py`, `ibge_sidra.py` | +| Tabular document | provider publishes HTML, XLSX, or CSV documents with recurring table cleanup patterns | `PublicWebSourceBase`, `TabularDocumentSourceBase`, `tabular.py`, `schema_helpers.py` | `cme_productslate_reference.py`, `ec_weekly_oil_bulletin.py`, `eurex_refdata_contracts.py`, `eurex_stats_daily.py`, `ezoic_adrevenue_daily.py`, `frb_term_structure.py`, `lch_cdsclear_daily.py` | +| Archive or batch source | data is distributed as yearly ZIPs or historical archive bundles | `PublicWebSourceBase`, `archive.py`, `schema_helpers.py` | `b3_historical_quotes.py`, `cftc_cot.py`, `cftc_swaps_weekly.py` | +| True outlier | workflow is too source-specific for a family helper | `PublicWebSourceBase` only | `dtcc_ppd.py`, `mof_jgb.py`, `philadelphia_spf.py` | + +If a source only shares HTTP setup and finalization, stop at +`PublicWebSourceBase`. Do not force it into a deeper family abstraction. + +## Required implementation checklist + +1. Read the target source, its matching test file, and any shared helper + modules it already depends on. +2. Preserve table names, schema contracts, entity-id shapes, sorting, and + `asof_utc` semantics unless the ticket explicitly changes them. +3. Define schemas with `table_schema()`, `daily_panel_schema()`, + `single_value_schema()`, or `event_table_schema()` rather than open-coding + `TableSchema` where the helper fits. +4. Use `_empty_frame()` and `_finalize()` so projection, entity filtering, + date filtering, and `asof_utc` handling stay consistent. +5. Update `alphaforge/data/public_web/registry.py` and + `alphaforge/data/public_web/__init__.py` when adding a new default source. +6. If the source is intended to be available from the package root, also update + `alphaforge/__init__.py`. +7. Add or update the matching tests in `tests/public_web/` and adapter tests in + `tests/` when higher-level routing changes. +8. Update docs after the implementation is stable: + `docs/api/`, `docs/getting-started/`, `docs/guides/`, and the mirrored plan + file in `doc/plan/` when ticket state or scope changed. + +## Defensive parsing rules + +- Expect provider archives to have header drift, renamed files, and historical + batch exceptions. +- Prefer `discover_archive_fetches(...)` or `iter_yearly_archive_fetches(...)` + over open-coded URL lists when a source repeatedly downloads archive files. + They keep deterministic artifact names and handle query-string download links + more robustly than ad hoc string filtering. +- Normalize date columns through shared helpers such as `ensure_date_utc()`, + `ensure_utc()`, `resolved_date_series()`, and `_asof_utc()`. +- Treat empty upstream payloads as a valid case and return schema-correct empty + frames instead of ad hoc partial frames. +- Keep artifact naming deterministic when caching HTTP responses. +- Do not invent entity ids, table names, or aliases. Verify them from the + provider payload, registry, or existing tests. + +## Validation + +Targeted source tests should fail first for the intended reason, then pass after +the implementation lands. + +Recommended validation commands: + +```bash +/Users/steveyang/miniforge3/bin/python -m pytest tests/public_web/test_.py -q +/Users/steveyang/miniforge3/bin/python -m pytest tests/public_web -k 'not live_sources' -q +/Users/steveyang/miniforge3/bin/python -m ruff check . +/Users/steveyang/miniforge3/bin/python -m mkdocs build --strict +``` + +Run adapter regressions in `tests/` when the source is reachable through +`alphaforge.data.sources` or a PIT transform. + +## Registry and discovery + +- `default_public_web_sources()` is the default constructor registry for the + public-web pack. +- `tests/public_web/test_source_test_mapping.py` enforces that every concrete + source module has a matching test module. +- `tests/public_web/test_live_sources.py` is opt-in and should stay resilient + to provider outages and empty live responses. + +## Coordination + +Follow the implementation workflow in `AGENTS.md` for Linear-driven work: +review upstream tickets first, announce the active ticket on screen, implement +with TDD, update docs, leave a Linear close-out note, mark the ticket done, and +only then sync the mirrored plan table. diff --git a/docs/guides/research-recipes.md b/docs/guides/research-recipes.md new file mode 100644 index 0000000..112ae3e --- /dev/null +++ b/docs/guides/research-recipes.md @@ -0,0 +1,66 @@ +# Research Recipes + +This guide records notebook-shaped patterns that Alphaforge should support +without bespoke helper layers. + +## Volatility recipe + +The canonical volatility workflow is: + +1. build a `DataContext` from adapters with `DataContext.from_adapters(...)` +2. group reusable features with `FeatureRequestGroup` +3. use built-in market templates for lagged returns and trailing volatility +4. keep custom code focused on the target definition rather than feature + plumbing + +The built-in feature family now covers the common market-price pieces: + +- `LagReturnsTemplate` for lagged simple or log returns +- `RollingVolatilityTemplate` for trailing realized volatility windows + +Example feature group: + +```python +from alphaforge.features import LagReturnsTemplate, RollingVolatilityTemplate +from alphaforge.features.dataset_spec import FeatureRequest, FeatureRequestGroup + +features = [ + FeatureRequestGroup( + key="volatility", + tags={"recipe": "volatility"}, + requests=[ + FeatureRequest( + template=LagReturnsTemplate(), + key="returns", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "lags": [1, 5, 10], + }, + ), + FeatureRequest( + template=RollingVolatilityTemplate(), + key="trailing_vol", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "windows": [5, 10, 21], + "lag": 1, + "annualization_factor": 252, + }, + ), + ], + ) +] +``` + +This removes the repeated notebook work of: + +- fetching prices manually for each feature family +- writing ad hoc lag-return helper cells +- rebuilding the trailing-volatility calculation in each notebook +- hand-annotating request metadata in the catalog + +For a full runnable example, see `examples/volatility_dataset_recipe.py`. diff --git a/docs/guides/source-operations.md b/docs/guides/source-operations.md new file mode 100644 index 0000000..774ca63 --- /dev/null +++ b/docs/guides/source-operations.md @@ -0,0 +1,92 @@ +# Source Operations + +This guide covers the shared operational surfaces for source health monitoring +and recurring archive-backed ingestion. + +## Release-aware health reports + +Use `SourceHealthPolicy` and `assess_source_health(...)` when you need a +source-level health decision at a specific `asof`. + +When a source has a typed `release_rule`, Alphaforge now exposes structured +release-aware diagnostics: + +- `expected_next` +- `overdue` +- `overdue_days` +- `status` +- `weight_factor` + +Turn a set of health statuses into a report frame with +`build_health_report(...)`: + +```python +from alphaforge.pipeline.health import ( + SourceHealthPolicy, + assess_source_health, + build_health_report, +) +from alphaforge.time import FixedLagMonths + +policy = SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=31), + release_rule=FixedLagMonths(months=2), +) + +status = assess_source_health( + "macro", + latest_obs_date=pd.Timestamp("2025-01-31", tz="UTC"), + asof=pd.Timestamp("2025-04-05", tz="UTC"), + policy=policy, +) + +report = build_health_report({"macro": status}) +``` + +If you are already recording health through `SourceHealthTracker`, use +`tracker.report(asof)` to get the same dataframe surface from the configured +tracker. + +## Archive-backed ingestion planning + +Recurring public-web ingestion flows often need more than raw URLs. They need a +deterministic fetch plan with artifact names that can be reused in cache and +debug logs. + +`alphaforge.data.public_web.archive` now exposes planned fetch entries via +`ArchiveFetchPlanEntry` and helper builders: + +- `discover_archive_fetches(...)` +- `iter_yearly_archive_fetches(...)` + +These helpers: + +- keep deterministic `artifact_name` values +- preserve the resolved fetch `url` +- infer a `year` when it is present in the archive path +- handle query-string download links during archive discovery + +Example: + +```python +from alphaforge.data.public_web.archive import discover_archive_fetches + +planned = discover_archive_fetches( + html, + base_url="https://example.com/archive/index.html", + suffixes=[".zip", ".csv"], + years=[2024, 2025], + fallback_artifact_prefix="cftc_swaps_weekly", +) + +for entry in planned: + print(entry.url, entry.artifact_name, entry.year) +``` + +The archive-backed public-web loaders use this plan layer so recurring +ingestion flows and cache artifact names stay deterministic across runs. + +For CFTC archive-backed loaders specifically, broken archive downloads or ZIP +parse failures now fail fast instead of being silently skipped. That keeps +historical gaps observable and prevents partial year ranges from looking like a +successful empty or truncated fetch. diff --git a/docs/guides/temporal-semantics.md b/docs/guides/temporal-semantics.md new file mode 100644 index 0000000..836dd67 --- /dev/null +++ b/docs/guides/temporal-semantics.md @@ -0,0 +1,132 @@ +# Temporal Semantics + +Alphaforge's temporal core distinguishes four different concepts that are easy +to blur together in messy real-world data: + +| Concept | Meaning | Typical field | +| --- | --- | --- | +| Observation time | When the measured phenomenon happened | `obs_date` | +| Reference period | The labeled period the value summarizes | `RefPeriod` | +| Release date | When a source is expected or observed to publish the value | `release_time_utc`, `ReleaseRule` | +| Availability / as-of | The information set cutoff used for evaluation | `asof_utc` | + +These are related, but they are not interchangeable. A March observation can +belong to the `2025Q1` reference period, be released in May, and still be +unavailable for an April `asof`. + +## Reference periods + +The canonical typed surface for reference periods is +[`alphaforge.time.ref_period`](../api/time-ref-period.md): + +```python +from alphaforge.time import RefFreq, RefPeriod, coerce_ref_period +``` + +Use `RefPeriod` when the semantic object is a labeled period rather than a raw +timestamp: + +- `2025Q1` +- `2025-02` +- `2025` + +The library now standardizes three related operations: + +- parsing canonical ref keys such as `2025Q1` +- normalizing pandas periods and explicit observation dates through one typed path +- resolving the observation-date anchor explicitly as `"start"` or `"end"` + +```python +import pandas as pd +from alphaforge.time import RefFreq, RefPeriod + +quarter = RefPeriod.parse("2025Q1") +quarter.to_key() # "2025Q1" +quarter.obs_date(anchor="end") # 2025-03-31 UTC +quarter.obs_date(anchor="start") # 2025-01-01 UTC + +RefPeriod.parse(pd.Period("2025Q1", freq="Q")) +RefPeriod.parse("2025-03-31", freq=RefFreq.Q, obs_date_anchor="end") +RefPeriod.parse("2025-01-01", freq=RefFreq.Q, obs_date_anchor="start") +``` + +That keeps the normalization rule explicit: if you are converting an +observation date into a reference period, you must also say which frequency and +anchor semantics you mean. + +## Canonical release-rule API + +Release schedules now live in [`alphaforge.time`](../api/pit-release-rules.md), +not just under the PIT package: + +```python +from alphaforge.time import FixedLagMonths, QuarterlyRelease, ReleaseRule +``` + +Use release rules for expected-publication semantics: + +- catalog metadata +- release-aware missingness classification +- source health policies when cadence alone is too coarse + +For weekly schedules, `WeeklyRelease` now combines: + +- the configured lag anchor (`lag_days`), and +- the configured publication weekday (`release_weekday`) + +The expected release date is the first configured weekday on or after the lag +anchor, rather than a raw day offset that ignores the weekday field. + +Release rules are expectation objects, not a replacement for realized PIT +timestamps. If source data gives an observed `release_time_utc` or `asof_utc`, +that realized timing wins. + +## Missingness taxonomy + +The canonical missingness classifier also lives in +[`alphaforge.time`](../api/pit-missingness.md): + +```python +from alphaforge.time import MissingnessReason, classify_missingness +``` + +The current taxonomy is: + +- `STRUCTURAL`: the panel grid asks for a value that should not exist there + (for example, a quarterly series on a non-quarter-end monthly slot) +- `FUTURE`: the observation period has not ended yet +- `RAGGED_EDGE`: the value is not expected to have been released yet +- `TRUE_MISSING`: the value should have been available but is absent + +This keeps PIT panel logic, release-aware reasoning, and downstream data +quality checks on the same vocabulary. + +## Health semantics + +`alphaforge.pipeline.health.SourceHealthPolicy` can now use a `release_rule` to +evaluate staleness against the next expected release date instead of only the +time since the latest observation date. + +This matters for sources where: + +- observation cadence and publication cadence differ materially +- long release lags would make cadence-only health checks look stale too early +- monthly or quarterly data should stay `ok` until the next scheduled release + window actually passes + +## Compatibility + +The preferred imports are now: + +- `alphaforge.time.release_rules` +- `alphaforge.time.missingness` +- `alphaforge.time.ref_period` +- `from alphaforge.time import ...` + +Compatibility shims remain in place for: + +- `alphaforge.pit.release_rules` +- `alphaforge.pit.missingness` + +Those PIT paths still work, but they are now legacy entry points rather than +the canonical temporal-semantics surface. diff --git a/examples/first_rate_cross_asset_5m.py b/examples/first_rate_cross_asset_5m.py index 04097ce..42d09fa 100644 --- a/examples/first_rate_cross_asset_5m.py +++ b/examples/first_rate_cross_asset_5m.py @@ -2,8 +2,8 @@ from __future__ import annotations -from pathlib import Path import sys +from pathlib import Path REPO_ROOT = Path(__file__).resolve().parents[1] if str(REPO_ROOT) not in sys.path: @@ -13,7 +13,6 @@ from alphaforge import ( # noqa: E402 FirstRateBarsConfig, - Query, build_first_rate_bars_context, ) @@ -24,34 +23,28 @@ def main() -> None: cfg = FirstRateBarsConfig.from_base_dir(data_root) ctx = build_first_rate_bars_context(cfg) - fx = ctx.fetch( - Query( - table="fx.contract_price_5m", - columns=["bar_start_utc", "close", "volume"], - entities=["AUDUSD"], - start=pd.Timestamp("2024-01-01T00:00:00Z"), - end=pd.Timestamp("2024-01-03T23:59:59Z"), - ) + fx = ctx.load( + "fx.contract_price_5m", + columns=["bar_start_utc", "close", "volume"], + entities=["AUDUSD"], + start=pd.Timestamp("2024-01-01T00:00:00Z"), + end=pd.Timestamp("2024-01-03T23:59:59Z"), ).data.sort_values(["series_key", "obs_date"]) - crypto = ctx.fetch( - Query( - table="crypto.contract_price_5m", - columns=["bar_start_utc", "close", "volume"], - entities=["BTC"], - start=pd.Timestamp("2024-01-01T00:00:00Z"), - end=pd.Timestamp("2024-01-03T23:59:59Z"), - ) + crypto = ctx.load( + "crypto.contract_price_5m", + columns=["bar_start_utc", "close", "volume"], + entities=["BTC"], + start=pd.Timestamp("2024-01-01T00:00:00Z"), + end=pd.Timestamp("2024-01-03T23:59:59Z"), ).data.sort_values(["series_key", "obs_date"]) - index = ctx.fetch( - Query( - table="index.level_5m", - columns=["bar_start_utc", "close"], - entities=["DAX"], - start=pd.Timestamp("2024-01-01T00:00:00Z"), - end=pd.Timestamp("2024-01-03T23:59:59Z"), - ) + index = ctx.load( + "index.level_5m", + columns=["bar_start_utc", "close"], + entities=["DAX"], + start=pd.Timestamp("2024-01-01T00:00:00Z"), + end=pd.Timestamp("2024-01-03T23:59:59Z"), ).data.sort_values(["series_key", "obs_date"]) print("FX 5-minute bars") diff --git a/examples/first_rate_futures_continuous.py b/examples/first_rate_futures_continuous.py index 0b6ab20..7716649 100644 --- a/examples/first_rate_futures_continuous.py +++ b/examples/first_rate_futures_continuous.py @@ -2,8 +2,8 @@ from __future__ import annotations -from pathlib import Path import sys +from pathlib import Path import pandas as pd @@ -14,7 +14,6 @@ from alphaforge import ( # noqa: E402 FirstRateFuturesConfig, FirstRateFuturesLoader, - Query, build_first_rate_futures_context, ) @@ -28,24 +27,20 @@ def main() -> None: ctx = build_first_rate_futures_context(cfg) - eod = ctx.fetch( - Query( - table="futures.continuous_eod_research", - columns=["open", "high", "low", "close", "volume", "active_contract_id"], - entities=["CL", "BZ"], # WTI and Brent - start=pd.Timestamp("2024-01-01T00:00:00Z"), - end=pd.Timestamp("2024-03-31T23:59:59Z"), - ) + eod = ctx.load( + "futures.continuous_eod_research", + columns=["open", "high", "low", "close", "volume", "active_contract_id"], + entities=["CL", "BZ"], # WTI and Brent + start=pd.Timestamp("2024-01-01T00:00:00Z"), + end=pd.Timestamp("2024-03-31T23:59:59Z"), ).data.sort_values(["series_key", "obs_date"]) - intraday = ctx.fetch( - Query( - table="futures.continuous_5m_execution", - columns=["close", "volume", "active_contract_id", "roll_flag"], - entities=["CL", "BZ"], - start=pd.Timestamp("2024-03-01T00:00:00Z"), - end=pd.Timestamp("2024-03-05T23:59:59Z"), - ) + intraday = ctx.load( + "futures.continuous_5m_execution", + columns=["close", "volume", "active_contract_id", "roll_flag"], + entities=["CL", "BZ"], + start=pd.Timestamp("2024-03-01T00:00:00Z"), + end=pd.Timestamp("2024-03-05T23:59:59Z"), ).data.sort_values(["series_key", "obs_date"]) print("Continuous EOD research series") diff --git a/examples/pit_splice_importance.py b/examples/pit_splice_importance.py index 0a82453..226a3ea 100644 --- a/examples/pit_splice_importance.py +++ b/examples/pit_splice_importance.py @@ -20,7 +20,6 @@ from alphaforge.features.template import SliceSpec from alphaforge.pit import PITAccessor from alphaforge.pit.transforms import PITTransformSpec -from alphaforge.store.duckdb_parquet import DuckDBParquetStore from alphaforge.time.calendar import TradingCalendar @@ -87,8 +86,7 @@ def run_example(root: str | Path) -> dict[str, Any]: root_path = Path(root) root_path.mkdir(parents=True, exist_ok=True) - store = DuckDBParquetStore(root=str(root_path)) - pit = PITAccessor(store.conn()) + pit = PITAccessor.open(root_path) pit.upsert_pit_observations( pd.DataFrame( { diff --git a/examples/public_web_macro_derivs.py b/examples/public_web_macro_derivs.py index 350034f..2a7f6b2 100644 --- a/examples/public_web_macro_derivs.py +++ b/examples/public_web_macro_derivs.py @@ -64,6 +64,7 @@ def main() -> None: calendars={}, store=LocalParquetStore(str(repo_root / "alphaforge_demo_store/public_web")), ) + # This example exercises the public-web raw-loader/DataSource path directly. q_dtcc = Query( table="dtcc.ppd.daily", diff --git a/examples/short_rate_kim_dataset.py b/examples/short_rate_kim_dataset.py index a719db3..b9d3f29 100644 --- a/examples/short_rate_kim_dataset.py +++ b/examples/short_rate_kim_dataset.py @@ -1,4 +1,4 @@ -"""Build the reduced Kim-Orphanides / Kim-Wright dataset with alphaforge.""" +"""Build the reduced Kim-Orphanides / Kim-Wright dataset with legacy loaders.""" from __future__ import annotations @@ -23,6 +23,7 @@ def main() -> None: if not fred_api_key: raise RuntimeError("Set FRED_API_KEY before running this example.") + # This dataset helper still consumes the legacy/raw-loader DataSource shape. ctx = DataContext( sources={ "fred": FREDDataSource(api_key=fred_api_key), diff --git a/examples/volatility_dataset_recipe.py b/examples/volatility_dataset_recipe.py new file mode 100644 index 0000000..b7fdc32 --- /dev/null +++ b/examples/volatility_dataset_recipe.py @@ -0,0 +1,197 @@ +"""Notebook-shaped volatility dataset recipe using canonical adapter loading.""" + +from __future__ import annotations + +import sys +from pathlib import Path + +import numpy as np +import pandas as pd + +REPO_ROOT = Path(__file__).resolve().parents[1] +if str(REPO_ROOT) not in sys.path: + sys.path.insert(0, str(REPO_ROOT)) + +from alphaforge import ( # noqa: E402 + DataContext, + LagReturnsTemplate, + RollingVolatilityTemplate, +) +from alphaforge.data.adapter import SourceAdapterBase # noqa: E402 +from alphaforge.data.query import Query # noqa: E402 +from alphaforge.data.types import FetchResult # noqa: E402 +from alphaforge.features.dataset_builder import build_dataset # noqa: E402 +from alphaforge.features.dataset_spec import ( # noqa: E402 + DatasetSpec, + FeatureRequest, + FeatureRequestGroup, + JoinPolicy, + MissingnessPolicy, + TargetRequest, + TimeSpec, + UniverseSpec, +) +from alphaforge.features.target_template import TargetFrame # noqa: E402 +from alphaforge.features.template import SliceSpec # noqa: E402 +from alphaforge.time.calendar import TradingCalendar # noqa: E402 + + +class InMemoryMarketAdapter(SourceAdapterBase): + source_name = "market" + datasets = frozenset({"market.ohlcv"}) + + def __init__(self, frame: pd.DataFrame) -> None: + self._frame = frame.copy() + + def fetch(self, query: Query, *, max_staleness=None) -> FetchResult: + frame = self._frame.copy() + if query.entities is not None: + frame = frame[frame["series_key"].isin(query.entities)] + + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.start is not None: + frame = frame[obs >= query.start] + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.end is not None: + frame = frame[obs <= query.end] + + keep = ["series_key", "obs_date"] + [ + column for column in query.columns if column in frame.columns + ] + return FetchResult( + data=frame[keep].reset_index(drop=True), + source=self.source_name, + dataset=query.table, + is_pit=False, + cached_at=None, + ) + + def list_entities(self, dataset: str) -> list[str]: + return sorted(self._frame["series_key"].unique()) + + +class NextDaySquaredLogReturnTarget: + name = "next_day_squared_log_return" + version = "1.0" + param_space = {} + + def fit(self, ctx, params, fit_slice): + return None + + def transform(self, ctx, params, slice: SliceSpec, state): + result = ctx.load( + "market.ohlcv", + columns=["close"], + start=slice.start, + end=slice.end, + entities=slice.entities, + asof=slice.asof, + grid=slice.grid, + source="market", + ) + frame = result.data.copy() + calendar = ctx.calendars["XNYS"] + frame["ts_utc"] = [ + calendar.session_close_utc(ts) + for ts in pd.to_datetime(frame["obs_date"], utc=True) + ] + prices = ( + frame.set_index(["ts_utc", "series_key"])["close"] + .rename_axis(index=["ts_utc", "entity_id"]) + .sort_index() + .astype(float) + ) + logret = np.log(prices).groupby(level="entity_id").diff() + target = (logret.groupby(level="entity_id").shift(-1) ** 2).rename("target") + return TargetFrame( + y=target, + meta={"definition": "next-period squared log return"}, + ) + + +def _market_frame() -> pd.DataFrame: + calendar = TradingCalendar("XNYS", tz="UTC") + dates = calendar.sessions("2024-01-02", "2024-02-16") + rng = np.random.default_rng(7) + rows: list[dict[str, object]] = [] + for entity in ["AAA", "BBB"]: + prices = 100.0 + np.cumsum(rng.normal(0.0, 1.0, size=len(dates))) + for obs_date, close in zip(dates, prices, strict=True): + rows.append( + { + "series_key": entity, + "obs_date": obs_date, + "close": float(close), + "volume": float(rng.integers(900, 1100)), + } + ) + return pd.DataFrame(rows) + + +def main() -> None: + calendar = TradingCalendar("XNYS", tz="UTC") + ctx = DataContext.from_adapters( + InMemoryMarketAdapter(_market_frame()), + calendars={"XNYS": calendar}, + store=None, + ) + + features = [ + FeatureRequestGroup( + key="volatility", + tags={"recipe": "volatility", "asset_class": "equity"}, + requests=[ + FeatureRequest( + template=LagReturnsTemplate(), + key="returns", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "lags": [1, 5, 10], + }, + ), + FeatureRequest( + template=RollingVolatilityTemplate(), + key="trailing_vol", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "windows": [5, 10], + "lag": 1, + "annualization_factor": 252, + }, + ), + ], + ) + ] + + spec = DatasetSpec( + universe=UniverseSpec(entities=["AAA", "BBB"]), + time=TimeSpec( + start=pd.Timestamp("2024-01-02T00:00:00Z"), + end=pd.Timestamp("2024-02-16T00:00:00Z"), + calendar="XNYS", + grid="B", + ), + features=features, + target=TargetRequest( + template=NextDaySquaredLogReturnTarget(), + horizon=1, + name="next_day_sq_logret", + ), + join_policy=JoinPolicy(how="inner", sort_index=True), + missingness=MissingnessPolicy(final_row_policy="drop_if_any_nan"), + name="volatility_recipe", + ) + + artifact = build_dataset(ctx, spec, persist=False) + print("X shape:", artifact.X.shape) + print("y non-null:", int(artifact.y.notna().sum())) + print("Catalog columns:", sorted(artifact.catalog.columns.tolist())) + print(artifact.catalog[["request_key", "template_name", "group_path"]].head()) + + +if __name__ == "__main__": + main() diff --git a/mkdocs.yml b/mkdocs.yml index 415bd99..163e6f3 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -39,12 +39,19 @@ nav: - "Quickstart: Dataset": getting-started/quickstart-dataset.md - "Quickstart: Public Web": getting-started/quickstart-public-web.md - Guides: + - Core Platform Architecture: guides/core-platform-architecture.md + - Core Platform Migration: guides/core-platform-migration.md + - Core Platform Contracts And Benchmarks: guides/contracts-and-benchmarks.md - Dataset Spec: guides/dataset-spec.md - Data Sources: guides/data-sources.md + - Source Operations: guides/source-operations.md - First Rate Futures: guides/first-rate-futures.md - PIT Data: guides/pit.md + - Research Recipes: guides/research-recipes.md + - Temporal Semantics: guides/temporal-semantics.md - PIT API Contract: guides/pit-api-contract.md - PIT Migrations: guides/pit-migrations.md + - Public Web Source Authoring: guides/public-web-source-authoring.md - Development: guides/development.md - API: - Package: api/package.md @@ -54,12 +61,15 @@ nav: - Data Transforms: api/transforms.md - Dataset Builder: api/dataset-builder.md - Dataset Spec: api/dataset-spec.md + - Market Templates: api/market-templates.md - PIT Accessor: api/pit-accessor.md - - PIT Release Rules: api/pit-release-rules.md + - PIT Ref Queries: api/pit-queries.md + - Time Release Rules: api/pit-release-rules.md + - Time Ref Periods: api/time-ref-period.md - PIT Vintage Views: api/pit-views.md - PIT Vintage Resolvers: api/pit-resolvers.md - PIT Panel Builder: api/pit-panel.md - - PIT Missingness: api/pit-missingness.md + - Time Missingness: api/pit-missingness.md - PIT Vintage Utilities: api/pit-vintage.md - PIT Transforms: api/pit-transforms.md - PIT Pipelines: api/pit-pipelines.md diff --git a/pyproject.toml b/pyproject.toml index eaceaa4..c76eb8d 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -15,6 +15,7 @@ dependencies = [ "numpy>=1.23", "openpyxl>=3.1", "pandas>=2.0", + "pytz>=2024.1", "python-dateutil>=2.9.0.post0", "pyarrow>=14.0.1", "fredapi>=0.5.2", @@ -33,6 +34,7 @@ dev = [ "ruff>=0.4.0", "twine>=5.0.0", "types-python-dateutil>=2.9.0.20240316", + "types-pytz>=2025.2.0.20250516", "types-requests>=2.31.0.20240406", "types-PyYAML>=6.0.12.20250915", ] @@ -59,7 +61,7 @@ include-package-data = true exclude-package-data = { "*" = ["alphaforge_demo_store/**"] } [tool.setuptools.packages.find] -include = ["alphaforge", "alphaforge.*"] +include = ["alphaforge", "alphaforge.*", "benchmarks", "benchmarks.*"] [tool.setuptools.package-data] alphaforge = ["**/*.yaml"] diff --git a/tests/contracts/test_nowcast_pit_contract.py b/tests/contracts/test_nowcast_pit_contract.py new file mode 100644 index 0000000..9b9b643 --- /dev/null +++ b/tests/contracts/test_nowcast_pit_contract.py @@ -0,0 +1,76 @@ +from __future__ import annotations + +import pandas as pd + +from alphaforge import PITAccessor, RefSnapshotQuery +from alphaforge.store.duckdb_parquet import DuckDBParquetStore +from alphaforge.time.ref_period import RefPeriod + + +def _make_accessor(tmp_path) -> PITAccessor: + store = DuckDBParquetStore(root=str(tmp_path)) + return PITAccessor(store.conn()) + + +def _sample_nowcast_df() -> pd.DataFrame: + return pd.DataFrame( + { + "series_key": ["GDP", "GDP", "CPI", "CPI"], + "obs_date": [ + pd.Timestamp("2024-12-31"), + pd.Timestamp("2024-12-31"), + pd.Timestamp("2024-12-31"), + pd.Timestamp("2024-12-31"), + ], + "asof_utc": [ + pd.Timestamp("2025-01-10", tz="UTC"), + pd.Timestamp("2025-02-10", tz="UTC"), + pd.Timestamp("2025-01-15", tz="UTC"), + pd.Timestamp("2025-02-15", tz="UTC"), + ], + "value": [1.0, 1.1, 3.0, 3.1], + } + ) + + +def test_nowcast_style_ref_queries_and_panel_builder_contract(tmp_path) -> None: + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_nowcast_df()) + + snap = pit.snapshot_ref( + RefSnapshotQuery( + series_key="GDP", + asof=pd.Timestamp("2025-03-01", tz="UTC"), + start_ref="2024Q4", + end_ref="2024Q4", + ) + ) + panel = pit.build_snapshot_panel_long( + [ + { + "series_key": "GDP", + "alias": "gdp", + "start_ref": "2024Q4", + "end_ref": "2024Q4", + "freq": "Q", + }, + { + "series_key": "CPI", + "alias": "cpi", + "start_ref": "2024Q4", + "end_ref": "2024Q4", + "freq": "Q", + }, + ], + asof=pd.Timestamp("2025-03-01", tz="UTC"), + align="quarter_end", + ) + + assert list(snap.index) == [RefPeriod.parse("2024Q4")] + assert snap.loc[RefPeriod.parse("2024Q4")] == 1.1 + assert {"series_alias", "source_obs_date", "source_asof_utc", "value"} <= set( + panel.columns + ) + assert set(panel["series_alias"]) == {"gdp", "cpi"} + assert panel["source_asof_utc"].notna().all() + diff --git a/tests/contracts/test_operations_contract.py b/tests/contracts/test_operations_contract.py new file mode 100644 index 0000000..328ef9e --- /dev/null +++ b/tests/contracts/test_operations_contract.py @@ -0,0 +1,59 @@ +from __future__ import annotations + +import pandas as pd + +from alphaforge.data.public_web.archive import ( + discover_archive_fetches, + iter_yearly_archive_fetches, +) +from alphaforge.pipeline.health import ( + SourceHealthPolicy, + assess_source_health, + build_health_report, +) +from alphaforge.time.release_rules import FixedLagMonths + + +def test_operations_contract_for_health_and_archive_planning() -> None: + policy = SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=31), + release_rule=FixedLagMonths(months=2), + grace_period=pd.Timedelta(days=2), + stale_threshold=pd.Timedelta(days=14), + dead_threshold=pd.Timedelta(days=42), + weight_decay_half_life=pd.Timedelta(days=7), + ) + status = assess_source_health( + "macro", + latest_obs_date=pd.Timestamp("2025-01-31", tz="UTC"), + asof=pd.Timestamp("2025-04-05", tz="UTC"), + policy=policy, + ) + report = build_health_report({"macro": status}) + + planned = discover_archive_fetches( + 'report', + base_url="https://example.com/archive/index.html", + suffixes=[".zip"], + years=[2021], + fallback_artifact_prefix="example", + ) + yearly = iter_yearly_archive_fetches( + start_year=2009, + end_year=2012, + url_template="https://example.com/data_{year}.zip", + first_year=2006, + yearly_first_year=2010, + historical_url="https://example.com/historical.zip", + historical_last_year=2011, + fallback_artifact_prefix="example", + ) + + assert report.loc[0, "source_name"] == "macro" + assert report.loc[0, "expected_next"] == pd.Timestamp("2025-04-01", tz="UTC") + assert report.loc[0, "overdue_days"] == 4.0 + assert planned[0].artifact_name == "report_2021.zip" + assert planned[0].year == 2021 + assert yearly[-1].artifact_name == "data_2012.zip" + assert yearly[-1].year == 2012 + diff --git a/tests/contracts/test_pit_benchmarks.py b/tests/contracts/test_pit_benchmarks.py new file mode 100644 index 0000000..e714964 --- /dev/null +++ b/tests/contracts/test_pit_benchmarks.py @@ -0,0 +1,16 @@ +from __future__ import annotations + +from benchmarks.pit import run_pit_contract_benchmarks + + +def test_pit_benchmark_harness_returns_named_metrics() -> None: + metrics = run_pit_contract_benchmarks( + iterations=1, + periods=8, + series_count=2, + revisions_per_period=2, + ) + + assert metrics["rows_loaded"] > 0 + assert metrics["snapshot_ref_median_ms"] >= 0.0 + assert metrics["snapshot_panel_long_median_ms"] >= 0.0 diff --git a/tests/contracts/test_volatility_dataset_contract.py b/tests/contracts/test_volatility_dataset_contract.py new file mode 100644 index 0000000..e3d8a8c --- /dev/null +++ b/tests/contracts/test_volatility_dataset_contract.py @@ -0,0 +1,142 @@ +from __future__ import annotations + +import pandas as pd + +from alphaforge.data.adapter import SourceAdapterBase +from alphaforge.data.context import DataContext +from alphaforge.data.query import Query +from alphaforge.data.types import FetchResult +from alphaforge.features import LagReturnsTemplate, RollingVolatilityTemplate +from alphaforge.features.dataset_builder import build_dataset +from alphaforge.features.dataset_spec import ( + DatasetSpec, + FeatureRequest, + FeatureRequestGroup, + JoinPolicy, + MissingnessPolicy, + TargetRequest, + TimeSpec, + UniverseSpec, +) +from alphaforge.time.calendar import TradingCalendar + + +class _InMemoryMarketAdapter(SourceAdapterBase): + source_name = "market" + datasets = frozenset({"market.ohlcv"}) + + def __init__(self, frame: pd.DataFrame) -> None: + self._frame = frame.copy() + + def fetch(self, query: Query, *, max_staleness=None) -> FetchResult: + frame = self._frame.copy() + if query.entities is not None: + frame = frame[frame["series_key"].isin(query.entities)] + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.start is not None: + frame = frame[obs >= query.start] + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.end is not None: + frame = frame[obs <= query.end] + keep = ["series_key", "obs_date"] + [ + column for column in query.columns if column in frame.columns + ] + return FetchResult( + data=frame[keep].reset_index(drop=True), + source=self.source_name, + dataset=query.table, + is_pit=False, + cached_at=None, + ) + + def list_entities(self, dataset: str) -> list[str]: + return sorted(self._frame["series_key"].unique()) + + +def _market_frame() -> pd.DataFrame: + dates = pd.date_range("2024-01-02", periods=8, freq="B", tz="UTC") + rows = [] + for entity, closes in { + "AAA": [100.0, 101.0, 103.0, 102.0, 104.0, 105.0, 106.0, 108.0], + "BBB": [50.0, 49.0, 50.0, 51.0, 52.0, 53.0, 54.0, 56.0], + }.items(): + for obs_date, close in zip(dates, closes, strict=True): + rows.append({"series_key": entity, "obs_date": obs_date, "close": close}) + return pd.DataFrame(rows) + + +def test_volatility_recipe_contract_uses_canonical_short_path() -> None: + ctx = DataContext.from_adapters( + _InMemoryMarketAdapter(_market_frame()), + calendars={"XNYS": TradingCalendar("XNYS", tz="UTC")}, + store=None, + ) + + loaded = ctx.load("market.ohlcv", columns=["close"], entities=["AAA"]) + assert loaded.source == "market" + assert loaded.dataset == "market.ohlcv" + + spec = DatasetSpec( + universe=UniverseSpec(entities=["AAA", "BBB"]), + time=TimeSpec( + start=pd.Timestamp("2024-01-02T00:00:00Z"), + end=pd.Timestamp("2024-01-12T00:00:00Z"), + calendar="XNYS", + grid="B", + ), + features=[ + FeatureRequestGroup( + key="volatility", + tags={"recipe": "volatility"}, + requests=[ + FeatureRequest( + template=LagReturnsTemplate(), + key="returns", + params={ + "dataset": "market.ohlcv", + "source": "market", + "lags": [1, 2], + }, + ), + FeatureRequest( + template=RollingVolatilityTemplate(), + key="trailing_vol", + params={ + "dataset": "market.ohlcv", + "source": "market", + "windows": [3], + "lag": 1, + "annualization_factor": 252, + }, + ), + ], + ) + ], + target=TargetRequest( + template=RollingVolatilityTemplate(), + params={ + "dataset": "market.ohlcv", + "source": "market", + "windows": [3], + "lag": 0, + "annualization_factor": 252, + }, + name="realized_vol_target", + ), + join_policy=JoinPolicy(how="inner", sort_index=True), + missingness=MissingnessPolicy(final_row_policy="keep"), + name="volatility_contract", + ) + + artifact = build_dataset(ctx, spec, persist=False) + + assert not artifact.X.empty + assert set(artifact.catalog["request_key"]) == { + "volatility/returns", + "volatility/trailing_vol", + } + assert set(artifact.catalog["template_name"]) == { + "lag_returns", + "rolling_volatility", + } + diff --git a/tests/fixtures/public_web/cftc_cot/sample_disagg.csv b/tests/fixtures/public_web/cftc_cot/sample_disagg.csv new file mode 100644 index 0000000..7a1c5df --- /dev/null +++ b/tests/fixtures/public_web/cftc_cot/sample_disagg.csv @@ -0,0 +1,4 @@ +Market_and_Exchange_Names,Report_Date_as_YYYY-MM-DD,CFTC_Contract_Market_Code,Open_Interest_All,Prod_Merc_Positions_Long_All,Prod_Merc_Positions_Short_All,Change_in_Prod_Merc_Long_All,Change_in_Prod_Merc_Short_All,Swap_Positions_Long_All,Swap__Positions_Short_All,Swap__Positions_Spread_All,Change_in_Swap_Long_All,Change_in_Swap_Short_All,M_Money_Positions_Long_All,M_Money_Positions_Short_All,M_Money_Positions_Spread_All,Change_in_M_Money_Long_All,Change_in_M_Money_Short_All,Other_Rept_Positions_Long_All,Other_Rept_Positions_Short_All,Other_Rept_Positions_Spread_All,Change_in_Other_Rept_Long_All,Change_in_Other_Rept_Short_All +WHEAT-SRW - CHICAGO BOARD OF TRADE,2026-01-06,001602,480000,47000,112000,4000,-5000,93000,20000,21000,5000,-1500,118000,110000,111000,7000,-1500,27000,44000,36000,600,1600 +WHEAT-SRW - CHICAGO BOARD OF TRADE,2026-01-13,001602,485000,52000,104000,3000,-3000,88000,22000,26000,-2000,1000,111000,112000,112000,-4000,2000,26000,43000,35000,-500,800 +GOLD - COMMODITY EXCHANGE INC.,2026-01-06,088691,520000,150000,48000,8000,-2000,62000,11000,9000,2500,-500,138000,124000,21000,4500,-3000,24000,17000,7000,1200,-800 diff --git a/tests/public_web/test_cftc_cot.py b/tests/public_web/test_cftc_cot.py index d6bc1b9..da0a9f6 100644 --- a/tests/public_web/test_cftc_cot.py +++ b/tests/public_web/test_cftc_cot.py @@ -7,7 +7,11 @@ import pandas as pd import pytest -from alphaforge.data.public_web.cftc_cot import CFTCCoTSource, _publication_date +from alphaforge.data.public_web.cftc_cot import ( + CFTCCoTSource, + CFTCDisaggregatedCoTSource, + _publication_date, +) from alphaforge.data.query import Query _FIXTURE_DIR = Path(__file__).resolve().parents[1] / "fixtures/public_web/cftc_cot" @@ -21,6 +25,13 @@ def _csv_to_zip_bytes(csv_path: Path) -> bytes: return buf.getvalue() +def _csv_text_to_zip_bytes(filename: str, csv_text: str) -> bytes: + buf = io.BytesIO() + with zipfile.ZipFile(buf, "w") as zf: + zf.writestr(filename, csv_text) + return buf.getvalue() + + class _MockHttpClient: """Return pre-built ZIP bytes for any URL request.""" @@ -31,7 +42,17 @@ def get_bytes(self, *, url: str, source: str, artifact_name: str, **kw) -> bytes return self._payload -def _make_source() -> CFTCCoTSource: +class _SequenceHttpClient: + """Return a per-URL payload so tests can model partial archive failures.""" + + def __init__(self, payloads: dict[str, bytes]) -> None: + self._payloads = payloads + + def get_bytes(self, *, url: str, source: str, artifact_name: str, **kw) -> bytes: + return self._payloads[url] + + +def _make_tff_source() -> CFTCCoTSource: csv_path = _FIXTURE_DIR / "sample.csv" zip_bytes = _csv_to_zip_bytes(csv_path) http = _MockHttpClient(zip_bytes) @@ -41,9 +62,43 @@ def _make_source() -> CFTCCoTSource: ) +def _make_disagg_source() -> CFTCDisaggregatedCoTSource: + csv_path = _FIXTURE_DIR / "sample_disagg.csv" + zip_bytes = _csv_to_zip_bytes(csv_path) + http = _MockHttpClient(zip_bytes) + return CFTCDisaggregatedCoTSource( + http_client=http, + file_urls=["file:///dummy.zip"], + ) + + +def _make_tff_legacy_header_source() -> CFTCCoTSource: + csv_path = _FIXTURE_DIR / "sample.csv" + csv_text = csv_path.read_text() + csv_text = csv_text.replace( + "Report_Date_as_YYYY-MM-DD", + "Report_Date_as_MM_DD_YYYY", + ) + zip_bytes = _csv_text_to_zip_bytes("legacy_tff.csv", csv_text) + http = _MockHttpClient(zip_bytes) + return CFTCCoTSource(http_client=http, file_urls=["file:///dummy.zip"]) + + +def _make_disagg_legacy_header_source() -> CFTCDisaggregatedCoTSource: + csv_path = _FIXTURE_DIR / "sample_disagg.csv" + csv_text = csv_path.read_text() + csv_text = csv_text.replace( + "Report_Date_as_YYYY-MM-DD", + "Report_Date_as_MM_DD_YYYY", + ) + zip_bytes = _csv_text_to_zip_bytes("legacy_disagg.csv", csv_text) + http = _MockHttpClient(zip_bytes) + return CFTCDisaggregatedCoTSource(http_client=http, file_urls=["file:///dummy.zip"]) + + class TestCFTCCoTSource: def test_schemas(self) -> None: - source = _make_source() + source = _make_tff_source() schemas = source.schemas() assert "cftc.cot.tff" in schemas schema = schemas["cftc.cot.tff"] @@ -53,7 +108,7 @@ def test_schemas(self) -> None: assert "short_positions" in schema.required_columns def test_fetch_returns_rows(self) -> None: - source = _make_source() + source = _make_tff_source() df = source.fetch( Query( table="cftc.cot.tff", @@ -72,7 +127,7 @@ def test_fetch_returns_rows(self) -> None: }.issubset(df.columns) def test_entity_id_format(self) -> None: - source = _make_source() + source = _make_tff_source() df = source.fetch( Query( table="cftc.cot.tff", @@ -93,7 +148,7 @@ def test_entity_id_format(self) -> None: assert parts[-1] == "cftc" def test_entity_filter(self) -> None: - source = _make_source() + source = _make_tff_source() df = source.fetch( Query( table="cftc.cot.tff", @@ -107,7 +162,7 @@ def test_entity_filter(self) -> None: assert df["entity_id"].eq("futures.vix.lev_money.cftc").all() def test_time_filter(self) -> None: - source = _make_source() + source = _make_tff_source() # Only the first VIX row has report_date 2026-01-06 → pub_date 2026-01-09 df = source.fetch( Query( @@ -126,7 +181,7 @@ def test_time_filter(self) -> None: assert all(d <= pd.Timestamp("2026-01-12") for d in dates) def test_date_column_is_utc(self) -> None: - source = _make_source() + source = _make_tff_source() df = source.fetch( Query( table="cftc.cot.tff", @@ -139,7 +194,7 @@ def test_date_column_is_utc(self) -> None: def test_publication_date_is_friday(self) -> None: """Publication date should always be a Friday (3 bdays after Tuesday report).""" - source = _make_source() + source = _make_tff_source() df = source.fetch( Query( table="cftc.cot.tff", @@ -157,7 +212,7 @@ def test_publication_date_is_friday(self) -> None: ).all(), f"Expected all Fridays, got day_of_week={day_of_week.unique()}" def test_all_trader_categories(self) -> None: - source = _make_source() + source = _make_tff_source() df = source.fetch( Query( table="cftc.cot.tff", @@ -178,8 +233,30 @@ def test_all_trader_categories(self) -> None: assert "futures.vix.asset_mgr.cftc" in entities assert "futures.vix.other_rept.cftc" in entities + def test_legacy_report_date_header_is_parsed(self) -> None: + source = _make_tff_legacy_header_source() + df = source.fetch( + Query( + table="cftc.cot.tff", + columns=["long_positions"], + start=pd.Timestamp("2026-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + entities=["futures.vix.lev_money.cftc"], + ) + ) + assert not df.empty + assert pd.Timestamp(df["date"].iloc[0]) == pd.Timestamp("2026-01-09", tz="UTC") + + def test_year_urls_use_historical_batch_before_2010(self) -> None: + source = CFTCCoTSource(file_urls=None) + assert source._year_urls(2008, 2018) == [ + "https://www.cftc.gov/files/dea/history/fin_fut_txt_2006_2016.zip", + "https://www.cftc.gov/files/dea/history/fut_fin_txt_2017.zip", + "https://www.cftc.gov/files/dea/history/fut_fin_txt_2018.zip", + ] + def test_unknown_table_raises(self) -> None: - source = _make_source() + source = _make_tff_source() with pytest.raises(ValueError, match="Unknown table"): source.fetch( Query( @@ -190,6 +267,125 @@ def test_unknown_table_raises(self) -> None: ) ) + def test_archive_failures_raise_instead_of_silently_skipping_years(self) -> None: + csv_path = _FIXTURE_DIR / "sample.csv" + good_zip = _csv_to_zip_bytes(csv_path) + source = CFTCCoTSource( + http_client=_SequenceHttpClient( + { + "file:///broken_2025.zip": b"not-a-zip", + "file:///good_2026.zip": good_zip, + } + ), + file_urls=["file:///broken_2025.zip", "file:///good_2026.zip"], + ) + + with pytest.raises(RuntimeError, match="file:///broken_2025.zip"): + source.fetch( + Query( + table="cftc.cot.tff", + columns=["long_positions"], + start=pd.Timestamp("2025-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + ) + ) + + +class TestCFTCDisaggregatedCoTSource: + def test_schemas(self) -> None: + source = _make_disagg_source() + schemas = source.schemas() + assert "cftc.cot.disagg" in schemas + schema = schemas["cftc.cot.disagg"] + assert schema.native_freq == "W" + assert schema.time_semantics == "interval_end" + assert "long_positions" in schema.required_columns + assert "short_positions" in schema.required_columns + + def test_fetch_returns_rows(self) -> None: + source = _make_disagg_source() + df = source.fetch( + Query( + table="cftc.cot.disagg", + columns=["long_positions", "short_positions"], + start=pd.Timestamp("2026-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + ) + ) + assert not df.empty + assert { + "date", + "entity_id", + "asof_utc", + "long_positions", + "short_positions", + }.issubset(df.columns) + + def test_entity_ids_use_commodity_contract_codes(self) -> None: + source = _make_disagg_source() + df = source.fetch( + Query( + table="cftc.cot.disagg", + columns=["long_positions", "short_positions"], + start=pd.Timestamp("2026-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + ) + ) + entities = set(df["entity_id"].unique()) + assert "futures.wheat_srw.prod_merc.cftc" in entities + assert "futures.gold.swap.cftc" in entities + + def test_entity_filter(self) -> None: + source = _make_disagg_source() + df = source.fetch( + Query( + table="cftc.cot.disagg", + columns=["long_positions", "short_positions"], + start=pd.Timestamp("2026-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + entities=["futures.wheat_srw.m_money.cftc"], + ) + ) + assert not df.empty + assert df["entity_id"].eq("futures.wheat_srw.m_money.cftc").all() + + def test_missing_spread_columns_are_nan(self) -> None: + source = _make_disagg_source() + df = source.fetch( + Query( + table="cftc.cot.disagg", + columns=["spread_positions", "trader_category"], + start=pd.Timestamp("2026-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + entities=["futures.wheat_srw.prod_merc.cftc"], + ) + ) + assert not df.empty + assert df["trader_category"].eq("prod_merc").all() + assert df["spread_positions"].isna().all() + + def test_legacy_report_date_header_is_parsed(self) -> None: + source = _make_disagg_legacy_header_source() + df = source.fetch( + Query( + table="cftc.cot.disagg", + columns=["long_positions"], + start=pd.Timestamp("2026-01-01", tz="UTC"), + end=pd.Timestamp("2026-02-01", tz="UTC"), + entities=["futures.wheat_srw.prod_merc.cftc"], + ) + ) + assert not df.empty + assert pd.Timestamp(df["date"].iloc[0]) == pd.Timestamp("2026-01-09", tz="UTC") + + def test_year_urls_use_historical_batch_before_2010(self) -> None: + source = CFTCDisaggregatedCoTSource(file_urls=None) + assert source._year_urls(2008, 2018) == [ + "https://www.cftc.gov/files/dea/history/fut_disagg_txt_hist_2006_2016.zip", + "https://www.cftc.gov/files/dea/history/fut_disagg_txt_2017.zip", + "https://www.cftc.gov/files/dea/history/fut_disagg_txt_2018.zip", + ] + class TestPublicationDate: def test_tuesday_to_friday(self) -> None: diff --git a/tests/public_web/test_public_web_foundation.py b/tests/public_web/test_public_web_foundation.py new file mode 100644 index 0000000..c95a05f --- /dev/null +++ b/tests/public_web/test_public_web_foundation.py @@ -0,0 +1,250 @@ +from __future__ import annotations + +import pandas as pd +import pytest + +from alphaforge.data.public_web.archive import ( + ArchiveFetchPlanEntry, + discover_archive_fetches, + discover_archive_links, + filter_urls_for_years, + iter_yearly_archive_fetches, + iter_yearly_archive_urls, +) +from alphaforge.data.public_web.base import PublicWebSourceBase +from alphaforge.data.public_web.finalize import ( + empty_frame_for_schema, + finalize_public_frame, +) +from alphaforge.data.public_web.registry_api import RegistryApiSourceBase +from alphaforge.data.public_web.schema_helpers import ( + event_table_schema, + single_value_schema, +) +from alphaforge.data.public_web.tabular import ( + artifact_name_from_url, + candidate_tables, + resolved_date_series, + resolved_text_series, +) +from alphaforge.data.query import Query + + +class _DummySource(PublicWebSourceBase): + name = "dummy" + TABLE = "dummy_series" + + def schemas(self): + return {self.TABLE: single_value_schema(self.TABLE)} + + +class _DummyRegistrySource(RegistryApiSourceBase): + name = "dummy_registry" + TABLE = "dummy_registry_series" + + def __init__(self, *, registry_entries=None) -> None: + super().__init__(http_client=None) + self._init_registry( + "dummy_registry.yaml", + registry_entries=registry_entries, + registry_path=None, + ) + + def schemas(self): + return {self.TABLE: single_value_schema(self.TABLE)} + + +def test_empty_frame_for_schema_uses_schema_columns() -> None: + schema = single_value_schema("dummy_series") + + out = empty_frame_for_schema(schema) + + assert list(out.columns) == ["date", "entity_id", "asof_utc", "value"] + assert out.empty + + +def test_event_table_schema_preserves_event_metadata() -> None: + schema = event_table_schema( + "dummy.events", + required_columns=["value"], + native_freq="D", + time_semantics="point", + expected_cadence_days=1, + ) + + assert schema.time_column == "ts_utc" + assert schema.event_time_column == "ts_utc" + assert schema.native_freq == "D" + assert schema.time_semantics == "point" + assert schema.expected_cadence_days == 1 + + +def test_finalize_public_frame_filters_projects_and_sorts() -> None: + schema = single_value_schema("dummy_series") + frame = pd.DataFrame( + { + "date": pd.to_datetime(["2020-02-29", "2020-01-31"], utc=True), + "entity_id": ["series_b", "series_a"], + "asof_utc": pd.to_datetime(["2020-03-02", "2020-03-01"], utc=True), + "value": [2.0, 1.0], + "ignored": [9, 8], + } + ) + + out = finalize_public_frame( + frame, + q=Query( + table="dummy_series", + columns=["value"], + start=pd.Timestamp("2020-01-01"), + end=pd.Timestamp("2020-12-31"), + entities=["series_a", "series_b"], + asof=pd.Timestamp("2020-03-02"), + ), + schema=schema, + ) + + assert list(out.columns) == ["date", "entity_id", "asof_utc", "value"] + assert out["entity_id"].tolist() == ["series_a", "series_b"] + assert out["value"].tolist() == [1.0, 2.0] + + +def test_public_web_source_base_builds_schema_empty_frame_from_records() -> None: + src = _DummySource() + schema = src.schemas()[src.TABLE] + + out = src._frame_from_records([], schema=schema) + + assert list(out.columns) == ["date", "entity_id", "asof_utc", "value"] + assert out.empty + + +def test_registry_api_source_iter_entity_configs_requires_entities() -> None: + src = _DummyRegistrySource(registry_entries=[{"entity_id": "TEST", "foo": "bar"}]) + + with pytest.raises(ValueError, match="requires entities"): + list( + src._iter_entity_configs( + Query(table=src.TABLE, columns=["value"]), + error_message="dummy registry requires entities", + ) + ) + + +def test_registry_api_source_iter_entity_configs_skips_unknown_entities() -> None: + src = _DummyRegistrySource(registry_entries=[{"entity_id": "TEST", "foo": "bar"}]) + + pairs = list( + src._iter_entity_configs( + Query(table=src.TABLE, columns=["value"], entities=["TEST", "MISSING"]), + error_message="dummy registry requires entities", + ) + ) + + assert len(pairs) == 1 + assert pairs[0][0] == "TEST" + assert pairs[0][1]["foo"] == "bar" + + +def test_tabular_candidate_tables_filters_by_any_and_all_columns() -> None: + matching = pd.DataFrame({"volume": [1], "date": ["2020-01-01"]}) + missing = pd.DataFrame({"price": [1]}) + + out = candidate_tables( + [matching, missing], + any_of=["volume", "open_interest"], + all_of=["date"], + ) + + assert out == [matching] + + +def test_tabular_resolved_helpers_apply_defaults_and_normalization() -> None: + frame = pd.DataFrame({"trading_day": ["2020-01-02"], "group": ["Equity Index"]}) + + dates = resolved_date_series( + frame, + ["date", "trading_day"], + default_date=pd.Timestamp("2020-01-01", tz="UTC"), + ) + groups = resolved_text_series( + frame, + ["product_group", "group"], + default="unknown", + case="lower", + space_replacement="_", + ) + + assert str(dates.iloc[0]) == "2020-01-02 00:00:00+00:00" + assert groups.iloc[0] == "equity_index" + + +def test_tabular_artifact_name_from_url_strips_query_string() -> None: + assert artifact_name_from_url("https://example.com/path/data.csv?x=1", "fallback") == "data.csv" + + +def test_archive_helpers_discover_filter_and_plan_urls() -> None: + html = ( + 'a' + 'b' + ) + links = discover_archive_links( + html, + base_url="https://example.com/archive/index.html", + suffixes=[".zip"], + ) + + filtered = filter_urls_for_years(links, [2021]) + planned = discover_archive_fetches( + html, + base_url="https://example.com/archive/index.html", + suffixes=[".zip"], + years=[2021], + ) + yearly = iter_yearly_archive_urls( + start_year=2009, + end_year=2012, + url_template="https://example.com/data_{year}.zip", + first_year=2006, + yearly_first_year=2010, + historical_url="https://example.com/historical.zip", + historical_last_year=2011, + ) + yearly_planned = iter_yearly_archive_fetches( + start_year=2009, + end_year=2012, + url_template="https://example.com/data_{year}.zip", + first_year=2006, + yearly_first_year=2010, + historical_url="https://example.com/historical.zip", + historical_last_year=2011, + ) + + assert links == [ + "https://example.com/files/report_2020.zip?download=1", + "https://example.com/files/report_2021.zip", + ] + assert filtered == ["https://example.com/files/report_2021.zip"] + assert planned == [ + ArchiveFetchPlanEntry( + url="https://example.com/files/report_2021.zip", + artifact_name="report_2021.zip", + year=2021, + ) + ] + assert yearly == [ + "https://example.com/historical.zip", + "https://example.com/data_2012.zip", + ] + assert yearly_planned == [ + ArchiveFetchPlanEntry( + url="https://example.com/historical.zip", + artifact_name="historical.zip", + year=None, + ), + ArchiveFetchPlanEntry( + url="https://example.com/data_2012.zip", + artifact_name="data_2012.zip", + year=2012, + ), + ] diff --git a/tests/public_web/test_source_test_mapping.py b/tests/public_web/test_source_test_mapping.py index ba184e4..20c572f 100644 --- a/tests/public_web/test_source_test_mapping.py +++ b/tests/public_web/test_source_test_mapping.py @@ -7,17 +7,24 @@ NON_SOURCE_MODULES = { "__init__", + "archive", + "base", + "finalize", "http", "parsing", + "registry_api", "utils", "registry", "registry_loader", + "schema_helpers", + "tabular", } NON_MAPPING_TEST_FILES = { "__init__.py", "_fake_http.py", "test_live_sources.py", + "test_public_web_foundation.py", "test_source_test_mapping.py", } diff --git a/tests/test_cftc_dtcc_adapter.py b/tests/test_cftc_dtcc_adapter.py index 7294115..8977d63 100644 --- a/tests/test_cftc_dtcc_adapter.py +++ b/tests/test_cftc_dtcc_adapter.py @@ -6,6 +6,8 @@ from __future__ import annotations +from dataclasses import dataclass + import duckdb import pandas as pd import pytest @@ -32,6 +34,24 @@ def _raw_cot_df(): ) +def _raw_disagg_cot_df(): + """Mimics CFTCDisaggregatedCoTSource.fetch() output.""" + return pd.DataFrame( + { + "date": pd.to_datetime(["2025-01-10", "2025-01-10", "2025-01-17", "2025-01-17"]), + "entity_id": [ + "futures.wheat_srw.m_money.cftc", + "futures.gold.swap.cftc", + "futures.wheat_srw.m_money.cftc", + "futures.gold.swap.cftc", + ], + "long_positions": [3000.0, 5000.0, 3200.0, 5100.0], + "short_positions": [2800.0, 4200.0, 2900.0, 4300.0], + "open_interest": [15000.0, 25000.0, 15500.0, 25200.0], + } + ) + + def _raw_dtcc_df(): """Mimics DTCCPPDSource.fetch() output.""" return pd.DataFrame( @@ -49,6 +69,66 @@ def _raw_dtcc_df(): ) +def _raw_dtcc_family_df(): + """Daily DTCC output with distinct FX and IRS family entity ids.""" + return pd.DataFrame( + { + "date": pd.to_datetime( + [ + "2025-01-02", + "2025-01-02", + "2025-01-02", + "2025-01-02", + "2025-01-03", + "2025-01-03", + ] + ), + "entity_id": [ + "dtccppd.fx.fx_forward.usd.1m", + "dtccppd.fx.fx_swap.eur.3m", + "dtccppd.rates.interest_rate_swap.usd.5y", + "dtccppd.rates.cross_currency_swap.usd.5y", + "dtccppd.fx.fx_swap.gbp.6m", + "dtccppd.rates.interest_rate_swap.eur.10y", + ], + "asof_utc": pd.to_datetime( + [ + "2025-01-02T15:00:00Z", + "2025-01-02T15:00:00Z", + "2025-01-02T15:00:00Z", + "2025-01-02T15:00:00Z", + "2025-01-03T15:00:00Z", + "2025-01-03T15:00:00Z", + ] + ), + "trade_count": [100, 90, 60, 40, 120, 55], + "notional_sum": [10e6, 15e6, 50e6, 35e6, 12e6, 42e6], + "price_mean": [1.085, 0.0025, 0.032, 0.015, 0.0040, 0.028], + "price_std": [0.01, 0.002, 0.001, 0.0015, 0.0025, 0.0012], + "notional_median": [5e6, 7.5e6, 25e6, 17.5e6, 6e6, 21e6], + "trade_count_large": [4, 5, 3, 2, 6, 3], + "dv01_proxy_sum": [0.0, 0.0, 120_000.0, 80_000.0, 0.0, 140_000.0], + } + ) + + +@dataclass +class _StubDTCCSource: + table_frames: dict[str, pd.DataFrame] + queries: list[Query] + + def __init__(self, table_frames: dict[str, pd.DataFrame]) -> None: + self.table_frames = table_frames + self.queries = [] + + def schemas(self): + return {} + + def fetch(self, q: Query) -> pd.DataFrame: + self.queries.append(q) + return self.table_frames[q.table].copy() + + # --------------------------------------------------------------------------- # Transform tests (functions moved from positioning) # --------------------------------------------------------------------------- @@ -123,6 +203,19 @@ def test_series_key_format(self): result = dtcc_daily_to_pit_observations(df) assert any("dtcc.ppd.daily." in k for k in result["series_key"]) + def test_transform_allows_custom_prefix_and_source_name(self): + from alphaforge.data.transforms.dtcc_pit import dtcc_daily_to_pit_observations + + df = _raw_dtcc_family_df() + result = dtcc_daily_to_pit_observations( + df, + key_prefix="dtcc.fx.", + source_name="dtcc_ppd_fx", + ) + + assert result["series_key"].str.startswith("dtcc.fx.").all() + assert result["source"].eq("dtcc_ppd_fx").all() + class TestMeltToPitFormat: def test_basic_melt(self): @@ -169,7 +262,21 @@ def dtcc_adapter(tmp_path): conn = duckdb.connect(str(tmp_path / "cache.duckdb")) return DTCCAdapter( - raw_fetcher=lambda start, end: _raw_dtcc_df(), + source=_StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_df()}), + cache_conn=conn, + ) + + +@pytest.fixture +def cftc_multi_adapter(tmp_path): + from alphaforge.data.sources.cftc import CFTCAdapter + + conn = duckdb.connect(str(tmp_path / "cache.duckdb")) + return CFTCAdapter( + raw_fetchers={ + "cot.tff": lambda start, end: _raw_cot_df(), + "cot.disagg": lambda start, end: _raw_disagg_cot_df(), + }, cache_conn=conn, ) @@ -195,6 +302,156 @@ def test_source_name(self, dtcc_adapter): def test_datasets(self, dtcc_adapter): assert "dtcc.ppd" in dtcc_adapter.datasets + def test_fetch_uses_internal_raw_source_query(self, tmp_path): + from alphaforge.data.sources.dtcc import DTCCAdapter + + conn = duckdb.connect(str(tmp_path / "cache.duckdb")) + source = _StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_df()}) + adapter = DTCCAdapter(source=source, cache_conn=conn) + + q = Query( + table="dtcc.ppd", + columns=["value"], + entities=["dtcc.ppd.daily.dtccppd.fx.eur.trade_count"], + start="2025-01-01", + end="2025-01-31", + asof="2025-01-20", + ) + result = adapter.fetch(q) + + assert len(source.queries) == 1 + raw_query = source.queries[0] + assert raw_query.table == "dtcc.ppd.daily" + assert raw_query.start == pd.Timestamp("2025-01-01", tz="UTC") + assert raw_query.end == pd.Timestamp("2025-01-31", tz="UTC") + assert not result.data.empty + + def test_prefetch_uses_internal_raw_source(self, tmp_path): + from alphaforge.data.sources.dtcc import DTCCAdapter + + conn = duckdb.connect(str(tmp_path / "cache.duckdb")) + source = _StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_df()}) + adapter = DTCCAdapter(source=source, cache_conn=conn) + + manifest = adapter.prefetch( + "dtcc.ppd", + asof_range=(pd.Timestamp("2025-01-01").date(), pd.Timestamp("2025-01-31").date()), + ) + + assert len(source.queries) == 1 + assert source.queries[0].table == "dtcc.ppd.daily" + assert manifest.dataset == "dtcc.ppd" + assert manifest.source == "dtcc" + assert manifest.row_count > 0 + + +class TestDTCCAdapterBase: + def test_base_supports_subclass_dataset_contract(self, tmp_path): + from alphaforge.data.sources.dtcc import DTCCPPDAdapterBase + from alphaforge.data.transforms.dtcc_pit import dtcc_daily_to_pit_observations + + class DemoDTCCAdapter(DTCCPPDAdapterBase): + source_name = "dtcc_demo" + datasets = frozenset({"dtcc.demo"}) + + def _to_pit(self, dataset: str, raw_df: pd.DataFrame) -> pd.DataFrame: + assert dataset == "dtcc.demo" + return dtcc_daily_to_pit_observations(raw_df) + + conn = duckdb.connect(str(tmp_path / "cache.duckdb")) + source = _StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_df()}) + adapter = DemoDTCCAdapter(source=source, cache_conn=conn) + + result = adapter.fetch( + Query( + table="dtcc.demo", + columns=["value"], + entities=["dtcc.ppd.daily.dtccppd.fx.eur.trade_count"], + asof="2025-01-20", + ) + ) + + assert result.source == "dtcc_demo" + assert result.dataset == "dtcc.demo" + assert not result.data.empty + + +class TestDTCCProductFamilyAdapters: + def test_fx_adapter_filters_to_fx_entities(self, tmp_path): + from alphaforge.data.sources.dtcc import DTCCFXAdapter + + conn = duckdb.connect(str(tmp_path / "cache.duckdb")) + source = _StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_family_df()}) + adapter = DTCCFXAdapter(source=source, cache_conn=conn) + + result = adapter.fetch( + Query( + table="dtcc.fx", + columns=["value"], + entities=["dtcc.fx.dtccppd.fx.fx_forward.usd.1m.trade_count"], + asof="2025-01-20", + ) + ) + + assert result.source == "dtcc_fx" + assert not result.data.empty + assert result.data["series_key"].str.startswith("dtcc.fx.").all() + assert result.data["series_key"].str.contains(".fx.", regex=False).all() + assert result.data["source"].eq("dtcc_ppd_fx").all() + + def test_irs_adapter_filters_to_interest_rate_swaps(self, tmp_path): + from alphaforge.data.sources.dtcc import DTCCIRSAdapter + + conn = duckdb.connect(str(tmp_path / "cache.duckdb")) + source = _StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_family_df()}) + adapter = DTCCIRSAdapter(source=source, cache_conn=conn) + + result = adapter.fetch( + Query( + table="dtcc.irs", + columns=["value"], + entities=["dtcc.irs.dtccppd.rates.interest_rate_swap.usd.5y.trade_count"], + asof="2025-01-20", + ) + ) + + assert result.source == "dtcc_irs" + assert not result.data.empty + assert result.data["series_key"].str.startswith("dtcc.irs.").all() + assert result.data["series_key"].str.contains("interest_rate_swap").all() + assert not result.data["series_key"].str.contains("cross_currency_swap").any() + assert result.data["source"].eq("dtcc_ppd_irs").all() + + def test_family_adapters_route_through_data_context(self, tmp_path): + from alphaforge.data.context import DataContext + from alphaforge.data.sources.dtcc import DTCCFXAdapter, DTCCIRSAdapter + + source = _StubDTCCSource({"dtcc.ppd.daily": _raw_dtcc_family_df()}) + fx = DTCCFXAdapter(source=source, cache_conn=duckdb.connect(str(tmp_path / "fx.duckdb"))) + irs = DTCCIRSAdapter( + source=source, + cache_conn=duckdb.connect(str(tmp_path / "irs.duckdb")), + ) + + ctx = DataContext.from_adapters(fx, irs, calendars={}, store=None) + + fx_result = ctx.load( + "dtcc.fx", + columns=["value"], + entities=["dtcc.fx.dtccppd.fx.fx_forward.usd.1m.trade_count"], + asof="2025-01-20", + ) + irs_result = ctx.load( + "dtcc.irs", + columns=["value"], + entities=["dtcc.irs.dtccppd.rates.interest_rate_swap.usd.5y.trade_count"], + asof="2025-01-20", + ) + + assert ctx.default_sources == {"dtcc.fx": "dtcc_fx", "dtcc.irs": "dtcc_irs"} + assert fx_result.source == "dtcc_fx" + assert irs_result.source == "dtcc_irs" + # --------------------------------------------------------------------------- # Adapter fetch (bulk behavior) @@ -263,6 +520,46 @@ def counting_fetcher(start, end): assert result2.cached_at is not None assert len(result2.data) > 0 + def test_cache_hit_preserves_source_column(self, cftc_adapter): + """Cache-hit path must return the lineage 'source' column identical to cache-miss.""" + q = Query( + table="cot.tff", columns=["value"], + entities=["cftc.cot.tff.futures.eur.lev_money.cftc.net_positions"], + asof="2025-01-20", + ) + # First fetch (cache miss) — source column comes from cot_to_pit_observations + miss_result = cftc_adapter.fetch(q) + assert "source" in miss_result.data.columns, "cache-miss result must have 'source' column" + assert miss_result.data["source"].eq("cftc_cot").all(), ( + "cache-miss source must be 'cftc_cot' for cot.tff dataset" + ) + + # Second fetch for same entity — cache hit + hit_result = cftc_adapter.fetch(q) + assert hit_result.cached_at is not None, "second fetch should be a cache hit" + assert "source" in hit_result.data.columns, "cache-hit result must have 'source' column" + assert hit_result.data["source"].eq("cftc_cot").all(), ( + "cache-hit source must equal cache-miss source for cot.tff dataset" + ) + + def test_cache_hit_preserves_disagg_source_column(self, cftc_multi_adapter): + """Cache-hit path for disagg dataset must carry the correct 'cftc_cot_disagg' source.""" + q = Query( + table="cot.disagg", columns=["value"], + entities=["cftc.cot.disagg.futures.wheat_srw.m_money.cftc.net_positions"], + asof="2025-01-20", + ) + # First fetch (cache miss) + miss_result = cftc_multi_adapter.fetch(q) + assert "source" in miss_result.data.columns + assert miss_result.data["source"].eq("cftc_cot_disagg").all() + + # Second fetch (cache hit) + hit_result = cftc_multi_adapter.fetch(q) + assert hit_result.cached_at is not None + assert "source" in hit_result.data.columns + assert hit_result.data["source"].eq("cftc_cot_disagg").all() + class TestCFTCAdapterPrefetch: def test_prefetch_returns_manifest(self, cftc_adapter): @@ -271,3 +568,30 @@ def test_prefetch_returns_manifest(self, cftc_adapter): assert manifest.dataset == "cot.tff" assert manifest.row_count > 0 assert len(manifest.entity_keys) > 0 + + +class TestCFTCAdapterMultipleDatasets: + def test_declares_disagg_dataset(self, cftc_multi_adapter): + assert "cot.tff" in cftc_multi_adapter.datasets + assert "cot.disagg" in cftc_multi_adapter.datasets + + def test_fetch_disagg_dataset_uses_dataset_prefix(self, cftc_multi_adapter): + q = Query( + table="cot.disagg", + columns=["value"], + entities=["cftc.cot.disagg.futures.wheat_srw.m_money.cftc.net_positions"], + asof="2025-01-20", + ) + result = cftc_multi_adapter.fetch(q) + assert not result.data.empty + assert ( + result.data["series_key"] + == "cftc.cot.disagg.futures.wheat_srw.m_money.cftc.net_positions" + ).all() + assert result.data["source"].eq("cftc_cot_disagg").all() + + def test_prefetch_disagg_returns_manifest(self, cftc_multi_adapter): + manifest = cftc_multi_adapter.prefetch("cot.disagg") + assert manifest.source == "cftc" + assert manifest.dataset == "cot.disagg" + assert manifest.row_count > 0 diff --git a/tests/test_data_context_routing.py b/tests/test_data_context_routing.py index 7b99142..5a9d404 100644 --- a/tests/test_data_context_routing.py +++ b/tests/test_data_context_routing.py @@ -23,21 +23,37 @@ class FakeCFTCAdapter(SourceAdapterBase): def __init__(self): self.fetch_calls = [] + self.fetch_many_calls = [] - def fetch(self, query, *, max_staleness=None): - self.fetch_calls.append(query) - df = pd.DataFrame( - { - "obs_date": pd.to_datetime(["2025-01-31"]), - "asof_utc": pd.to_datetime(["2025-02-07"]), - "value": [42.0], - "series_key": ["cot.tff.eur.net"], - } + def _result_for_query(self, query): + entity = ( + query.entities[0] + if query.entities + else "cot.tff.eur.net" ) return FetchResult( - data=df, source="cftc", dataset="cot.tff", is_pit=True, cached_at=None, + data=pd.DataFrame( + { + "obs_date": pd.to_datetime(["2025-01-31"]), + "asof_utc": pd.to_datetime(["2025-02-07"]), + "value": [42.0], + "series_key": [entity], + } + ), + source="cftc", + dataset=query.table, + is_pit=True, + cached_at=None, ) + def fetch(self, query, *, max_staleness=None): + self.fetch_calls.append((query, max_staleness)) + return self._result_for_query(query) + + def fetch_many(self, queries, *, max_staleness=None): + self.fetch_many_calls.append((list(queries), max_staleness)) + return [self._result_for_query(query) for query in queries] + def list_entities(self, dataset): return ["cot.tff.eur.net", "cot.tff.gbp.net", "cot.tff.jpy.net"] @@ -48,20 +64,32 @@ class FakeBloombergAdapter(SourceAdapterBase): def __init__(self): self.fetch_calls = [] + self.fetch_many_calls = [] - def fetch(self, query, *, max_staleness=None): - self.fetch_calls.append(query) - df = pd.DataFrame( - { - "obs_date": pd.to_datetime(["2025-01-31"]), - "value": [43.0], - "series_key": ["cot.tff.eur.net"], - } - ) + def _result_for_query(self, query): + entity = query.entities[0] if query.entities else "cot.tff.eur.net" return FetchResult( - data=df, source="bloomberg", dataset="cot.tff", is_pit=True, cached_at=None, + data=pd.DataFrame( + { + "obs_date": pd.to_datetime(["2025-01-31"]), + "value": [43.0], + "series_key": [entity], + } + ), + source="bloomberg", + dataset=query.table, + is_pit=True, + cached_at=None, ) + def fetch(self, query, *, max_staleness=None): + self.fetch_calls.append((query, max_staleness)) + return self._result_for_query(query) + + def fetch_many(self, queries, *, max_staleness=None): + self.fetch_many_calls.append((list(queries), max_staleness)) + return [self._result_for_query(query) for query in queries] + def list_entities(self, dataset): if dataset == "cot.tff": return ["cot.tff.eur.net", "cot.tff.gbp.net"] @@ -74,21 +102,33 @@ class FakeFREDAdapter(SourceAdapterBase): def __init__(self): self.fetch_calls = [] + self.fetch_many_calls = [] - def fetch(self, query, *, max_staleness=None): - self.fetch_calls.append(query) - df = pd.DataFrame( - { - "obs_date": pd.to_datetime(["2025-01-31"]), - "asof_utc": pd.to_datetime(["2025-02-15"]), - "value": [3.1], - "series_key": ["GDP"], - } - ) + def _result_for_query(self, query): + entity = query.entities[0] if query.entities else "GDP" return FetchResult( - data=df, source="fred", dataset="gdp", is_pit=True, cached_at=None, + data=pd.DataFrame( + { + "obs_date": pd.to_datetime(["2025-01-31"]), + "asof_utc": pd.to_datetime(["2025-02-15"]), + "value": [3.1], + "series_key": [entity], + } + ), + source="fred", + dataset=query.table, + is_pit=True, + cached_at=None, ) + def fetch(self, query, *, max_staleness=None): + self.fetch_calls.append((query, max_staleness)) + return self._result_for_query(query) + + def fetch_many(self, queries, *, max_staleness=None): + self.fetch_many_calls.append((list(queries), max_staleness)) + return [self._result_for_query(query) for query in queries] + def list_entities(self, dataset): return ["GDP", "GDPC1"] @@ -160,21 +200,77 @@ def test_unknown_source_raises(self, ctx): class TestContextRouting: + def test_from_adapters_derives_defaults_and_load_helper(self, cftc, fred): + from alphaforge.data.context import DataContext + + ctx = DataContext.from_adapters(cftc, fred, calendars={}, store=None) + + assert ctx.default_sources == {"cot.tff": "cftc", "gdp": "fred"} + + result = ctx.load("gdp", columns=["value"], entities=["GDP"]) + + assert result.source == "fred" + assert len(fred.fetch_calls) == 1 + + def test_from_adapters_requires_default_for_ambiguous_datasets(self, cftc, bbg): + from alphaforge.data.context import DataContext + + ctx = DataContext.from_adapters(cftc, bbg, calendars={}, store=None) + + with pytest.raises(KeyError, match="Multiple adapters serve dataset 'cot.tff'"): + ctx.load("cot.tff", columns=["value"], entities=["cot.tff.eur.net"]) + def test_fetch_routes_to_adapter(self, ctx, fred): q = Query(table="gdp", columns=["value"], entities=["GDP"]) result = ctx.fetch(q) assert result.source == "fred" assert len(fred.fetch_calls) == 1 - def test_fetch_many_delegates(self, ctx, cftc, fred): + def test_fetch_many_batches_by_adapter_and_preserves_input_order( + self, ctx, cftc, fred + ): queries = [ Query(table="cot.tff", columns=["value"], entities=["cot.tff.eur.net"]), Query(table="gdp", columns=["value"], entities=["GDP"]), + Query(table="cot.tff", columns=["value"], entities=["cot.tff.jpy.net"]), ] results = ctx.fetch_many(queries) - assert len(results) == 2 - assert len(cftc.fetch_calls) == 1 - assert len(fred.fetch_calls) == 1 + assert len(results) == 3 + assert [result.source for result in results] == ["cftc", "fred", "cftc"] + assert [result.data["series_key"].iloc[0] for result in results] == [ + "cot.tff.eur.net", + "GDP", + "cot.tff.jpy.net", + ] + assert len(cftc.fetch_many_calls) == 1 + assert len(fred.fetch_many_calls) == 1 + assert len(cftc.fetch_calls) == 0 + assert len(fred.fetch_calls) == 0 + + def test_fetch_many_with_explicit_source_uses_single_adapter_batch( + self, ctx, bbg, cftc + ): + queries = [ + Query(table="cot.tff", columns=["value"], entities=["cot.tff.eur.net"]), + Query(table="cot.tff", columns=["value"], entities=["cot.tff.gbp.net"]), + ] + results = ctx.fetch_many(queries, source="bloomberg") + assert [result.source for result in results] == ["bloomberg", "bloomberg"] + assert len(bbg.fetch_many_calls) == 1 + assert len(bbg.fetch_calls) == 0 + assert len(cftc.fetch_many_calls) == 0 + + def test_fetch_many_forwards_max_staleness_to_batch_fetches( + self, ctx, cftc, fred + ): + threshold = pd.Timedelta(days=2) + queries = [ + Query(table="cot.tff", columns=["value"], entities=["cot.tff.eur.net"]), + Query(table="gdp", columns=["value"], entities=["GDP"]), + ] + ctx.fetch_many(queries, max_staleness=threshold) + assert cftc.fetch_many_calls[0][1] == threshold + assert fred.fetch_many_calls[0][1] == threshold def test_prefetch_delegates(self, ctx, cftc): manifest = ctx.prefetch("cot.tff") @@ -218,6 +314,74 @@ def test_old_sources_dict_still_works(self, tmp_path): df = ctx.sources["dummy"].fetch(q) assert len(df) == 1 + def test_canonical_fetch_does_not_fallback_to_legacy_sources(self, tmp_path, cftc): + from alphaforge.data.context import DataContext + from alphaforge.store.duckdb_parquet import DuckDBParquetStore + + class LegacyOnlySource: + name = "cftc" + + def __init__(self): + self.fetch_calls = [] + + def schemas(self): + return {} + + def fetch(self, q): + self.fetch_calls.append(q) + return pd.DataFrame( + { + "date": pd.to_datetime(["2025-01-31"]), + "entity_id": ["legacy.entity"], + "value": [999.0], + } + ) + + legacy = LegacyOnlySource() + ctx = DataContext( + sources={"cftc": legacy}, + calendars={}, + store=DuckDBParquetStore(root=str(tmp_path / "store")), + adapters={"cftc": cftc}, + default_sources={"cot.tff": "cftc"}, + ) + + result = ctx.fetch( + Query(table="cot.tff", columns=["value"], entities=["cot.tff.eur.net"]) + ) + + assert result.source == "cftc" + assert len(cftc.fetch_calls) == 1 + assert len(legacy.fetch_calls) == 0 + + def test_canonical_fetch_requires_adapter_registration(self, tmp_path): + from conftest import DummySource, MemoryStore + + ohlcv = pd.DataFrame( + { + "date": pd.to_datetime(["2020-01-02"]), + "entity_id": ["AAA"], + "close": [100.0], + } + ) + macro = pd.DataFrame( + { + "date": pd.to_datetime(["2020-01-31"]), + "entity_id": ["CPI"], + "value": [1.0], + } + ) + from alphaforge.data.context import DataContext + + ctx = DataContext( + sources={"dummy": DummySource(ohlcv, macro)}, + calendars={}, + store=MemoryStore(), + ) + + with pytest.raises(KeyError, match="No adapters registered"): + ctx.fetch(Query(table="market.ohlcv", columns=["close"], entities=["AAA"])) + def test_pit_accessor_still_works(self, ctx): """ctx.pit should still be a PITAccessor (when DuckDB store).""" from alphaforge.pit.accessor import PITAccessor diff --git a/tests/test_dataset_spec_composition.py b/tests/test_dataset_spec_composition.py new file mode 100644 index 0000000..88b4f20 --- /dev/null +++ b/tests/test_dataset_spec_composition.py @@ -0,0 +1,143 @@ +import json + +import pandas as pd +import pytest + +from alphaforge.data.context import DataContext +from alphaforge.features.dataset_builder import build_dataset +from alphaforge.features.dataset_spec import ( + DatasetSpec, + FeatureRequest, + FeatureRequestGroup, + JoinPolicy, + MissingnessPolicy, + SliceOverride, + TargetRequest, + TimeSpec, + UniverseSpec, +) +from alphaforge.features.frame import FeatureFrame +from alphaforge.features.template import SliceSpec +from alphaforge.time.calendar import TradingCalendar + + +class _RecordingTemplate: + version = "1.0" + param_space = {} + + def __init__(self, name: str): + self.name = name + self.seen_slices: list[SliceSpec] = [] + + def requires(self, params): + return [] + + def transform(self, ctx, params, slice: SliceSpec, state): + del params, state + self.seen_slices.append(slice) + cal = ctx.calendars["XNYS"] + sessions = cal.sessions(str(slice.start.date()), str(slice.end.date())) + idx = pd.MultiIndex.from_product( + [pd.DatetimeIndex([sessions[0]]), [slice.entities[0]]], + names=["ts_utc", "entity_id"], + ) + X = pd.DataFrame({self.name: [1.0]}, index=idx) + catalog = pd.DataFrame( + [{"feature_id": self.name, "family": "test"}] + ).set_index("feature_id") + return FeatureFrame(X=X, catalog=catalog, meta={}) + + +class _TinyTarget: + name = "target" + version = "1.0" + param_space = {} + + def transform(self, ctx, params, slice: SliceSpec, state): + del params, state + cal = ctx.calendars["XNYS"] + sessions = cal.sessions(str(slice.start.date()), str(slice.end.date())) + return pd.Series( + [0.1] * len(sessions), + index=pd.DatetimeIndex(sessions), + name="target", + ) + + +def test_feature_request_group_composes_keys_tags_and_slice_overrides(): + cal = TradingCalendar("XNYS", tz="UTC") + ctx = DataContext(sources={}, calendars={"XNYS": cal}, store=None) + + gdp = _RecordingTemplate("gdp_level") + cpi = _RecordingTemplate("cpi_level") + nested_asof = pd.Timestamp("2024-01-03T16:00:00Z") + + spec = DatasetSpec( + universe=UniverseSpec(entities=["AAA"]), + time=TimeSpec( + start=pd.Timestamp("2024-01-08T00:00:00Z"), + end=pd.Timestamp("2024-01-10T00:00:00Z"), + calendar="XNYS", + grid="B", + asof=pd.Timestamp("2024-01-10T16:00:00Z"), + ), + target=TargetRequest(template=_TinyTarget()), + features=[ + FeatureRequestGroup( + key="macro", + tags={"family": "macro", "stage": "group"}, + slice_override=SliceOverride(lookback=pd.Timedelta(days=7)), + requests=[ + FeatureRequest( + template=gdp, + key="gdp", + tags={"series": "gdp"}, + ), + FeatureRequestGroup( + key="inflation", + tags={"stage": "nested"}, + requests=[ + FeatureRequest( + template=cpi, + key="cpi", + tags={"series": "cpi"}, + slice_override=SliceOverride(asof=nested_asof), + ) + ], + ), + ], + ) + ], + missingness=MissingnessPolicy(final_row_policy="keep"), + ) + + artifact = build_dataset(ctx, spec, persist=False) + catalog = artifact.catalog + + assert gdp.seen_slices[0].start == pd.Timestamp("2024-01-01T00:00:00Z") + assert gdp.seen_slices[0].asof == pd.Timestamp("2024-01-10T16:00:00Z") + assert cpi.seen_slices[0].start == pd.Timestamp("2024-01-01T00:00:00Z") + assert cpi.seen_slices[0].asof == nested_asof + + assert catalog.loc["gdp_level", "request_key"] == "macro/gdp" + assert catalog.loc["cpi_level", "request_key"] == "macro/inflation/cpi" + + assert json.loads(catalog.loc["gdp_level", "tags_json"]) == { + "family": "macro", + "stage": "group", + "series": "gdp", + } + assert json.loads(catalog.loc["cpi_level", "tags_json"]) == { + "family": "macro", + "stage": "nested", + "series": "cpi", + } + + +def test_dataset_contract_validates_join_and_missingness_policies(): + with pytest.raises(ValueError, match="JoinPolicy.how"): + JoinPolicy(how="left") + + with pytest.raises(ValueError, match="MissingnessPolicy.final_row_policy"): + MissingnessPolicy(final_row_policy="drop_some") + diff --git a/tests/test_first_rate_bars.py b/tests/test_first_rate_bars.py index e915cc4..869fd3c 100644 --- a/tests/test_first_rate_bars.py +++ b/tests/test_first_rate_bars.py @@ -4,7 +4,7 @@ import pandas as pd -from alphaforge import FirstRateBarsConfig, Query, build_first_rate_bars_context +from alphaforge import FirstRateBarsConfig, build_first_rate_bars_context def _write_rows(path: Path, rows: list[str]) -> None: @@ -50,14 +50,12 @@ def test_first_rate_bars_context_discovers_entities_and_fetches_rows(tmp_path) - assert adapter.list_entities("crypto.contract_price_5m") == ["BTC"] assert adapter.list_entities("index.level_5m") == ["DAX"] - fx = ctx.fetch( - Query( - table="fx.contract_price_5m", - columns=["bar_start_utc", "close", "volume"], - entities=["AUDUSD"], - start="2010-01-03T22:05:00Z", - end="2010-01-03T22:10:00Z", - ) + fx = ctx.load( + "fx.contract_price_5m", + columns=["bar_start_utc", "close", "volume"], + entities=["AUDUSD"], + start="2010-01-03T22:05:00Z", + end="2010-01-03T22:10:00Z", ) assert fx.source == "first_rate_bars" assert fx.dataset == "fx.contract_price_5m" @@ -69,22 +67,18 @@ def test_first_rate_bars_context_discovers_entities_and_fetches_rows(tmp_path) - ] assert fx.data["bar_start_utc"].tolist()[0] == pd.Timestamp("2010-01-03 22:00:00+00:00") - crypto = ctx.fetch( - Query( - table="crypto.contract_price_5m", - columns=["close", "volume"], - entities=["BTC"], - ) + crypto = ctx.load( + "crypto.contract_price_5m", + columns=["close", "volume"], + entities=["BTC"], ) assert crypto.data["series_key"].tolist() == ["BTC", "BTC"] assert crypto.data["close"].tolist() == [93.183, 93.24] - index = ctx.fetch( - Query( - table="index.level_5m", - columns=["bar_start_utc", "close"], - entities=["DAX"], - ) + index = ctx.load( + "index.level_5m", + columns=["bar_start_utc", "close"], + entities=["DAX"], ) assert index.data["series_key"].tolist() == ["DAX", "DAX"] assert "volume" not in index.data.columns diff --git a/tests/test_first_rate_futures_loader.py b/tests/test_first_rate_futures_loader.py index 4e22ad1..523bbed 100644 --- a/tests/test_first_rate_futures_loader.py +++ b/tests/test_first_rate_futures_loader.py @@ -5,7 +5,6 @@ import pandas as pd import pytest -from alphaforge import Query from alphaforge.futures import ( FirstRateFuturesConfig, FirstRateFuturesLoader, @@ -187,14 +186,12 @@ def test_context_adapter_reads_manifest_and_fetches_data(tmp_path) -> None: assert adapter.list_entities("futures.continuous_eod_research") == ["ES"] - result = ctx.fetch( - Query( - table="futures.continuous_eod_research", - columns=["close", "active_contract_id"], - entities=["ES"], - start="2024-06-17T00:00:00Z", - end="2024-06-21T00:00:00Z", - ) + result = ctx.load( + "futures.continuous_eod_research", + columns=["close", "active_contract_id"], + entities=["ES"], + start="2024-06-17T00:00:00Z", + end="2024-06-21T00:00:00Z", ) assert result.dataset == "futures.continuous_eod_research" diff --git a/tests/test_market_templates.py b/tests/test_market_templates.py new file mode 100644 index 0000000..33652e9 --- /dev/null +++ b/tests/test_market_templates.py @@ -0,0 +1,246 @@ +from __future__ import annotations + +import numpy as np +import pandas as pd + +from alphaforge.data.adapter import SourceAdapterBase +from alphaforge.data.context import DataContext +from alphaforge.data.query import Query +from alphaforge.data.types import FetchResult +from alphaforge.features import LagReturnsTemplate, RollingVolatilityTemplate +from alphaforge.features.dataset_builder import build_dataset +from alphaforge.features.dataset_spec import ( + DatasetSpec, + FeatureRequest, + FeatureRequestGroup, + JoinPolicy, + MissingnessPolicy, + TargetRequest, + TimeSpec, + UniverseSpec, +) +from alphaforge.features.template import SliceSpec +from alphaforge.time.calendar import TradingCalendar + + +class InMemoryMarketAdapter(SourceAdapterBase): + source_name = "market" + datasets = frozenset({"market.ohlcv"}) + + def __init__(self, frame: pd.DataFrame) -> None: + self._frame = frame.copy() + self.fetch_calls: list[Query] = [] + + def fetch(self, query: Query, *, max_staleness=None) -> FetchResult: + self.fetch_calls.append(query) + frame = self._frame.copy() + if query.entities is not None: + frame = frame[frame["series_key"].isin(query.entities)] + + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.start is not None: + frame = frame[obs >= query.start] + obs = pd.to_datetime(frame["obs_date"], utc=True) + if query.end is not None: + frame = frame[obs <= query.end] + + keep = ["series_key", "obs_date"] + [ + column for column in query.columns if column in frame.columns + ] + return FetchResult( + data=frame[keep].reset_index(drop=True), + source=self.source_name, + dataset=query.table, + is_pit=False, + cached_at=None, + ) + + def list_entities(self, dataset: str) -> list[str]: + return sorted(self._frame["series_key"].unique()) + + +def _market_frame() -> pd.DataFrame: + dates = pd.date_range("2024-01-02", periods=6, freq="B", tz="UTC") + rows: list[dict[str, object]] = [] + values = { + "AAA": [100.0, 101.0, 103.0, 102.0, 104.0, 105.0], + "BBB": [50.0, 49.0, 50.0, 51.0, 52.0, 53.0], + } + for entity, closes in values.items(): + for obs_date, close in zip(dates, closes, strict=True): + rows.append( + { + "series_key": entity, + "obs_date": obs_date, + "close": close, + "volume": 1000.0, + } + ) + return pd.DataFrame(rows) + + +def _ctx() -> tuple[DataContext, InMemoryMarketAdapter]: + adapter = InMemoryMarketAdapter(_market_frame()) + ctx = DataContext.from_adapters( + adapter, + calendars={"XNYS": TradingCalendar("XNYS", tz="UTC")}, + store=None, + ) + return ctx, adapter + + +def _price_index(frame: pd.DataFrame, *, asof: pd.Timestamp | None = None) -> pd.Series: + out = frame.copy() + out["ts_utc"] = pd.to_datetime(out["obs_date"], utc=True) + if asof is not None: + out = out[out["ts_utc"] <= asof] + return ( + out.set_index(["ts_utc", "series_key"])["close"] + .rename_axis(index=["ts_utc", "entity_id"]) + .sort_index() + ) + + +def test_lag_returns_template_uses_adapter_load_and_local_asof_filter() -> None: + ctx, adapter = _ctx() + asof = pd.Timestamp("2024-01-08T00:00:00Z") + template = LagReturnsTemplate() + + ff = template.transform( + ctx, + { + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "lags": [1, 2], + }, + SliceSpec( + start=pd.Timestamp("2024-01-02T00:00:00Z"), + end=pd.Timestamp("2024-01-12T00:00:00Z"), + entities=["AAA", "BBB"], + asof=asof, + grid="B", + ), + None, + ) + + assert len(adapter.fetch_calls) == 1 + assert adapter.fetch_calls[0].table == "market.ohlcv" + assert set(ff.catalog["lag"]) == {1, 2} + assert all(ff.X.index.get_level_values("ts_utc") <= asof) + + prices = _price_index(_market_frame(), asof=asof).astype(float) + logret = np.log(prices).groupby(level="entity_id").diff() + expected_lag_1 = logret.groupby(level="entity_id").shift(1).rename("expected") + feature_id = ff.catalog[ff.catalog["lag"] == 1].index[0] + got = ff.X[feature_id].rename("expected") + pd.testing.assert_series_equal(got, expected_lag_1, check_names=True) + + +def test_rolling_volatility_template_builds_annualized_shifted_windows() -> None: + ctx, _ = _ctx() + template = RollingVolatilityTemplate() + + ff = template.transform( + ctx, + { + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "windows": [3], + "lag": 1, + "annualization_factor": 252, + }, + SliceSpec( + start=pd.Timestamp("2024-01-02T00:00:00Z"), + end=pd.Timestamp("2024-01-12T00:00:00Z"), + entities=["AAA", "BBB"], + asof=None, + grid="B", + ), + None, + ) + + prices = _price_index(_market_frame()).astype(float) + logret = np.log(prices).groupby(level="entity_id").diff() + expected = logret.groupby(level="entity_id").transform( + lambda values: values.rolling(window=3, min_periods=3).std() + ) + expected = expected.groupby(level="entity_id").shift(1) * np.sqrt(252.0) + + feature_id = ff.catalog[ff.catalog["window"] == 3].index[0] + got = ff.X[feature_id].rename(expected.name) + pd.testing.assert_series_equal(got, expected, check_names=False) + + +def test_dataset_spec_recipe_can_use_built_in_market_templates() -> None: + ctx, _ = _ctx() + + spec = DatasetSpec( + universe=UniverseSpec(entities=["AAA", "BBB"]), + time=TimeSpec( + start=pd.Timestamp("2024-01-02T00:00:00Z"), + end=pd.Timestamp("2024-01-12T00:00:00Z"), + calendar="XNYS", + grid="B", + ), + features=[ + FeatureRequestGroup( + key="volatility", + tags={"recipe": "volatility"}, + requests=[ + FeatureRequest( + template=LagReturnsTemplate(), + key="returns", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "lags": [1, 2], + }, + ), + FeatureRequest( + template=RollingVolatilityTemplate(), + key="trailing_vol", + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "windows": [3, 5], + "lag": 1, + "annualization_factor": 252, + }, + ), + ], + ) + ], + target=TargetRequest( + template=RollingVolatilityTemplate(), + params={ + "dataset": "market.ohlcv", + "source": "market", + "price_col": "close", + "windows": [3], + "lag": 0, + "annualization_factor": 252, + }, + name="realized_vol_target", + ), + join_policy=JoinPolicy(how="inner", sort_index=True), + missingness=MissingnessPolicy(final_row_policy="keep"), + name="volatility_recipe", + ) + + artifact = build_dataset(ctx, spec, persist=False) + + assert not artifact.X.empty + assert artifact.X.shape[1] == 4 + assert artifact.y.name == "realized_vol_target" + assert set(artifact.catalog["request_key"]) == { + "volatility/returns", + "volatility/trailing_vol", + } + assert set(artifact.catalog["template_name"]) == { + "lag_returns", + "rolling_volatility", + } diff --git a/tests/test_pipeline_health.py b/tests/test_pipeline_health.py index de22af4..8b1dee0 100644 --- a/tests/test_pipeline_health.py +++ b/tests/test_pipeline_health.py @@ -8,7 +8,9 @@ DTCC_PPD_HEALTH_POLICY, SourceHealthPolicy, assess_source_health, + build_health_report, ) +from alphaforge.pit.release_rules import FixedLagMonths def _ts(days_ago: float) -> pd.Timestamp: @@ -110,3 +112,115 @@ def test_expected_next(self) -> None: expected = latest + pd.Timedelta(days=7) assert h.expected_next is not None assert abs((h.expected_next - expected).total_seconds()) < 1 + + +class TestReleaseAwareHealth: + def test_release_rule_defers_health_until_expected_release(self) -> None: + policy = SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=31), + release_rule=FixedLagMonths(months=2), + grace_period=pd.Timedelta(days=2), + stale_threshold=pd.Timedelta(days=14), + dead_threshold=pd.Timedelta(days=42), + weight_decay_half_life=pd.Timedelta(days=7), + ) + latest = pd.Timestamp("2025-01-31", tz="UTC") + asof = pd.Timestamp("2025-03-15", tz="UTC") + + h = assess_source_health("macro", latest, asof, policy) + + assert h.status == "ok" + assert h.expected_next == pd.Timestamp("2025-04-01", tz="UTC") + assert h.overdue_days == 0.0 + + def test_release_rule_uses_release_delay_for_late_status(self) -> None: + policy = SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=31), + release_rule=FixedLagMonths(months=2), + grace_period=pd.Timedelta(days=2), + stale_threshold=pd.Timedelta(days=14), + dead_threshold=pd.Timedelta(days=42), + weight_decay_half_life=pd.Timedelta(days=7), + ) + latest = pd.Timestamp("2025-01-31", tz="UTC") + asof = pd.Timestamp("2025-04-05", tz="UTC") + + h = assess_source_health("macro", latest, asof, policy) + + assert h.status == "late" + assert h.expected_next == pd.Timestamp("2025-04-01", tz="UTC") + assert h.overdue_days == 4.0 + + def test_weekly_release_weekday_changes_expected_next(self) -> None: + from alphaforge.time.release_rules import WeeklyRelease + + latest = pd.Timestamp("2025-01-04", tz="UTC") + asof = pd.Timestamp("2025-01-07", tz="UTC") + + monday = assess_source_health( + "weekly", + latest, + asof, + SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=7), + release_rule=WeeklyRelease(release_weekday="Monday", lag_days=0), + grace_period=pd.Timedelta(days=0), + stale_threshold=pd.Timedelta(days=7), + dead_threshold=pd.Timedelta(days=21), + weight_decay_half_life=pd.Timedelta(days=7), + ), + ) + thursday = assess_source_health( + "weekly", + latest, + asof, + SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=7), + release_rule=WeeklyRelease(release_weekday="Thursday", lag_days=0), + grace_period=pd.Timedelta(days=0), + stale_threshold=pd.Timedelta(days=7), + dead_threshold=pd.Timedelta(days=21), + weight_decay_half_life=pd.Timedelta(days=7), + ), + ) + + assert monday.expected_next == pd.Timestamp("2025-01-13", tz="UTC") + assert thursday.expected_next == pd.Timestamp("2025-01-16", tz="UTC") + + +class TestHealthReport: + def test_build_health_report_keeps_release_aware_diagnostics(self) -> None: + macro_policy = SourceHealthPolicy( + expected_cadence=pd.Timedelta(days=31), + release_rule=FixedLagMonths(months=2), + grace_period=pd.Timedelta(days=2), + stale_threshold=pd.Timedelta(days=14), + dead_threshold=pd.Timedelta(days=42), + weight_decay_half_life=pd.Timedelta(days=7), + ) + asof = pd.Timestamp("2025-04-05", tz="UTC") + + report = build_health_report( + { + "macro": assess_source_health( + "macro", + pd.Timestamp("2025-01-31", tz="UTC"), + asof, + macro_policy, + ), + "weekly": assess_source_health( + "weekly", + pd.Timestamp("2025-03-28", tz="UTC"), + asof, + _POLICY, + ), + } + ) + + assert report["source_name"].tolist() == ["macro", "weekly"] + assert "expected_next" in report.columns + assert "overdue_days" in report.columns + macro_row = report.set_index("source_name").loc["macro"] + assert macro_row["status"] == "late" + assert macro_row["expected_next"] == pd.Timestamp("2025-04-01", tz="UTC") + assert macro_row["overdue_days"] == 4.0 diff --git a/tests/test_pit_accessor.py b/tests/test_pit_accessor.py index 265ffd6..5831112 100644 --- a/tests/test_pit_accessor.py +++ b/tests/test_pit_accessor.py @@ -45,6 +45,16 @@ def test_upsert_idempotency(tmp_path): assert count == len(df) +def test_open_bootstraps_store_root(tmp_path): + pit = PITAccessor.open(tmp_path) + df = _sample_df() + pit.upsert_pit_observations(df) + + snap = pit.get_snapshot("GDP", pd.Timestamp("2025-03-01", tz="UTC")) + + assert snap.loc[pd.Timestamp("2024-12-31", tz="UTC")] == 1.1 + + def test_snapshot_selection(tmp_path): pit = _make_accessor(tmp_path) df = _sample_df() diff --git a/tests/test_pit_explainability.py b/tests/test_pit_explainability.py new file mode 100644 index 0000000..43094fa --- /dev/null +++ b/tests/test_pit_explainability.py @@ -0,0 +1,105 @@ +import pandas as pd + +from alphaforge.pit.accessor import PITAccessor +from alphaforge.pit.models import PITExpressionGraphSpec, PITExpressionNode +from alphaforge.pit.transforms import PITTransformSpec +from alphaforge.store.duckdb_parquet import DuckDBParquetStore + + +def _make_accessor(tmp_path) -> PITAccessor: + store = DuckDBParquetStore(root=str(tmp_path)) + return PITAccessor(store.conn()) + + +def _sample_cross_df() -> pd.DataFrame: + return pd.DataFrame( + { + "series_key": ["GDP", "GDP", "CPI", "CPI"], + "obs_date": [ + pd.Timestamp("2024-01-31"), + pd.Timestamp("2024-02-29"), + pd.Timestamp("2024-01-31"), + pd.Timestamp("2024-02-29"), + ], + "asof_utc": [ + pd.Timestamp("2024-03-10", tz="UTC"), + pd.Timestamp("2024-03-10", tz="UTC"), + pd.Timestamp("2024-03-05", tz="UTC"), + pd.Timestamp("2024-03-05", tz="UTC"), + ], + "value": [3.0, 3.5, 1.0, 1.1], + } + ) + + +def test_get_series_lineage_and_summary_for_transform_output(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_cross_df()) + + spec = PITTransformSpec( + input_series_key="GDP", + output_series_key="GDP_MINUS_CPI", + op="binary", + params={"right_series_key": "CPI", "operator": "sub", "join": "inner"}, + engine="python", + ) + pit.apply_transform(spec, overwrite=True) + + lineage = pit.get_series_lineage("GDP_MINUS_CPI") + + assert not lineage.empty + assert {"transform_id", "input_series_keys", "max_source_asof_utc", "causality_status"}.issubset( + lineage.columns + ) + assert lineage["lineage_kind"].eq("transform").all() + assert lineage["causality_status"].eq("ok").all() + assert all(keys == ("GDP", "CPI") for keys in lineage["input_series_keys"]) + assert (lineage["max_source_asof_utc"] <= lineage["asof_utc"]).all() + + summary = pit.explain_series("GDP_MINUS_CPI") + assert summary["derived_row_count"] == len(lineage) + assert summary["input_series_keys"] == ["CPI", "GDP"] + assert summary["transform_ids"] == [spec.transform_id()] + assert summary["causality_safe"] is True + + +def test_get_series_lineage_and_summary_for_expression_graph_output(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_cross_df()) + + graph = PITExpressionGraphSpec( + graph_id="macro/gdp_expr", + nodes=( + PITExpressionNode( + name="spread", + output_series_key="GDP_MINUS_CPI_EXPR", + expression="gdp - cpi", + inputs={"gdp": "GDP", "cpi": "CPI"}, + join="inner", + ), + ), + ) + pit.apply_expression_graph(graph, overwrite=True) + + lineage = pit.get_series_lineage("GDP_MINUS_CPI_EXPR") + + assert not lineage.empty + assert lineage["lineage_kind"].eq("expression_graph").all() + assert lineage["graph_id"].eq("macro/gdp_expr").all() + assert lineage["node_name"].eq("spread").all() + assert lineage["causality_status"].eq("ok").all() + + summary = pit.explain_series("GDP_MINUS_CPI_EXPR") + assert summary["graph_ids"] == ["macro/gdp_expr"] + assert summary["causality_safe"] is True + + +def test_explain_series_handles_raw_series_without_lineage(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_cross_df()) + + summary = pit.explain_series("GDP") + + assert summary["derived_row_count"] == 0 + assert summary["lineage_kinds"] == ["raw"] + assert summary["causality_status_counts"] == {"raw": 2} diff --git a/tests/test_pit_gdp_derived.py b/tests/test_pit_gdp_derived.py index f31c263..7a58c6e 100644 --- a/tests/test_pit_gdp_derived.py +++ b/tests/test_pit_gdp_derived.py @@ -73,7 +73,12 @@ def test_get_snapshot_multi_matches_single_snapshot_union(tmp_path): ], ignore_index=True, )[["series_key", "obs_date", "value"]].sort_values(["series_key", "obs_date"]).reset_index(drop=True) - pd.testing.assert_frame_equal(multi, expected, check_dtype=False) + pd.testing.assert_frame_equal( + multi[["series_key", "obs_date", "value"]], + expected, + check_dtype=False, + ) + assert multi["source_asof_utc"].notna().all() def test_get_revision_path_multi_matches_single_requests(tmp_path): diff --git a/tests/test_pit_ref_queries.py b/tests/test_pit_ref_queries.py new file mode 100644 index 0000000..feda646 --- /dev/null +++ b/tests/test_pit_ref_queries.py @@ -0,0 +1,115 @@ +import pandas as pd + +from alphaforge import RefRevisionQuery as TopLevelRefRevisionQuery +from alphaforge import RefSnapshotQuery as TopLevelRefSnapshotQuery +from alphaforge.pit import PITAccessor, RefRevisionQuery, RefSnapshotQuery +from alphaforge.pit.ref_entity import make_ref_entity_id +from alphaforge.store.duckdb_parquet import DuckDBParquetStore +from alphaforge.time.ref_period import RefFreq, RefPeriod + + +def _make_accessor(tmp_path) -> PITAccessor: + store = DuckDBParquetStore(root=str(tmp_path)) + return PITAccessor(store.conn()) + + +def _sample_quarterly_df() -> pd.DataFrame: + return pd.DataFrame( + { + "series_key": ["GDP", "GDP", "GDP", "GDP"], + "obs_date": [ + pd.Timestamp("2024-12-31"), + pd.Timestamp("2024-12-31"), + pd.Timestamp("2025-03-31"), + pd.Timestamp("2025-03-31"), + ], + "asof_utc": [ + pd.Timestamp("2025-01-10", tz="UTC"), + pd.Timestamp("2025-02-10", tz="UTC"), + pd.Timestamp("2025-04-10", tz="UTC"), + pd.Timestamp("2025-05-10", tz="UTC"), + ], + "value": [1.0, 1.1, 2.0, 2.1], + } + ) + + +def _sample_monthly_start_anchor_df() -> pd.DataFrame: + return pd.DataFrame( + { + "series_key": ["CPI", "CPI"], + "obs_date": [ + pd.Timestamp("2025-01-01"), + pd.Timestamp("2025-02-01"), + ], + "asof_utc": [ + pd.Timestamp("2025-01-15", tz="UTC"), + pd.Timestamp("2025-02-15", tz="UTC"), + ], + "value": [3.0, 3.1], + } + ) + + +def test_public_ref_query_exports_are_canonical() -> None: + assert TopLevelRefSnapshotQuery is RefSnapshotQuery + assert TopLevelRefRevisionQuery is RefRevisionQuery + + +def test_snapshot_ref_query_returns_ref_period_index(tmp_path) -> None: + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_quarterly_df()) + + snap = pit.snapshot_ref( + RefSnapshotQuery( + series_key="GDP", + asof=pd.Timestamp("2025-06-01", tz="UTC"), + start_ref="2024Q4", + end_ref=pd.Period("2025Q1", freq="Q"), + ) + ) + + assert snap.index.name == "ref_period" + assert list(snap.index) == [RefPeriod.parse("2024Q4"), RefPeriod.parse("2025Q1")] + assert snap.loc[RefPeriod.parse("2024Q4")] == 1.1 + assert snap.loc[RefPeriod.parse("2025Q1")] == 2.1 + + +def test_snapshot_ref_query_accepts_explicit_obs_date_anchor(tmp_path) -> None: + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_monthly_start_anchor_df()) + + snap = pit.snapshot_ref( + RefSnapshotQuery( + series_key="CPI", + asof=pd.Timestamp("2025-03-01", tz="UTC"), + start_ref="2025-01-01", + end_ref="2025-02-01", + freq=RefFreq.M, + obs_date_anchor="start", + ) + ) + + assert list(snap.index) == [RefPeriod.parse("2025-01"), RefPeriod.parse("2025-02")] + assert snap.loc[RefPeriod.parse("2025-01")] == 3.0 + assert snap.loc[RefPeriod.parse("2025-02")] == 3.1 + + +def test_revisions_ref_query_accepts_mapping_and_sets_ref_entity_name(tmp_path) -> None: + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_quarterly_df()) + + timeline = pit.revisions_ref( + { + "series_key": "GDP", + "ref": pd.Period("2024Q4", freq="Q"), + "end_asof": pd.Timestamp("2025-02-15", tz="UTC"), + } + ) + + assert list(timeline.index) == [ + pd.Timestamp("2025-01-10", tz="UTC"), + pd.Timestamp("2025-02-10", tz="UTC"), + ] + assert list(timeline.values) == [1.0, 1.1] + assert timeline.name == make_ref_entity_id("GDP", RefPeriod.parse("2024Q4")) diff --git a/tests/test_pit_release_helpers.py b/tests/test_pit_release_helpers.py index 1998b19..ff2337f 100644 --- a/tests/test_pit_release_helpers.py +++ b/tests/test_pit_release_helpers.py @@ -74,3 +74,12 @@ def test_release_helpers_validate_ref_frequency(tmp_path): with pytest.raises(PITContractError, match="frequency"): pit.list_release_stream("GDP", "2024Q4", freq=RefFreq.M) + + +def test_release_helpers_accept_pandas_period_refs(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_release_df()) + + stream = pit.list_release_stream("GDP", pd.Period("2024Q4", freq="Q")) + + assert stream["ref_key"].tolist() == ["2024Q4", "2024Q4", "2024Q4"] diff --git a/tests/test_pit_union_vintages_snapshot_panel.py b/tests/test_pit_union_vintages_snapshot_panel.py index 373a532..f7b5aa3 100644 --- a/tests/test_pit_union_vintages_snapshot_panel.py +++ b/tests/test_pit_union_vintages_snapshot_panel.py @@ -107,3 +107,83 @@ def test_build_snapshot_panel_respects_asof_cut(tmp_path): asof=pd.Timestamp("2024-03-15", tz="UTC"), ) assert panel.index.max() == pd.Timestamp("2024-02-29", tz="UTC") + + +def test_get_snapshot_multi_includes_source_asof_metadata(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_df()) + + batch = pit.get_snapshot_multi( + ["GDP", "CPI"], + pd.Timestamp("2024-04-30", tz="UTC"), + ) + + assert {"series_key", "obs_date", "source_asof_utc", "value"} == set(batch.columns) + assert batch["source_asof_utc"].notna().all() + + +def test_build_snapshot_panel_long_preserves_source_metadata(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations(_sample_df()) + + panel = pit.build_snapshot_panel_long( + [ + {"series_key": "GDP", "alias": "gdp"}, + {"series_key": "CPI", "alias": "cpi", "release_policy": "latest"}, + ], + asof=pd.Timestamp("2024-04-30", tz="UTC"), + align="month_end", + ) + + assert { + "series_key", + "series_alias", + "obs_date", + "source_obs_date", + "source_asof_utc", + "value", + } == set(panel.columns) + assert set(panel["series_alias"]) == {"gdp", "cpi"} + assert panel["source_asof_utc"].notna().all() + assert panel["obs_date"].dt.tz is not None + + +def test_build_snapshot_panel_long_accepts_ref_anchor_specs(tmp_path): + pit = _make_accessor(tmp_path) + pit.upsert_pit_observations( + pd.DataFrame( + { + "series_key": ["CPI", "CPI"], + "obs_date": [pd.Timestamp("2025-01-01"), pd.Timestamp("2025-02-01")], + "asof_utc": [ + pd.Timestamp("2025-01-15", tz="UTC"), + pd.Timestamp("2025-02-15", tz="UTC"), + ], + "value": [3.0, 3.1], + } + ) + ) + + panel = pit.build_snapshot_panel_long( + [ + { + "series_key": "CPI", + "alias": "cpi", + "start_ref": "2025-01-01", + "end_ref": "2025-02-01", + "freq": "M", + "obs_date_anchor": "start", + } + ], + asof=pd.Timestamp("2025-03-01", tz="UTC"), + align="month_end", + ) + + assert list(panel["source_obs_date"]) == [ + pd.Timestamp("2025-01-01", tz="UTC"), + pd.Timestamp("2025-02-01", tz="UTC"), + ] + assert list(panel["obs_date"]) == [ + pd.Timestamp("2025-01-31", tz="UTC"), + pd.Timestamp("2025-02-28", tz="UTC"), + ] diff --git a/tests/test_public_web_registry_exports.py b/tests/test_public_web_registry_exports.py new file mode 100644 index 0000000..2c0ea4f --- /dev/null +++ b/tests/test_public_web_registry_exports.py @@ -0,0 +1,18 @@ +from __future__ import annotations + +import alphaforge +from alphaforge.data.public_web import MOFJGBYieldCurveSource + + +def test_root_exports_default_public_web_registry() -> None: + sources = alphaforge.default_public_web_sources() + + assert "mof_jgb_yields" in sources + assert "philadelphia_spf" in sources + assert "cftc_cot_disagg" in sources + assert isinstance(sources["mof_jgb_yields"], MOFJGBYieldCurveSource) + assert "mof.jgb.yields" in sources["mof_jgb_yields"].schemas() + + +def test_root_exports_mof_source_constructor() -> None: + assert alphaforge.MOFJGBYieldCurveSource is MOFJGBYieldCurveSource diff --git a/tests/test_ref_period.py b/tests/test_ref_period.py index 3f70a16..58cc473 100644 --- a/tests/test_ref_period.py +++ b/tests/test_ref_period.py @@ -1,4 +1,5 @@ import pandas as pd +import pytest from alphaforge.time.ref_period import RefFreq, RefPeriod @@ -51,3 +52,43 @@ def test_refperiod_round_trip_from_obs_end(): ref = RefPeriod.parse("2024Q4") round_trip = RefPeriod.from_obs_date_end(ref.end_obs_date(), freq=RefFreq.Q) assert round_trip == ref + + +def test_refperiod_normalizes_pandas_period_inputs(): + assert RefPeriod.parse(pd.Period("2025Q1", freq="Q")).to_key() == "2025Q1" + assert RefPeriod.parse(pd.Period("2025-02", freq="M")).to_key() == "2025-02" + assert RefPeriod.parse(pd.Period("2025", freq="Y")).to_key() == "2025" + + +def test_refperiod_normalizes_obs_dates_with_explicit_freq_and_anchor(): + assert RefPeriod.parse( + pd.Timestamp("2025-03-31"), + freq=RefFreq.Q, + obs_date_anchor="end", + ).to_key() == "2025Q1" + assert RefPeriod.parse( + pd.Timestamp("2025-01-01"), + freq=RefFreq.Q, + obs_date_anchor="start", + ).to_key() == "2025Q1" + assert RefPeriod.parse( + "2025-12-31", + freq=RefFreq.A, + obs_date_anchor="end", + ).to_key() == "2025" + + +def test_refperiod_rejects_obs_dates_that_do_not_match_anchor(): + with pytest.raises(ValueError, match="requested reference period"): + RefPeriod.parse( + pd.Timestamp("2025-02-01"), + freq=RefFreq.Q, + obs_date_anchor="start", + ) + + +def test_refperiod_start_and_anchored_obs_dates(): + ref = RefPeriod.parse("2025Q2") + assert ref.start_obs_date() == pd.Timestamp("2025-04-01", tz="UTC") + assert ref.obs_date(anchor="start") == pd.Timestamp("2025-04-01", tz="UTC") + assert ref.obs_date(anchor="end") == pd.Timestamp("2025-06-30", tz="UTC") diff --git a/tests/test_release_rules.py b/tests/test_release_rules.py index a7c555c..848a9e5 100644 --- a/tests/test_release_rules.py +++ b/tests/test_release_rules.py @@ -70,6 +70,18 @@ def test_five_day_lag(self): result = rule.expected_release_date(date(2025, 1, 4)) assert result == date(2025, 1, 9) + def test_release_weekday_changes_result_when_needed(self): + monday = WeeklyRelease(release_weekday="Monday", lag_days=0) + thursday = WeeklyRelease(release_weekday="Thursday", lag_days=0) + + assert monday.expected_release_date(date(2025, 1, 4)) == date(2025, 1, 6) + assert thursday.expected_release_date(date(2025, 1, 4)) == date(2025, 1, 9) + + def test_respects_weekday_after_lag_offset(self): + rule = WeeklyRelease(release_weekday="Monday", lag_days=5) + result = rule.expected_release_date(date(2025, 1, 4)) + assert result == date(2025, 1, 13) + class TestCustomRule: def test_serialization(self): diff --git a/tests/test_temporal_semantics_api.py b/tests/test_temporal_semantics_api.py new file mode 100644 index 0000000..23a4c83 --- /dev/null +++ b/tests/test_temporal_semantics_api.py @@ -0,0 +1,26 @@ +"""Tests for the canonical temporal semantics public surface.""" + +from __future__ import annotations + +from datetime import date + +from alphaforge import MissingnessReason, ReleaseRule, classify_missingness +from alphaforge.pit.missingness import MissingnessReason as PitMissingnessReason +from alphaforge.pit.release_rules import ReleaseRule as PitReleaseRule +from alphaforge.time import FixedLagMonths + + +def test_top_level_temporal_semantics_exports_are_canonical() -> None: + assert ReleaseRule is PitReleaseRule + assert MissingnessReason is PitMissingnessReason + + +def test_time_release_rules_drive_core_missingness_classifier() -> None: + result = classify_missingness( + obs_date=date(2025, 3, 31), + asof_date=date(2025, 4, 15), + series_frequency="M", + release_rule=FixedLagMonths(months=2), + ) + + assert result == MissingnessReason.RAGGED_EDGE