Skip to content

Add evaluation metrics protocol, PIT primitives, and pipeline infrastructure - #8

Merged
steveya merged 5 commits into
mainfrom
feature/evaluation-layer-and-pit-primitives
Mar 22, 2026
Merged

steveya merged 5 commits into
mainfrom
feature/evaluation-layer-and-pit-primitives

Conversation

@steveya

@steveya steveya commented Mar 22, 2026

Copy link
Copy Markdown
Owner

Summary

  • Introduces alphaforge.evaluation package with runtime-checkable MetricFn protocol and 5 built-in metric implementations (RMSE, MAE, DirectionalAccuracy, MAPE, MeanError)
  • Adds PIT layer primitives: release rules (QuarterlyRelease), vintage resolvers (RealtimeResolver, LatestResolver, FrozenResolver), views, panel builder, and missingness detection
  • Adds pipeline infrastructure: protocols, health checks, weight management, tracking
  • Adds supporting modules: config, logging, registry
  • Comprehensive tests and API documentation for all new modules

Test plan

  • All new tests pass (pytest tests/)
  • Existing tests unaffected
  • MetricFn protocol verified with isinstance checks
  • Release rules verified against known BEA publication dates
  • Vintage resolvers verified with synthetic PIT data

🤖 Generated with Claude Code

…ructure

Introduces the alphaforge.evaluation package with a runtime-checkable
MetricFn protocol and five built-in implementations (RMSE, MAE,
DirectionalAccuracy, MAPE, MeanError) for pluggable metric computation.

Also adds PIT layer primitives (release rules, vintage resolvers, views,
panel builder, missingness detection), pipeline infrastructure (protocols,
health checks, weight management, tracking), and supporting modules
(config, logging, registry).

Includes comprehensive tests and API documentation for all new modules.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@steveya

steveya commented Mar 22, 2026

Copy link
Copy Markdown
Owner Author

@copilot fix the CI/CD until they succeed.

Copilot AI commented Mar 22, 2026

Copy link
Copy Markdown
Contributor

@steveya I've opened a new pull request, #9, to work on those changes. Once the pull request is ready, I'll request review from you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e2423982e3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

torch.tensor(deriv, dtype=x.dtype).ravel() * grad_output.ravel()
).sum()

return grad_x, None, None, None, None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Return gradients for Numpy wrapper parameters

NumpyDifferentiableWrapper registers trainable nn.Parameters, but _NumpyAutograd.backward always returns None for non-input arguments, so those parameters never receive gradients during loss.backward(). In any training loop that tries to optimize param_names (the advertised use case for this module), optimizer steps become no-ops for wrapper parameters and tuning silently fails.

Useful? React with 👍 / 👎.

Comment on lines +34 to +35
"SELECT MAX(obs_date) FROM pit_observations WHERE source = ?",
[source_name],

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Restrict health lookup to data available at as-of

assess(..., asof=...) currently calls a query that takes MAX(obs_date) over all rows for the source, without filtering by asof_utc <= asof. When historical health is computed after newer observations have already been ingested, this will pull future observations into past assessments, causing look-ahead leakage and overly optimistic source-health weights in backtests/replays.

Useful? React with 👍 / 👎.

Copilot AI and others added 3 commits March 22, 2026 01:06
Fix CI: resolve ruff and mypy failures introduced by evaluation/PIT/pipeline PR
@steveya
steveya merged commit eef73e6 into main Mar 22, 2026
7 checks passed
@steveya
steveya deleted the feature/evaluation-layer-and-pit-primitives branch March 22, 2026 01:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants