Add evaluation metrics protocol, PIT primitives, and pipeline infrastructure - #8
Conversation
…ructure Introduces the alphaforge.evaluation package with a runtime-checkable MetricFn protocol and five built-in implementations (RMSE, MAE, DirectionalAccuracy, MAPE, MeanError) for pluggable metric computation. Also adds PIT layer primitives (release rules, vintage resolvers, views, panel builder, missingness detection), pipeline infrastructure (protocols, health checks, weight management, tracking), and supporting modules (config, logging, registry). Includes comprehensive tests and API documentation for all new modules. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
@copilot fix the CI/CD until they succeed. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e2423982e3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| torch.tensor(deriv, dtype=x.dtype).ravel() * grad_output.ravel() | ||
| ).sum() | ||
|
|
||
| return grad_x, None, None, None, None |
There was a problem hiding this comment.
Return gradients for Numpy wrapper parameters
NumpyDifferentiableWrapper registers trainable nn.Parameters, but _NumpyAutograd.backward always returns None for non-input arguments, so those parameters never receive gradients during loss.backward(). In any training loop that tries to optimize param_names (the advertised use case for this module), optimizer steps become no-ops for wrapper parameters and tuning silently fails.
Useful? React with 👍 / 👎.
| "SELECT MAX(obs_date) FROM pit_observations WHERE source = ?", | ||
| [source_name], |
There was a problem hiding this comment.
Restrict health lookup to data available at as-of
assess(..., asof=...) currently calls a query that takes MAX(obs_date) over all rows for the source, without filtering by asof_utc <= asof. When historical health is computed after newer observations have already been ingested, this will pull future observations into past assessments, causing look-ahead leakage and overly optimistic source-health weights in backtests/replays.
Useful? React with 👍 / 👎.
Co-authored-by: steveya <7390489+steveya@users.noreply.github.com> Agent-Logs-Url: https://github.com/steveya/alphaforge/sessions/0aec6e90-fcdb-4dec-b86a-ee5298e82006
…-doc-sync guard Co-authored-by: steveya <7390489+steveya@users.noreply.github.com> Agent-Logs-Url: https://github.com/steveya/alphaforge/sessions/fbc4346e-186b-475f-b68d-321ce09b1cf7
Fix CI: resolve ruff and mypy failures introduced by evaluation/PIT/pipeline PR
Summary
alphaforge.evaluationpackage with runtime-checkableMetricFnprotocol and 5 built-in metric implementations (RMSE, MAE, DirectionalAccuracy, MAPE, MeanError)QuarterlyRelease), vintage resolvers (RealtimeResolver,LatestResolver,FrozenResolver), views, panel builder, and missingness detectionTest plan
pytest tests/)MetricFnprotocol verified withisinstancechecks🤖 Generated with Claude Code