Skip to content

PSR and DSR are inflated by about sqrt(252), and four related defects in the validation package #35

Description

@raohemdutt

The validation package is the part of algosystem the engine work needs most, and its headline statistics are wrong as shipped. This was found on 2026-09-25 and noted in AlgoGators/trade-ngin#113; I am filing it here so it has an owner. Anyone quoting a PSR, DSR or PBO from algosystem should read this first.

1. PSR and DSR mix units

algosystem/validation/domain/statistics/psr_dsr.py annualises the Sharpe (line 106) and then multiplies by sqrt(T - 1) with T in days (lines 209 and 347; also in results.py). The statistic is defined on the per-period Sharpe. So z is inflated by about sqrt(252), and nearly every positive-Sharpe series comes out significant.

Worked numbers (normal returns, skew 0, kurtosis 3):

case correct algosystem today
PSR, annual Sharpe 0.5, 1,000 days 0.840 (z 1.00) 1.000000 (z 14.90)
PSR, annual Sharpe 0.5, 2,520 days 0.943 (z 1.58) 1.000000 (z 23.66)
PSR, annual Sharpe 0.3, 2,520 days 0.829 1.000000
DSR, best annual Sharpe 0.5, 10 trials with Sharpe std 0.2, 2,520 days 0.721 1.0000 (z 8.76)

How the first row is computed: per-period Sharpe = 0.5 / sqrt(252) = 0.0315; z = 0.0315 x sqrt(999) = 1.00; PSR = Phi(1.00) = 0.840. algosystem puts the annual 0.5 into the same formula: z = 0.5 x sqrt(999) / sqrt(1 + 0.5 x 0.5^2) = 15.80 / 1.061 = 14.90.

Workaround until fixed: call with annualize=1.0 and per-period Sharpes.

The same code is in the monorepo copy (libs/algosystem, lines 106, 208, 343). Fix it in the canonical copy named by the audit issue (#34).

2. The tests pin the wrong values

The shipped tests assert current outputs, so they pass today and would fail on a correct fix. Replace them with known-answer tests: the four rows above, and for compute_pbo a planted overfit case and a dominant-variant case (it has no known-answer test at all).

3. The DSR effective-trial count is a heuristic

rho = 1 - std/range of the trial Sharpes (near line 318) stands in for the correlation or clustering estimate. Either estimate it properly or take the raw trial count as an argument and document which.

4. The detector's deflated_sharpe field is not the DSR

In detect_overfitting.py (near line 169) it is a permutation z-score. Rename it.

5. SignalAnalyzer feeds the wrong series

It computes PSR and DSR on the input market series, not on the strategy's returns (signal_analyzer.py:92-97), and builds the PBO matrix synthetically from scaled block returns (:116-126), which is not the variants' return series. compute_pbo itself is fine when given a real matrix.

6. The walk-forward

It starts each out-of-sample slice cold, drops its purge silently, and is rolling only. For a strategy with a 256-bar warm-up it would score mostly warm-up. At minimum: raise when a requested purge cannot be applied, and document the cold start. Also, monte_carlo_trades reports a degenerate terminal P&L distribution; only its drawdown half carries information.

Acceptance

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions