The validation package is the part of algosystem the engine work needs most, and its headline statistics are wrong as shipped. This was found on 2026-09-25 and noted in AlgoGators/trade-ngin#113; I am filing it here so it has an owner. Anyone quoting a PSR, DSR or PBO from algosystem should read this first.
1. PSR and DSR mix units
algosystem/validation/domain/statistics/psr_dsr.py annualises the Sharpe (line 106) and then multiplies by sqrt(T - 1) with T in days (lines 209 and 347; also in results.py). The statistic is defined on the per-period Sharpe. So z is inflated by about sqrt(252), and nearly every positive-Sharpe series comes out significant.
Worked numbers (normal returns, skew 0, kurtosis 3):
| case |
correct |
algosystem today |
| PSR, annual Sharpe 0.5, 1,000 days |
0.840 (z 1.00) |
1.000000 (z 14.90) |
| PSR, annual Sharpe 0.5, 2,520 days |
0.943 (z 1.58) |
1.000000 (z 23.66) |
| PSR, annual Sharpe 0.3, 2,520 days |
0.829 |
1.000000 |
| DSR, best annual Sharpe 0.5, 10 trials with Sharpe std 0.2, 2,520 days |
0.721 |
1.0000 (z 8.76) |
How the first row is computed: per-period Sharpe = 0.5 / sqrt(252) = 0.0315; z = 0.0315 x sqrt(999) = 1.00; PSR = Phi(1.00) = 0.840. algosystem puts the annual 0.5 into the same formula: z = 0.5 x sqrt(999) / sqrt(1 + 0.5 x 0.5^2) = 15.80 / 1.061 = 14.90.
Workaround until fixed: call with annualize=1.0 and per-period Sharpes.
The same code is in the monorepo copy (libs/algosystem, lines 106, 208, 343). Fix it in the canonical copy named by the audit issue (#34).
2. The tests pin the wrong values
The shipped tests assert current outputs, so they pass today and would fail on a correct fix. Replace them with known-answer tests: the four rows above, and for compute_pbo a planted overfit case and a dominant-variant case (it has no known-answer test at all).
3. The DSR effective-trial count is a heuristic
rho = 1 - std/range of the trial Sharpes (near line 318) stands in for the correlation or clustering estimate. Either estimate it properly or take the raw trial count as an argument and document which.
4. The detector's deflated_sharpe field is not the DSR
In detect_overfitting.py (near line 169) it is a permutation z-score. Rename it.
5. SignalAnalyzer feeds the wrong series
It computes PSR and DSR on the input market series, not on the strategy's returns (signal_analyzer.py:92-97), and builds the PBO matrix synthetically from scaled block returns (:116-126), which is not the variants' return series. compute_pbo itself is fine when given a real matrix.
6. The walk-forward
It starts each out-of-sample slice cold, drops its purge silently, and is rolling only. For a strategy with a 256-bar warm-up it would score mostly warm-up. At minimum: raise when a requested purge cannot be applied, and document the cold start. Also, monte_carlo_trades reports a degenerate terminal P&L distribution; only its drawdown half carries information.
Acceptance
The validation package is the part of algosystem the engine work needs most, and its headline statistics are wrong as shipped. This was found on 2026-09-25 and noted in AlgoGators/trade-ngin#113; I am filing it here so it has an owner. Anyone quoting a PSR, DSR or PBO from algosystem should read this first.
1. PSR and DSR mix units
algosystem/validation/domain/statistics/psr_dsr.pyannualises the Sharpe (line 106) and then multiplies bysqrt(T - 1)with T in days (lines 209 and 347; also inresults.py). The statistic is defined on the per-period Sharpe. So z is inflated by about sqrt(252), and nearly every positive-Sharpe series comes out significant.Worked numbers (normal returns, skew 0, kurtosis 3):
How the first row is computed: per-period Sharpe = 0.5 / sqrt(252) = 0.0315; z = 0.0315 x sqrt(999) = 1.00; PSR = Phi(1.00) = 0.840. algosystem puts the annual 0.5 into the same formula: z = 0.5 x sqrt(999) / sqrt(1 + 0.5 x 0.5^2) = 15.80 / 1.061 = 14.90.
Workaround until fixed: call with
annualize=1.0and per-period Sharpes.The same code is in the monorepo copy (
libs/algosystem, lines 106, 208, 343). Fix it in the canonical copy named by the audit issue (#34).2. The tests pin the wrong values
The shipped tests assert current outputs, so they pass today and would fail on a correct fix. Replace them with known-answer tests: the four rows above, and for
compute_pboa planted overfit case and a dominant-variant case (it has no known-answer test at all).3. The DSR effective-trial count is a heuristic
rho = 1 - std/rangeof the trial Sharpes (near line 318) stands in for the correlation or clustering estimate. Either estimate it properly or take the raw trial count as an argument and document which.4. The detector's
deflated_sharpefield is not the DSRIn
detect_overfitting.py(near line 169) it is a permutation z-score. Rename it.5.
SignalAnalyzerfeeds the wrong seriesIt computes PSR and DSR on the input market series, not on the strategy's returns (
signal_analyzer.py:92-97), and builds the PBO matrix synthetically from scaled block returns (:116-126), which is not the variants' return series.compute_pboitself is fine when given a real matrix.6. The walk-forward
It starts each out-of-sample slice cold, drops its purge silently, and is rolling only. For a strategy with a 256-bar warm-up it would score mostly warm-up. At minimum: raise when a requested purge cannot be applied, and document the cold start. Also,
monte_carlo_tradesreports a degenerate terminal P&L distribution; only its drawdown half carries information.Acceptance