This is direction and a starting design, not a finished spec. It starts when the audit issue (#34) has named the canonical copy.
What I want: someone runs a backtest in trade-ngin, then runs one algosystem command naming that run and a benchmark, and gets the tearsheet with the benchmark-relative numbers: correlation, beta, alpha, tracking error, information ratio, up and down capture. If the benchmark they want is not there, they add it. The same works for a live book.
Task list
What exists today
algosystem backtest and algosystem tearsheet take a CSV or a series and --benchmark <alias or CSV file>. The metrics include alpha, beta, correlation, tracking error, information ratio and up and down capture, through quantstats.
- The benchmark catalogue (
algosystem/marketdata/domain/benchmark.py) is a fixed list in code, default sp500: equity indices, Treasury yields, a few ETFs, international indices. Prices come from yfinance and are cached to parquet.
- Nothing reads the engine. The engine stores
backtest.equity_curve (run_id, portfolio_id, timestamp, equity) and trading.equity_curve for the live books.
- algosystem's own Postgres repository can write and delete, and models a
backtest schema with the engine's table names and different columns. It must never be pointed at the engine's database.
The work
W1. The reader. A new, separate, read-only adapter: given a run_id (and portfolio_id where a run has several), or a live portfolio_id and a date range, return the (timestamp, equity) series. It issues SELECT only, uses a database role that has no write grant, and shares no code path with the existing repository's save, delete or create_all. EquityCurve.from_series() already accepts the result.
Warm-up rows. The engine stores the warm-up period flat at initial capital (256 bars for the conservative futures book). Scored against a benchmark, those rows drag every number toward zero, so the reader drops them. AlgoGators/trade-ngin#144 adds a phase flag to each curve row; use it once it lands. Until then, drop the leading run of rows equal to initial capital and print how many were dropped.
W2. The commands. For example algosystem tearsheet --run-id <id> --benchmark dbmf, and the same from Python. The output states the run, the portfolio, the window after warm-up, the benchmark and its source.
W3. Benchmarks.
| need |
design to start from |
| the benchmarks a trend-following futures book is judged against |
add managed-futures funds to the catalogue (DBMF, KMLM and AQMIX were the ones used in the 2026-10-05 comparison) and say what a non-investable index such as SG Trend needs (a CSV; it has no ticker) |
| add one without editing library code |
a user catalogue file (alias, ticker or CSV path, description, category) merged over the built-in list; algosystem benchmarks add writes it |
| a benchmark that is not a price |
the Treasury-yield entries are yields, not total-return series. Either convert or refuse them as benchmarks with a clear message |
| reproducibility |
yfinance revises history. The tearsheet records the fetch date; a pinned CSV is the way to freeze a comparison |
Benchmarks live on the algosystem side. The engine does not fetch outside prices and stores no benchmark. If AlgoLens later needs the same benchmark series for the desk, that is a table loaded by data-ngin and a separate request to me.
W4. Conventions. One page: return type (simple, daily), annualisation (252), how the two series are aligned (dates both have; what happens to a day only one has), risk-free rate, and the formula for each relative metric. Include a worked example small enough to check by hand: a table of five days of book equity and benchmark price, the two return columns, the active return column, and the tracking error and information ratio computed from them. The engine's own Sharpe and drawdown in backtest.results use the engine's conventions; say where algosystem's differ and by how much on one real run. Do not change either to match.
W5. Tests. Against a scratch copy with seeded rows: a run with warm-up, a run with two portfolios, a live book, a benchmark with missing days, an unknown alias. One test asserts the reader's role cannot write.
Never do
- No write of any kind to the engine's database.
- No recomputation that is then presented as the engine's number. The tearsheet is algosystem's view and says so.
Done when
On a scratch copy, a conservative futures backtest run is compared with DBMF and with a user-added benchmark, and the worked example in W4 matches the command's output.
This is direction and a starting design, not a finished spec. It starts when the audit issue (#34) has named the canonical copy.
What I want: someone runs a backtest in trade-ngin, then runs one algosystem command naming that run and a benchmark, and gets the tearsheet with the benchmark-relative numbers: correlation, beta, alpha, tracking error, information ratio, up and down capture. If the benchmark they want is not there, they add it. The same works for a live book.
Task list
tearsheet,backtestand the Python API accept an engine runWhat exists today
algosystem backtestandalgosystem tearsheettake a CSV or a series and--benchmark <alias or CSV file>. The metrics include alpha, beta, correlation, tracking error, information ratio and up and down capture, through quantstats.algosystem/marketdata/domain/benchmark.py) is a fixed list in code, defaultsp500: equity indices, Treasury yields, a few ETFs, international indices. Prices come from yfinance and are cached to parquet.backtest.equity_curve (run_id, portfolio_id, timestamp, equity)andtrading.equity_curvefor the live books.backtestschema with the engine's table names and different columns. It must never be pointed at the engine's database.The work
W1. The reader. A new, separate, read-only adapter: given a
run_id(andportfolio_idwhere a run has several), or a liveportfolio_idand a date range, return the(timestamp, equity)series. It issuesSELECTonly, uses a database role that has no write grant, and shares no code path with the existing repository'ssave,deleteorcreate_all.EquityCurve.from_series()already accepts the result.Warm-up rows. The engine stores the warm-up period flat at initial capital (256 bars for the conservative futures book). Scored against a benchmark, those rows drag every number toward zero, so the reader drops them. AlgoGators/trade-ngin#144 adds a phase flag to each curve row; use it once it lands. Until then, drop the leading run of rows equal to initial capital and print how many were dropped.
W2. The commands. For example
algosystem tearsheet --run-id <id> --benchmark dbmf, and the same from Python. The output states the run, the portfolio, the window after warm-up, the benchmark and its source.W3. Benchmarks.
algosystem benchmarks addwrites itBenchmarks live on the algosystem side. The engine does not fetch outside prices and stores no benchmark. If AlgoLens later needs the same benchmark series for the desk, that is a table loaded by data-ngin and a separate request to me.
W4. Conventions. One page: return type (simple, daily), annualisation (252), how the two series are aligned (dates both have; what happens to a day only one has), risk-free rate, and the formula for each relative metric. Include a worked example small enough to check by hand: a table of five days of book equity and benchmark price, the two return columns, the active return column, and the tracking error and information ratio computed from them. The engine's own Sharpe and drawdown in
backtest.resultsuse the engine's conventions; say where algosystem's differ and by how much on one real run. Do not change either to match.W5. Tests. Against a scratch copy with seeded rows: a run with warm-up, a run with two portfolios, a live book, a benchmark with missing days, an unknown alias. One test asserts the reader's role cannot write.
Never do
Done when
On a scratch copy, a conservative futures backtest run is compared with DBMF and with a user-added benchmark, and the worked example in W4 matches the command's output.