Governed Historical + Synthetic Order-Book Validation for Market Surveillance
A research and validation platform that replays licensed historical order-book data, injects controlled synthetic attacks, and benchmarks surveillance detectors against reproducible ground truth.
Historical order flow rarely supplies reliable manipulation labels. Detector validation needs repeatable scenarios, separate ground truth and inspectable evidence.
LOB Arena combines immutable historical replay, bounded synthetic overlays and verified detector comparisons. Rules produce incidents; AI explains their evidence. Historical activity is never automatically labeled benign or abusive. Outputs are neither trading signals nor compliance decisions.
| Task | Guide |
|---|---|
| Run the application | Quickstart |
| Explore or build the static public website | Website and Pages guide |
| Understand ownership and data flow | Architecture |
| Check delivered versus planned work | Current status |
| Choose a workflow | Use cases |
| Train, evaluate or inspect learned models | ML lifecycle |
| Find a detailed specification | Documentation index |
Java owns the exchange, live REST/WebSocket controls, replay and agent orchestration. Python owns ingestion, offline ML, AI and cloud integration. Agents return bounded intents; they cannot mutate the exchange. MLflow indexes verified artifacts and never grants release approval. Prometheus/Grafana are optional read-only diagnostics.
The canonical architecture diagram shows these boundaries. Historical publication posters are not current design references.
See the sanitized screenshot gallery for historical demo captures; use the architecture above for current ownership.
git clone https://github.com/khab40/lob-arena.git
cd lob-arena
cp .env.example .env
docker compose up --buildOpen the UI at http://localhost:5173, Python AI API at http://localhost:8000,
or Java status at http://localhost:8081/api/kernel/status. The same-origin arena
WebSocket is ws://localhost:5173/ws/arena.
Default Compose builds Java, Python, the agent runner and frontend from source. Local Mock needs no Nebius credentials or GPU. Serverless access is disabled; the backend clears stale cloud settings unless explicitly enabled. See the quickstart for prerequisites, configuration and troubleshooting.
Agent-initiated training, scoring, model fixtures and frozen-runtime validation run on Nebius Serverless Jobs under the execution policy. Local orchestration, static checks and artifact inspection remain available.
Use the replay quickstart for LOBSTER/ITCH imports, control-versus-hybrid comparison, signatures and market-profile commands. The client validation runbook owns delivery checks; hybrid validation owns equivalence tests and trust boundaries.
The feature reference owns formulas, configuration, label isolation and causal prefix guarantees. The LightGBM runbook owns commands and artifacts; the ML lifecycle distinguishes implemented training and sequence materialization from planned Transformer/cascade/serving work. Software completion or synthetic recovery does not establish production quality.
Use kernel observability for the Prometheus/Grafana profiles and dashboards, MLflow operations for tracking, and Nebius deployment for opt-in endpoint/job configuration. Missing submit templates must remain explicitly pending; configuration examples are not execution authorization.
From a fresh checkout of the default main branch, run exactly:
make grader-smokeThis credential-free command installs locked dependencies when needed, launches the backend and frontend on local ephemeral ports, submits one fixed-seed Local Mock scenario, and validates backend health, the rendered frontend, detector output, results metrics, event data, and all eight artifacts. It uses temporary output and prints GRADER_OK only after every check succeeds. It does not require Docker, cloud credentials, a GPU, or access to Nebius services.
Video walkthrough and demo script cover the historical challenge demo.
The submission index links measured runtime/cost and frozen evidence. Those observations are not current estimates or learned-model qualification.
- Challenge submission index
- Manual Nebius Control Panel evidence (100-workload Job + 12 real Endpoint calls)
- Six-job production E2E evidence (1,200 workloads)
- Production L40S/vLLM Endpoint evidence (25 real calls)
- Representative scenario benchmark
- Frozen benchmark bundle
- Frozen Nebius deployment bundle
Freeze a new local evidence snapshot with ./scripts/freeze-release.sh; add --offline when Docker, the backend, or Nebius CLI is unavailable.
CI validates retained Python tests and Ruff, frontend lint/build, the authoritative Java 25 kernel and live control plane, deterministic CPU evaluation, agent workspace contracts, Compose config, application Docker images, and Gitleaks. It intentionally does not build long-running Nebius Endpoint/Job images and does not run GPU/vLLM inference.
The commands below describe developer checks. Agents must apply the execution policy above before running tests that train, score or exercise model runtime:
uv sync --project backend --dev --frozen
PYTHONPATH=. uv run --project backend ruff check backend serverless scripts
PYTHONPATH=. uv run --project backend pytest -c backend/pyproject.toml backend/tests
(cd frontend && corepack enable && pnpm install --frozen-lockfile && pnpm run lint && pnpm run build)
(cd java && ./gradlew clean check)
docker compose --env-file .env.example config --quiet
./scripts/check-secrets.shCommon dev commands:
make grader-smoke
make backend-dev
make frontend-dev
make backend-test
make serverless-benchmark
make secrets-plan
make secrets-checkKeep local fallback explicit. Never commit credentials, private endpoints,
signed URLs or unredacted cloud logs; never print .env.
Use docker compose config --quiet and ./scripts/check-secrets.sh.
Follow the documentation ownership rules.
