Skip to content

Latest commit

 

History

416 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LOB Arena

Governed Historical + Synthetic Order-Book Validation for Market Surveillance

LOB Arena banner

A research and validation platform that replays licensed historical order-book data, injects controlled synthetic attacks, and benchmarks surveillance detectors against reproducible ground truth.

Problem

Historical order flow rarely supplies reliable manipulation labels. Detector validation needs repeatable scenarios, separate ground truth and inspectable evidence.

Solution

LOB Arena combines immutable historical replay, bounded synthetic overlays and verified detector comparisons. Rules produce incidents; AI explains their evidence. Historical activity is never automatically labeled benign or abusive. Outputs are neither trading signals nor compliance decisions.

Task Guide
Run the application Quickstart
Explore or build the static public website Website and Pages guide
Understand ownership and data flow Architecture
Check delivered versus planned work Current status
Choose a workflow Use cases
Train, evaluate or inspect learned models ML lifecycle
Find a detailed specification Documentation index

Architecture

Java owns the exchange, live REST/WebSocket controls, replay and agent orchestration. Python owns ingestion, offline ML, AI and cloud integration. Agents return bounded intents; they cannot mutate the exchange. MLflow indexes verified artifacts and never grants release approval. Prometheus/Grafana are optional read-only diagnostics.

The canonical architecture diagram shows these boundaries. Historical publication posters are not current design references.

Screenshots

See the sanitized screenshot gallery for historical demo captures; use the architecture above for current ownership.

Quick start

git clone https://github.com/khab40/lob-arena.git
cd lob-arena
cp .env.example .env
docker compose up --build

Open the UI at http://localhost:5173, Python AI API at http://localhost:8000, or Java status at http://localhost:8081/api/kernel/status. The same-origin arena WebSocket is ws://localhost:5173/ws/arena.

Default Compose builds Java, Python, the agent runner and frontend from source. Local Mock needs no Nebius credentials or GPU. Serverless access is disabled; the backend clears stale cloud settings unless explicitly enabled. See the quickstart for prerequisites, configuration and troubleshooting.

Agent-initiated training, scoring, model fixtures and frozen-runtime validation run on Nebius Serverless Jobs under the execution policy. Local orchestration, static checks and artifact inspection remain available.

Historical and hybrid replay

Use the replay quickstart for LOBSTER/ITCH imports, control-versus-hybrid comparison, signatures and market-profile commands. The client validation runbook owns delivery checks; hybrid validation owns equivalence tests and trust boundaries.

Model-ready causal features

The feature reference owns formulas, configuration, label isolation and causal prefix guarantees. The LightGBM runbook owns commands and artifacts; the ML lifecycle distinguishes implemented training and sequence materialization from planned Transformer/cascade/serving work. Software completion or synthetic recovery does not establish production quality.

Observability and cloud setup

Use kernel observability for the Prometheus/Grafana profiles and dashboards, MLflow operations for tracking, and Nebius deployment for opt-in endpoint/job configuration. Missing submit templates must remain explicitly pending; configuration examples are not execution authorization.

Automated grader

From a fresh checkout of the default main branch, run exactly:

make grader-smoke

This credential-free command installs locked dependencies when needed, launches the backend and frontend on local ephemeral ports, submits one fixed-seed Local Mock scenario, and validates backend health, the rendered frontend, detector output, results metrics, event data, and all eight artifacts. It uses temporary output and prints GRADER_OK only after every check succeeds. It does not require Docker, cloud credentials, a GPU, or access to Nebius services.

Demo

Video walkthrough and demo script cover the historical challenge demo.

Evidence

The submission index links measured runtime/cost and frozen evidence. Those observations are not current estimates or learned-model qualification.

Freeze a new local evidence snapshot with ./scripts/freeze-release.sh; add --offline when Docker, the backend, or Nebius CLI is unavailable.

Development

CI validates retained Python tests and Ruff, frontend lint/build, the authoritative Java 25 kernel and live control plane, deterministic CPU evaluation, agent workspace contracts, Compose config, application Docker images, and Gitleaks. It intentionally does not build long-running Nebius Endpoint/Job images and does not run GPU/vLLM inference.

The commands below describe developer checks. Agents must apply the execution policy above before running tests that train, score or exercise model runtime:

uv sync --project backend --dev --frozen
PYTHONPATH=. uv run --project backend ruff check backend serverless scripts
PYTHONPATH=. uv run --project backend pytest -c backend/pyproject.toml backend/tests
(cd frontend && corepack enable && pnpm install --frozen-lockfile && pnpm run lint && pnpm run build)
(cd java && ./gradlew clean check)
docker compose --env-file .env.example config --quiet
./scripts/check-secrets.sh

Common dev commands:

make grader-smoke
make backend-dev
make frontend-dev
make backend-test
make serverless-benchmark
make secrets-plan
make secrets-check

Contributing

Keep local fallback explicit. Never commit credentials, private endpoints, signed URLs or unredacted cloud logs; never print .env. Use docker compose config --quiet and ./scripts/check-secrets.sh. Follow the documentation ownership rules.

About

A multi-agent platform that generates realistic synthetic limit-order-book activity and benchmarks market-surveillance systems against adaptive manipulation strategies.

Resources

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages