Project Ledger began as a production-oriented learning implementation for ingesting, preserving, transforming, and serving public SEC/XBRL financial data. The implementation was later generalized into this organization-neutral reference repository so another engineer can study, run, and extend the architecture without inheriting personal infrastructure, credentials, or project-specific history.
This repository is therefore not a byte-for-byte export of an operating environment. It is the reusable engineering core extracted from a broader system.
Build a financial-data pipeline that could move from deterministic local development to managed cloud execution without changing its core data contracts.
The system needed to:
- support deterministic development without depending on a live API
- preserve source payloads before transformation
- normalize semi-structured XBRL observations into a stable analytical grain
- prevent duplicate analytical records across repeated runs
- separate configuration from code and identity
- validate code, models, infrastructure, and pipeline behavior in CI
- expose clear extension points for orchestration, storage, governance, and observability
The reusable template includes:
- Python-based SEC Company Facts ingestion
- live and fixture-backed clients behind a common protocol
- governed issuer configuration and CIK normalization
- raw JSON preservation
- normalized JSONL financial observations
- atomic and repeatable local writes
- dbt staging and deduplicated fact models in DuckDB
- unit and pipeline tests
- strict type checking and linting
- Docker packaging
- an Airflow DAG example
- parameterized GCP infrastructure for Cloud Run Jobs, GCS, BigQuery, IAM, and Secret Manager references
- GitHub Actions validation across application, analytics, and infrastructure layers
The broader implementation from which this template was extracted also explored managed GCP execution, BigQuery serving, Airflow or Composer orchestration, workload identity, monitoring, market-price enrichment, machine-learning extensions, language-model-assisted analysis, and earnings-event processing. Those later extensions are described as future integration patterns rather than presented here as completed template features.
Depending entirely on live SEC requests would make tests slower, less deterministic, and more vulnerable to rate limits or network failures. The solution was a fixture-first client boundary. The same pipeline can consume committed payloads in CI or a rate-limited live client in an authorized runtime.
Transforming source payloads in place would remove evidence needed for debugging and future reprocessing. The pipeline therefore writes the original issuer payload separately from normalized observations.
Financial facts may contain repeated observations across forms and filing periods. The pipeline produces governed outputs atomically, while dbt applies a documented analytical grain and deduplication rule before exposing the mart.
Personal names, cloud identifiers, issuer selections, paths, and runtime contacts were moved into configuration or Terraform variables. This allows the implementation to be reused without editing application logic.
A passing unit-test suite alone would not establish readiness. CI validates Python formatting and typing, tests, a fixture pipeline smoke run, dbt models and data tests, and Terraform formatting and syntax.
The design explicitly accounts for:
- invalid or mismatched issuer identifiers
- malformed Company Facts payloads
- API rate limits and transient request failures
- missing environment variables
- partial writes
- duplicate analytical observations
- model contract drift
- unsafe hard-coded infrastructure values
- production dependencies leaking into deterministic CI
This project provides evidence of competence in:
- ingestion boundary design
- semi-structured data normalization
- analytical data modeling
- idempotency and reproducibility
- configuration and secrets separation
- automated testing and delivery gates
- orchestration design
- containerization
- infrastructure as code
- converting a project-specific system into a reusable engineering template
The repository is a reference implementation, not a hosted production service. It does not include:
- a live deployed environment
- service-level objectives or production alert routing
- backfill coordination across multiple workers
- a complete BigQuery loading implementation
- full schema-evolution automation
- cost benchmarks at production volume
- generated dashboards, runtime logs, or private operational evidence
These exclusions are deliberate. They prevent the public template from implying deployment evidence it does not contain.
The most valuable extensions would be:
- implement storage adapters for GCS and BigQuery
- add explicit schema contracts and migration checks
- add freshness, volume, error-rate, and cost observability
- add integration tests against an ephemeral cloud environment
- add backfill and replay controls
- publish synthetic-data dashboards over the governed mart
- document benchmark results and operational recovery exercises