Skip to content

Latest commit

 

History

History
83 lines (42 loc) · 4.85 KB

File metadata and controls

83 lines (42 loc) · 4.85 KB

Engineering Decisions

This document records the major architectural choices in the reusable Project Ledger template. It is intentionally concise. Production teams should convert decisions that materially affect their environment into formal ADRs with dates, owners, alternatives, and review status.

Decision 1: Use fixture-first development

Decision: The ingestion pipeline must run from committed SEC Company Facts fixtures as well as the live API.

Why: Deterministic fixtures make local development and CI independent of network access, API availability, and rate limits.

Trade-off: Fixtures can drift from current source behavior. The project should periodically refresh them and retain a separate live-integration validation lane.

Decision 2: Keep live and fixture clients behind one protocol

Decision: Pipeline code depends on a CompanyFactsClient boundary rather than a concrete HTTP implementation.

Why: The ingestion and normalization path can be exercised without conditional logic scattered throughout the pipeline.

Trade-off: The abstraction must remain narrow. Adding source-specific behavior to the protocol would reduce substitutability.

Decision 3: Preserve raw payloads before normalization

Decision: Store the complete source payload separately from normalized analytical observations.

Why: Raw preservation supports replay, debugging, auditability, and future transformation changes.

Trade-off: Raw storage increases retention and governance requirements. Production deployments must define lifecycle, classification, and access policies.

Decision 4: Define issuers through an allowlist

Decision: Process only explicitly configured issuer names and CIKs.

Why: An allowlist makes scope intentional, reduces unexpected volume, and creates a validation boundary between configuration and source payloads.

Trade-off: Expanding coverage requires a configuration change. Large-scale discovery would need a governed issuer registry rather than a static file.

Decision 5: Normalize to an observation-level analytical grain

Decision: Emit individual financial observations with issuer, taxonomy, concept, unit, period, filing, form, and value attributes.

Why: A narrow fact grain supports flexible downstream modeling while preserving the filing context required for deduplication.

Trade-off: XBRL semantics are complex. Production marts may need concept mapping, dimensional context, restatement rules, and company-specific exceptions.

Decision 6: Make repeated runs replace governed local outputs

Decision: Write raw and normalized files atomically and replace the governed output for the same configured run scope.

Why: This provides simple repeatability and avoids accidental append-only duplication in the reference implementation.

Trade-off: Replacement is not a complete production backfill strategy. Cloud implementations should use partition-aware loads, merge semantics, checkpoints, and immutable raw landing paths.

Decision 7: Separate application, analytics, and infrastructure validation

Decision: CI validates Python, the fixture pipeline, dbt, and Terraform in one delivery gate.

Why: A data platform can fail outside application unit tests. Cross-layer validation exposes contract and packaging errors earlier.

Trade-off: Validation time and dependency count increase. Production repositories may split fast pull-request checks from slower integration workflows.

Decision 8: Use DuckDB locally and keep cloud storage replaceable

Decision: Use DuckDB for the runnable local analytical path while documenting GCS and BigQuery as production adapters.

Why: DuckDB minimizes setup cost and makes the complete template executable by another engineer.

Trade-off: Local execution does not prove distributed scale, cloud IAM, or BigQuery cost behavior. Those require separate deployment evidence.

Decision 9: Parameterize infrastructure and keep secrets external

Decision: Cloud names and identifiers are Terraform variables; runtime contact information is supplied through environment variables or a secret reference.

Why: Code should not encode a person's identity, a private environment, or credentials.

Trade-off: Deployments require an explicit configuration process and external secret provisioning.

Decision 10: Generalize the implementation rather than publish operational history

Decision: Exclude personal sprint journals, private identifiers, generated operational evidence, and environment-specific artifacts from the reusable repository.

Why: The public artifact should demonstrate transferable engineering without leaking context or overstating what is deployed in the template.

Trade-off: The repository alone provides less evidence of long-running operation. The case study documents scope and limitations without exposing private material.