Skip to content

feat(netflow-db): make internal-side MAAD optional - #127

Merged
flamboh merged 2 commits into
mainfrom
feat/optional-internal-maad
Oct 1, 2026
Merged

flamboh merged 2 commits into
mainfrom
feat/optional-internal-maad

Conversation

@flamboh

@flamboh flamboh commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Note

🤖 Claude Opus 5.5 on behalf of Oliver

Closes #121.

Problem

MAAD on internal-side address sets isn't well defined, and computing it costs time and storage. Internal-side sets are the source addresses of internal-source traffic and the destination addresses of internal-destination traffic. The decision on #121: make these sets optional everywhere, and skip them by default.

What changes

Pipeline (tools/netflow-db)

  • New setting maad_internal_side, default false. Set it on a datasets.json entry, or at the top level of a pipeline config. A config's own datasets[] entries can't set it, which matches how locality works.
  • Skipped sets aren't computed at all. publish::write_buckets takes a MaadScopes value (None / ExceptInternalSide / All) and filters address sets before the MAAD worker pool runs. This applies to all three measures and to rollups. Address-count rows are unchanged, and all/all scopes are still computed.
  • Product identity: a skipping product adds "internal_side": false inside the maad result config. A product that computes every set keeps exactly the identity it had before. So:
    • A database built with one setting rejects runs with the other.
    • Coordinated --dataset runs must agree on the setting.
    • Existing databases can still be extended if their registry entry sets maad_internal_side: true.
  • datasets.maad_internal_side column: mirrors the setting for the dashboard, which can't read pipeline_product on D1. The pipeline adds the column (default 1) when it opens a pre-existing database.
  • verify reads the setting from product identity. It fails when the column disagrees or when a skipping product stores internal-side rows. With --require-maad-data, it now requires a MAAD row for every address set the product computes, so internal-side rows are required when opted in.
  • compare reports reference rows for skipped sets as skipped_reference_rows rather than reference-only rows. Candidate rows for skipped sets count as unexpected.
  • merge-shards refuses a skipping shard that stores internal-side rows. Mixed settings were already refused through product identity.

Scopes computed in default mode (per IP version and bucket)

Dataset Computed address sets
Normal (with locality) all→all source + destination, external→external source + destination, external→internal source, internal→external destination (6 of 10)
daily_active_sources (selection fixes src_locality = internal) All its traffic is internal→internal or internal→external. Useful MAAD: all→all source + destination, and internal→external destination (external peers). The external→* scopes are still computed but are always empty for this selection.

Dashboard (apps/web)

  • Drizzle migration 20261001083148_maad_internal_side adds datasets.maad_internal_side (DEFAULT true, so existing D1 rows read as "computed").
  • The local SQLite driver reads the column when present and treats legacy databases as 1.
  • /api/netflow/maad-status now returns { computed, internalSide }. The dataset page and the file page load it.
  • MAAD Dimensions and IP Address Spectrum cards:
    • When the selected address side is internal, the card shows "MAAD is not computed for internal addresses…" inside the chart area. The side and source controls stay usable.
    • Lateral traffic has both sides internal, so the whole card shows that message.
  • The file view shows the same message in place of that side's structure and spectrum panes.

Screenshots of both states were taken from the Playwright fixture: ingress with Destination selected, and lateral. They aren't embedded because the evidence upload tool wasn't available in this environment. tests/e2e/maad-internal-side.spec.ts asserts the exact copy for both states.

Review guide

Flows to exercise

  1. Build a CSV or nfcapd product with locality rules and no maad_internal_side. Then run verify --require-data --require-maad-data. It should pass, and address_maad_stats should have no internal source-side or destination-side rows.
  2. Set "maad_internal_side": true and rebuild to a new path. You should get all 10 scope/side combinations, and verify should pass.
  3. Point either config at the other database. It should fail with a product identity (config) conflict.
  4. Open the dashboard on the skip-mode database. Check ingress → Destination, egress → Source, lateral (whole card), transit/all (normal charts). Also check a file page with direction=egress.

Setup / test data: everything uses fixtures. The e2e fixture adds a second dataset row, playwright-external-maad, with maad_internal_side = 0 in the same Playwright database.

Decisions needing review

  • Identity is unchanged for opt-in products. Only the skipping product records internal_side. This keeps currently deployed databases, which computed every set, extendable by setting maad_internal_side: true on their registry entries. If you'd rather force a fresh product for every database, add the key unconditionally.
  • Existing registries switch to skip by default. A live feed whose entry doesn't set maad_internal_side: true will fail its next run on an identity conflict rather than mixing results. Add "maad_internal_side": true to keep extending an existing database.
  • Source of truth: Rust tooling treats product identity as authoritative. The datasets column is a mirror for the web, and verify checks the two agree. The column defaults to 1 everywhere (Rust DDL, legacy upgrade, Drizzle/D1), and the pipeline always writes it explicitly. merge-shards compares the datasets table by its columns rather than its raw CREATE TABLE text, so a shard upgraded in place merges with a fresh one.
  • --require-maad-data is stricter. It now runs an anti-join from address_count_stats to address_maad_stats (indexed lookups), so it costs more on large databases.
  • all/all scopes stay computed in both modes. For daily_active_sources, the all→all source side is effectively all internal addresses. Removing it would be a separate decision.

Verification

Automated:

  • bun run format, bun run lint, bun run typecheck: pass
  • bun run test:db: pass (219 tests)
  • bun run test:web: pass (212 tests)
  • bun run test:e2e: pass (26 tests)

New tests:

  • publish scope filtering
  • storage legacy-column upgrade and identity read
  • verify scope checks
  • a pipeline CLI end-to-end test (skip, opt-in, --no-maad, identity refusal, column mismatch, stray internal rows)
  • compare skip accounting
  • coordinated-run mismatch
  • web API, page-load, driver and helper tests
  • three e2e flows

Manual checks remaining:

  • Run a real-size build in skip mode and compare wall time and size against an opt-in build.
  • Apply the D1 migration on a staging stage.

Made with Claude Opus 5.5 in Claude Code.

Skip MAAD for internal-side address sets by default and add a
maad_internal_side dataset/config setting to opt in. Skipping products
record the setting in their identity; verify, compare, and merge-shards
expect skipped rows to be absent. The dashboard reads the setting from
datasets metadata and explains uncomputed internal sides.

Closes #121
Compare the datasets table structurally in merge-shards so a table
upgraded with ALTER TABLE matches a freshly created one, and default
maad_internal_side to 1 everywhere. Move column-mirror and config
checks into unit tests and trim the pipeline CLI test.
@flamboh
flamboh merged commit 11072e9 into main Oct 1, 2026
3 checks passed
@flamboh
flamboh deleted the feat/optional-internal-maad branch October 1, 2026 18:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider skipping MAAD for internal-side address sets

1 participant