Skip to content

Latest commit

 

History

History
221 lines (176 loc) · 13.4 KB

File metadata and controls

221 lines (176 loc) · 13.4 KB

Test Coverage Gap Analysis

This document tracks infrastructure test coverage gaps by Layer-1 module. The global infrastructure gate remains 60%. Module-% tables below are historical notes, not live gates. Live measured coverage and exemplar floors belong in docs/_generated/COUNTS.md.

Last verified: 2026-08-09 (source-bound guidance; current measurements are read from the generated coverage receipts below rather than copied into this document).

Current evidence: the 24 public-exemplar coverage rows and source-tree identities are maintained in docs/_generated/coverage_snapshot.json and checked by uv run python scripts/docgen/counts.py --check. The current infrastructure aggregate is reported by the live infrastructure coverage gate and its receipt; no stale aggregate is asserted here until that receipt is available.

Coverage oracle: full infrastructure gate:

COVERAGE_FILE=.coverage.infra uv run pytest tests/infra_tests/ \
  -n 2 --dist loadscope --benchmark-disable \
  --cov=infrastructure --cov-report=term-missing --cov-fail-under=60 \
  --durations=10 \
  -m "not requires_ollama and not requires_docker and not network and not slow and not bench and not benchmark and not performance" \
  --timeout=120

The uncached serial diagnostic oracle uses the same command with the xdist
flags removed; it is the comparison baseline for performance claims.

Historical infrastructure baseline (2026-07-22)

The following values are preserved as historical evidence from the dated oracle run. They are not current release claims and must not be copied into new documentation without a new receipt:

Overall infrastructure coverage: 83.38% (gate: >= 60%) Tests: 9066 passed on the 2026-07-22 oracle run. One existing NumPy overflow warning in the scientific stability edge-case test. Total statements measured: 52,607

The 2026-07-22 uncached local run used two macOS xdist workers with --dist loadscope and completed in 554.50 seconds. Higher-worker coverage is rejected on macOS by the shared runner guard because the prior work-stealing lane reproduced worker-replacement failures; Linux retains explicit higher worker opt-in.

The sorted module rows were taken from:

uv run coverage report --sort=cover -m

Module Documentation Inventory

All top-level code packages under infrastructure/ have the expected README.md, AGENTS.md, and SKILL.md files. The only top-level documentation exception is infrastructure/logrotate.d/, which is a config directory and carries README.md plus AGENTS.md but no skill-routing surface. Ignored/generated directories such as infrastructure/.benchmarks/ and infrastructure/__pycache__/ are not module packages.

The publication validation and advanced-literature skill surfaces changed in this pass; the generated skill/index checks are therefore part of the release gate rather than an assumption that routing metadata is unchanged.

Target Categories

Category Coverage expectation Action
First-party logic below 60% Add meaningful branch coverage when the branch can be driven with real files, real subprocesses, or deterministic fixtures. Document the next concrete branch gap until tested.
CLI/subprocess shims Require smoke or subprocess tests for command behavior. Do not treat in-process 0% as a defect by itself for thin __main__.py or dispatch shims.
Optional-tool or LLM-gated modules Exercise missing-tool, offline, and fallback behavior in the default suite. Keep live-tool paths behind their explicit gates unless CI installs the tool.
Publish/security/release paths Prefer dry-run, fake executable, or local fixture coverage. Do not require credentials, network publication, or destructive release actions in default tests.

Current Low Rows

CLI And Subprocess Shim Rows

These rows are low because coverage is collected in-process while the behavior is intentionally command-oriented. The target is subprocess smoke coverage for user-visible command behavior, not line-chasing through dispatch glue.

Module Coverage Reason / next target
autoresearch/cli.py 0.00% CLI dispatch shim; add subprocess smoke when flags change.
core/pipeline/multi_project_cli.py 0.00% Multi-project command wrapper; keep behavior covered through orchestration and subprocess command tests.
doctor/__main__.py and other __main__.py files 0.00% Entry-point wrappers; no separate unit tests needed unless import behavior changes.
documentation/active_projects_doc.py 0.00% Generated-doc command shim; validate through the active-projects doc generator and docs consistency gates.
methods/cli.py 0.00% Methods CLI shim; add subprocess smoke for changed commands.
prose/cli.py 0.00% Prose CLI behavior is subprocess-tested; 0% in-process is not itself a defect.
publishing/pypi_release.py 0.00% Release helper; default tests should stay on dry-run/local-fixture paths.
sia/cli.py 0.00% SIA command shim; command behavior belongs in subprocess CLI tests.

First-Party Logic Below 60%

All previously-below-60% first-party modules now have meaningful no-mock coverage from the 2026-07-22 COVERAGE-BASELINE-1 wave (145 new tests). The modules below are the remaining lower-coverage rows above the 60% gate that still have concrete branch gaps to close.

Module Measured coverage (2026-08-21, module-scoped pytest) Tests added Test file
project/workspace.py 88.89% (94%+ incl. new branch-gap tests) +6 tests/infra_tests/project/test_workspace_branches.py
publishing/transmission_page_check.py 85.71% (remaining miss is the subprocess-tested main() guard) covered by _additional tests/infra_tests/publishing/test_transmission_page_check_additional.py
rendering/docx_renderer.py 84.09% (remaining gaps are optional-dependency error branches) covered by _additional tests/infra_tests/rendering/test_docx_renderer_additional.py
rendering/_pipeline_summary.py 85.95% (95.25% incl. new branch-gap tests) +19 tests/infra_tests/rendering/test_pipeline_summary_branches.py
rendering/epub_renderer.py 63.92% covered by _additional tests/infra_tests/rendering/test_epub_renderer_additional.py
documentation/publication_records.py 74.06% covered by _additional tests/infra_tests/documentation/test_publication_records_additional.py

Note (2026-08-21): percentages in this table were re-measured with module-scoped pytest runs; earlier single-module figures predate the _additional test waves and understate current coverage. Re-derive before citing any number here.

Optional-Tool And Gated Rows

Module Coverage Gate / fallback target
llm/review/pipeline_runner.py 12.88% LLM review orchestration; default suite should cover offline/fallback summaries, live LLM remains gated.
sia/live_llm.py 25.81% Optional Ollama feedback path; keep offline failure handling in default tests.
llm/review/ollama_setup.py 51.52% Ollama setup path; cover missing binary/server and actionable install-message branches.
search/deep_research/gemini.py 51.02% External API backend; default target is config validation and missing-credential behavior.
validation/security_gate.py 54.47% Added missing-tool and severity aggregation tests; next target is parser coverage for each configured scanner output.
llm/utils/server.py 57.79% Server lifecycle path; default target is unavailable-server diagnostics and timeout handling.
validation/docs/lint_runner.py 57.97% Subprocess/tool-gated docs runner; cover missing mmdc/Chrome fallbacks and failed-tool aggregation.
search/deep_research/openai.py 56.99% External API backend; default target is missing-key, timeout, and malformed-response handling.

Recently Added Module Tests (2026-06-27)

Eight modules promoted out of the "First-Party Logic Below 60%" table. All now exceed the 60% gate; branch coverage was driven with real files, subprocess fixtures, and deterministic local paths — no mocks introduced.

Module Previous Current Tests added Test file
rendering/_combined_exports.py 15.70% 83.43% +24 tests/infra_tests/rendering/test_combined_exports.py
project/drift/runner.py 18.64% 100.00% +17 tests/infra_tests/project/test_thin_orchestrator_drift.py
doctor/detectors/layout.py 31.08% 95.95% +11 tests/infra_tests/doctor/test_detectors.py
core/install_commands.py 38.89% 100.00% +8 tests/infra_tests/core/test_install_commands.py
rendering/pipeline.py 39.37% 96.85% +11 tests/infra_tests/rendering/test_pipeline.py
core/runtime/env_deps.py 46.22% 84.03% +8 tests/infra_tests/core/test_env_deps.py
core/runtime/setup_checks.py 46.67% 85.71% +15 tests/infra_tests/core/test_setup_checks.py
project/working_render.py 46.67% 90.33% +25 tests/infra_tests/project/test_working_render.py

The six thin-orchestrator script violations recorded here have been moved into infrastructure/ (package discovery, stage-label resolution, filepath statistics, no-mocks scan-root resolution, Stage-00 setup aggregation, and stage-table generation from pipeline.yaml). Scripts in those paths now coordinate I/O; the algorithms live in tested Layer-1 modules. Re-run scripts/audit/check_template_drift.py rather than treating this paragraph as a live violation list.

Parity notes: infrastructure/docker/ has partial coverage via tests/infra_tests/rendering/test_dockerfile_gen.py (no dedicated tests/infra_tests/docker/ needed — not a Python package). infrastructure/logrotate.d/ is a config directory with zero test coverage by design; tests/infra_tests/gates/ and tests/infra_tests/git_hook_smoke/ are legitimate test locations for non-infrastructure code (gate scripts and git hooks).

Recently Added Module Tests (2026-06-26)

New subpackages from PUB-PLATFORM-1. Coverage measured via dry-run / local-fixture paths; live network and credential-gated paths are excluded from the default suite.

Module Current coverage Test file
publishing/pypi/adapter.py dry-run paths only tests/infra_tests/publishing/test_pypi.py (11 tests)
publishing/static_site/registry.py 100% (pure factory) tests/infra_tests/publishing/test_static_site.py (22 tests)
publishing/archival/orchestrate.py dry-run paths only tests/infra_tests/publishing/test_archival_module.py (57 tests)
publishing/registry.py 100% (pure registry) tests/infra_tests/publishing/test_registry.py (47 tests)

Next targets: orchestrate.load_credentials fallback for missing credentials file and _missing_credential_receipt return paths in providers — both driveable with tmp_path.

Recently Added Module Tests (2026-06-16)

Canonical tests were added or extended in the existing module test locations; no *_coverage.py, *_full.py, or duplicate supplement files were introduced.

Module Current coverage Test file
project/info.py 87.25% tests/infra_tests/project/test_info.py
project/workspace.py 51.11% tests/infra_tests/project/test_workspace.py
rendering/_pdf_section_titles.py 94.44% tests/infra_tests/rendering/test_pdf_section_titles.py
rendering/pdf_renderer.py 68.70% tests/infra_tests/rendering/test_pdf_renderer.py
validation/security_gate.py 53.85% tests/infra_tests/validation/test_security_gate.py
validation/plugin_export.py 68.64% tests/infra_tests/validation/test_plugin_export.py
autoresearch/reports.py 86.67% tests/infra_tests/autoresearch/test_autoresearch_plan_validation.py
benchmark/template_harness.py 61.25% tests/infra_tests/benchmark/test_template_benchmark_harness.py
reporting/executive_outputs.py 80.36% tests/infra_tests/reporting/test_executive_outputs.py

Coverage Gates

  • Infrastructure: >= 60%; current aggregate comes from the live coverage receipt.
  • Projects: >= 90% per project, with rotating-project exceptions documented in CI and project-local AGENTS.md files.
  • Per-module targets: documented here only; they are not CI gates.

Testing Standards

  • No mocks: every test uses real data, real files, real subprocess calls, or deterministic local fixtures. See infrastructure/validation/output/no_mock_enforcer.py.
  • Deterministic: fixed RNG seeds, MPLBACKEND=Agg, and hermetic subprocess environments through repository helpers.
  • Canonical test files: keep module tests under tests/infra_tests/<module>/ and avoid new supplement tiers such as *_coverage.py, *_full.py, or *_comprehensive.py unless the production subject is coverage reporting itself.

See Also