This document contains all tasks needed to build the Multi-Agent System that solves the post-EOI contract workflow automation challenge.
Purpose: These tasks guide the implementation of agents that will autonomously process property deals. The tasks themselves are not the final solution - the agents they create are.
Structure: Tasks are organized in 9 phases (0-8), from foundation setup through final polish. Each task includes context, objectives, constraints, acceptance criteria, and test commands.
-
Read:
spec/MAS_Brief.mdspec/judging-criteria.mddocs/architecture.mddocs/project_plan.md
-
Reference:
agent_docs/testing.md(for pytest usage patterns)
Create a requirements.txt that lists all runtime and dev dependencies needed to run the multi-agent system, parse PDFs/emails, and run tests.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints where applicable (e.g. helper scripts, if any)
- Must follow existing code conventions in
src/ - Use broadly compatible version specifiers (e.g.
^or>=, <) to avoid over‑pinning - Include both runtime deps (PDF parsing, dates, SQLite, HTTP/LLM client if used) and dev deps (
pytest,mypy/ruffif desired)
-
requirements.txtincludes libraries to cover: PDF parsing, email/date handling, SQLite, and testing (pytest) -
pip install -r requirements.txtsucceeds in a clean Python 3.11 environment with no resolution errors
- Create:
requirements.txt - Test with:
pip install -r requirements.txt
-
Read:
folder-structure.txt(orspec/+docs/architecture.md)
-
Reference:
src/layout described infolder-structure.txt
Create __init__.py files and any missing package folders so that src/, src/agents/, src/agents/prompts/, src/utils/, and src/orchestrator/ are valid Python packages.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints in any non‑empty modules
- Must follow existing code conventions in
src/ __init__.pyfiles should be minimal, only exporting top‑level symbols if clearly needed later
- The following packages import without errors:
src,src.agents,src.agents.prompts,src.utils,src.orchestrator - Running
python -c "import src, src.agents, src.utils, src.orchestrator"from project root exits with status code 0
-
Create:
src/__init__.pysrc/agents/__init__.pysrc/agents/prompts/__init__.pysrc/utils/__init__.pysrc/orchestrator/__init__.py
-
Test with:
python -c "import src, src.agents, src.utils, src.orchestrator"
-
Read:
ground-truth/eoi_extracted.jsonground-truth/v1_extracted.jsonground-truth/v2_extracted.jsonground-truth/v1_mismatches.jsonground-truth/expected_outputs.jsonemails_manifest.json
-
Reference:
agent_docs/testing.md
Create tests/conftest.py providing reusable pytest fixtures for EOI data, contract data (V1/V2), mismatches, email manifest, and expected outputs.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints for all fixtures
- Must follow existing code conventions in
tests/ - Fixtures should read JSON from disk at runtime, not embed JSON inline
- Paths must be robust (e.g. use
Path(__file__).parent.parentor similar)
- Pytest discovers fixtures without raising
ImportErrororFileNotFoundError - Running
pytest --maxfail=1 --disable-warnings -qin an otherwise empty tests suite completes without error (even if zero tests are collected)
- Create:
tests/conftest.py - Test with:
pytest --maxfail=1 --disable-warnings -q
-
Read:
agent_docs/extraction.mddata/source-of-truth/EOI_John_JaneSmith.pdfdata/contracts/CONTRACT_V1.pdfdata/contracts/CONTRACT_V2.pdf
-
Reference:
ground-truth/eoi_extracted.jsonground-truth/v1_extracted.jsonground-truth/v2_extracted.json
Create src/utils/pdf_parser.py with reusable functions to extract text (and optionally simple tables) from EOI and contract PDFs for downstream extraction logic.
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include docstrings and type hints
-
Must follow existing code conventions in
src/ -
Should provide at least:
read_pdf_text(path: Path | str) -> strread_pdf_pages(path: Path | str) -> list[str]
-
Must not embed any OneCorp‑specific values; logic must work for other similar PDFs
- Calling
read_pdf_texton the EOI and contract PDFs returns non‑empty strings containing obvious anchors (e.g. “Expression of Interest”, “CONTRACT OF SALE OF REAL ESTATE”) - Utility functions are pure (no global state) and handle missing files by raising clear exceptions
- Create:
src/utils/pdf_parser.py - Test with:
pytest tests/test_utils.py::test_pdf_parser_basic(to be added in Task 1.4)
-
Read:
agent_docs/emails.mddata/emails/incoming/*.txtdata/emails/templates/*.txtemails_manifest.json
-
Reference:
- MAS brief email requirements in
spec/MAS_Brief.md
- MAS brief email requirements in
Create src/utils/email_parser.py that parses raw .txt email files into a structured Python object (e.g. dataclass or typed dict) capturing headers, body, and attachments.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/ - Parsed model must at minimum expose:
subject,from_addr,to_addrs,cc_addrs,body,attachment_filenames, and original file path - Parser must be robust to minor formatting differences and extra whitespace
- Parsing each
data/emails/incoming/*.txtfile yields a structured object that matchesemails_manifest.jsonfor key fields (subject, from, to, attachments) - No parsing function assumes a particular email ID; behaviour is driven by patterns (e.g. header prefixes, blank line separating body)
- Create:
src/utils/email_parser.py - Test with:
pytest tests/test_utils.py::test_email_parser_against_manifest(to be added in Task 1.4)
-
Read:
agent_docs/emails.md(appointment date rules)data/emails/incoming/04_solicitor_approved.txtemails_manifest.json(noteextracted_data.appointment_datetime)
-
Reference:
agent_docs/state-machine.md(SLA timer rules)expected_outputs.jsonSLA sections
Create src/utils/date_resolver.py to convert human phrases like “Thursday at 11:30am” into concrete timezone‑aware datetimes, given a reference date and timezone.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/ - Should expose a function like:
def resolve_appointment_phrase(base_dt: datetime, phrase: str, tz: tzinfo | str) -> datetime: ... - Must support at least weekday names + time (“Thursday at 11:30am”), and be easy to extend
- Resolving the phrase from
04_solicitor_approved.txtrelative to its email timestamp produces2025-01-16T11:30:00+11:00as peremails_manifest.json - Function raises a clear error or returns
Nonewhen it cannot confidently resolve the phrase (no silent incorrect guesses)
- Create:
src/utils/date_resolver.py - Test with:
pytest tests/test_utils.py::test_date_resolver_appointment_phrase(to be added in Task 1.4)
-
Read:
tests/conftest.pysrc/utils/pdf_parser.pysrc/utils/email_parser.pysrc/utils/date_resolver.py
-
Reference:
agent_docs/testing.mdemails_manifest.json
Create tests/test_utils.py covering core behaviours of pdf_parser, email_parser, and date_resolver using the supplied dataset.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints for helper functions
- Must follow existing code conventions in
tests/ - Tests should rely on fixtures from
conftest.pywherever possible - Tests should assert general properties (contains anchors, correct structure, correct datetime), not exact entire file contents
- Tests validate that PDF utilities return non‑empty, reasonably long text for EOI and contracts
- Tests validate that
email_parseroutput aligns withemails_manifest.jsonfor all incoming emails - Tests validate that
date_resolvercorrectly resolves the solicitor appointment phrase to the expected datetime -
pytest tests/test_utils.py -qpasses
- Create:
tests/test_utils.py - Test with:
pytest tests/test_utils.py -q
-
Read:
agent_docs/extraction.mdspec/MAS_Brief.mdground-truth/eoi_extracted.jsonground-truth/v1_extracted.jsonground-truth/v2_extracted.json
-
Reference:
docs/architecture.md(Extractor agent role)
Create a robust LLM prompt in src/agents/prompts/extractor_prompt.md that describes how to extract EOI and contract fields into a standard JSON schema.
- Must use pattern-based logic (no hardcoded demo values)
- Must include clear instructions and JSON schema examples
- Must follow existing markdown style in
src/agents/prompts/ - Prompt must cover both EOI and CONTRACT document types and specify required fields (names, emails, prices, finance, deposits, solicitor, vendor, lot, address)
- Must emphasise: no guessing, preserve numeric formats, mark missing/uncertain fields explicitly
- Prompt includes a single canonical JSON schema compatible with
ground-truth/*.json - Prompt instructs the model to normalise finance terms (e.g.
is_subject_to_financeboolean +termsstring) - Prompt makes no reference to specific OneCorp client names or file names; it remains reusable
- Create:
src/agents/prompts/extractor_prompt.md - Test with:
(Manual / later integration)
pytest tests/test_extraction.py -qafter Tasks 2.2–2.4
-
Read:
src/utils/pdf_parser.pyground-truth/eoi_extracted.jsondata/source-of-truth/EOI_John_JaneSmith.pdf
-
Reference:
agent_docs/extraction.mdsrc/agents/prompts/extractor_prompt.md
Implement extract_eoi(pdf_path) in src/agents/extractor.py to parse the EOI PDF into a structured dict matching the EOI ground‑truth schema.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/agents/ - Implementation may use direct parsing (regex/string rules) or LLM, but the public interface must be deterministic and testable
- Return structure must align with
ground-truth/eoi_extracted.json(fields, nesting, data types)
-
extract_eoi("data/source-of-truth/EOI_John_JaneSmith.pdf")returns a dict whose keys and types matcheoi_extracted.json - All required fields (purchasers, property, pricing, finance, solicitor, deposits, introducer) are present
-
tests/test_extraction.py::test_eoi_extraction_matches_ground_truthpasses
- Create / Update:
src/agents/extractor.py(addextract_eoi) - Test with:
pytest tests/test_extraction.py::test_eoi_extraction_matches_ground_truth -q
-
Read:
src/utils/pdf_parser.pyground-truth/v1_extracted.jsonground-truth/v2_extracted.jsondata/contracts/CONTRACT_V1.pdfdata/contracts/CONTRACT_V2.pdf
-
Reference:
agent_docs/extraction.mdsrc/agents/prompts/extractor_prompt.md
Implement extract_contract(pdf_path) in src/agents/extractor.py that extracts contract data from both V1 and V2 into the shared contract schema.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/agents/ - Must correctly parse: purchasers, property, pricing, finance terms, solicitor, deposits, vendor
- Implementation must not rely on knowing whether it’s V1 or V2 up front; it should infer from content or simply parse generically
-
extract_contractapplied to V1 and V2 returns dicts whose structure matchesv1_extracted.jsonandv2_extracted.json - Numeric fields (prices, deposits) are parsed as numbers, not strings
-
tests/test_extraction.py::test_contract_extraction_v1_v2_match_ground_truthpasses
- Update:
src/agents/extractor.py(addextract_contract) - Test with:
pytest tests/test_extraction.py::test_contract_extraction_v1_v2_match_ground_truth -q
-
Read:
tests/conftest.pysrc/agents/extractor.pyground-truth/eoi_extracted.jsonground-truth/v1_extracted.jsonground-truth/v2_extracted.json
-
Reference:
agent_docs/testing.md
Create tests/test_extraction.py to validate that extract_eoi and extract_contract reproduce the ground‑truth JSON structures for EOI and both contracts.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints for helper functions
- Must follow existing code conventions in
tests/ - Tests should allow for benign differences (e.g. whitespace) but enforce exact values for key fields (names, emails, lot, price, finance)
- Tests verify that EOI extraction exactly matches canonical values for critical fields (e.g. purchaser names, emails, lot, total price, finance terms)
- Tests verify that V1 and V2 extractions match their ground‑truth JSONs
-
pytest tests/test_extraction.py -qpasses
- Create:
tests/test_extraction.py - Test with:
pytest tests/test_extraction.py -q
This phase implements a hybrid router that uses deterministic pattern matching for high-confidence classifications and falls back to LLM for ambiguous cases. This approach saves approximately 90% of API calls while maintaining accuracy.
-
Read:
agent_docs/emails.mdspec/MAS_Brief.mdemails_manifest.json
-
Reference:
docs/architecture.md(Router role)
Create src/agents/prompts/router_prompt.md that explains how to classify emails into event types (EOI_SIGNED, CONTRACT_FROM_VENDOR, SOLICITOR_APPROVED_WITH_APPOINTMENT, DOCUSIGN_RELEASED, DOCUSIGN_BUYER_SIGNED, DOCUSIGN_EXECUTED, etc.).
Note: This prompt is used as the LLM fallback path when deterministic pattern matching confidence is below 0.8. The prompt should be optimized for handling ambiguous cases that the deterministic classifier could not resolve with high confidence.
- Must use pattern-based logic (no hardcoded demo values)
- Must include clear instructions and event type definitions
- Must follow existing markdown style in
src/agents/prompts/ - Prompt must describe how to use sender, subject line, body, and attachments as features
- Must instruct the model to output a small JSON object with
event_type,confidencescore, and any extracted metadata (e.g. appointment phrase) - Prompt should emphasize handling edge cases and ambiguous patterns that deterministic matching couldn't resolve
- Prompt lists all event types present in
emails_manifest.json - Prompt describes at least one example for each event type based on the dataset (without copying raw email text)
- Prompt instructs extraction of appointment phrases when present
- Prompt includes guidance for handling ambiguous cases (e.g., emails with mixed signals)
- Prompt specifies the JSON output format including
event_type,confidence, andmetadata
- Create:
src/agents/prompts/router_prompt.md - Test with:
(Manual / later integration)
pytest tests/test_email_classification.py -qafter Tasks 3.2–3.5
-
Read:
src/utils/email_parser.pyemails_manifest.jsondata/emails/incoming/*.txt
-
Reference:
agent_docs/emails.mdsrc/agents/prompts/router_prompt.md
Implement a hybrid classify_email(parsed_email) in src/agents/router.py that:
- First attempts deterministic pattern matching with confidence scoring
- If confidence >= 0.8, returns the classification immediately (no LLM call)
- If confidence < 0.8, falls back to LLM using the
router_prompt.md - Returns a
ClassificationResultwithevent_type,confidencescore, andmethodused (deterministicvsllm)
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/agents/ - Deterministic logic should rely on patterns (sender domains, subject contains, body keywords, attachments) derived from
emails_manifest.jsonand templates - LLM fallback must use the prompt from
router_prompt.md - Must define a
ClassificationResultdataclass/TypedDict with:event_type: str- The classified event typeconfidence: float- Confidence score (0.0 to 1.0)method: Literal["deterministic", "llm"]- Which classification method was usedmetadata: dict- Any extracted metadata (e.g., appointment phrase)
- Confidence threshold for deterministic classification: 0.8
- All incoming emails in the dataset are classified into the correct
event_typeas specified inemails_manifest.json - For solicitor approval emails, the returned event includes the appointment phrase (e.g. "Thursday at 11:30am") for later resolution
- High-confidence emails (clear patterns) are classified deterministically without LLM calls
- Ambiguous emails (confidence < 0.8) trigger LLM fallback classification
-
ClassificationResultincludesevent_type,confidence,method, andmetadatafields -
pytest tests/test_email_classification.py::test_router_classifies_all_emailspasses -
pytest tests/test_email_classification.py::test_hybrid_classification_methodpasses
- Create / Update:
src/agents/router.py - Test with:
pytest tests/test_email_classification.py -q
-
Read:
emails_manifest.json(patterns for each event type)data/emails/incoming/*.txtagent_docs/emails.md
-
Reference:
- Task 3.2 (parent task)
Implement a confidence scoring algorithm in src/agents/router.py that calculates how confident the deterministic classifier is about its classification. The algorithm should return a score between 0.0 and 1.0.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/agents/ - Scoring factors should include:
- Sender domain match (e.g.,
@docusign.comfor DocuSign events) - Subject line keyword match (e.g., "Expression of Interest", "Contract", "Signed")
- Body keyword density (presence of key phrases)
- Attachment type match (e.g., PDF attachments for contracts)
- Pattern exclusivity (does the email match only one event type or multiple?)
- Sender domain match (e.g.,
- Higher scores when multiple strong signals align; lower scores when signals are mixed or weak
- Function
calculate_confidence(parsed_email, candidate_event_type) -> floatis implemented - Returns values in range [0.0, 1.0]
- Clear, unambiguous emails (e.g., DocuSign from
@docusign.comwith "Completed" subject) score >= 0.8 - Ambiguous emails (e.g., generic subjects, multiple possible interpretations) score < 0.8
- Confidence calculation is deterministic and testable
-
pytest tests/test_email_classification.py::test_confidence_scoringpasses
- Update:
src/agents/router.py(add confidence scoring logic) - Test with:
pytest tests/test_email_classification.py::test_confidence_scoring -q
-
Read:
src/agents/prompts/router_prompt.mdsrc/agents/router.py(existing deterministic logic)
-
Reference:
docs/architecture.md(LLM integration patterns)requirements.txt(LLM client library)
Implement the LLM fallback path in src/agents/router.py that is triggered when deterministic classification confidence is below 0.8.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/agents/ - LLM client should:
- Load the prompt template from
router_prompt.md - Format the prompt with the parsed email content
- Parse the LLM response JSON to extract
event_type,confidence, andmetadata - Handle LLM errors gracefully (timeout, invalid response, etc.)
- Load the prompt template from
- Must support configuration for LLM endpoint/API key via environment variables
- Should implement retry logic with exponential backoff for transient failures
- Function
classify_with_llm(parsed_email) -> ClassificationResultis implemented - LLM is only called when deterministic confidence < 0.8
- Prompt template is loaded from
router_prompt.mdat runtime - LLM response is parsed correctly into
ClassificationResult - Graceful error handling for LLM failures (returns error result, doesn't crash)
-
pytest tests/test_email_classification.py::test_llm_fallbackpasses (with mocked LLM)
- Update:
src/agents/router.py(add LLM fallback logic) - Test with:
pytest tests/test_email_classification.py::test_llm_fallback -q
-
Read:
src/agents/router.py(hybrid classification implementation)
-
Reference:
docs/architecture.md(observability requirements)
Add comprehensive logging and metrics collection to the hybrid router for monitoring classification performance, LLM usage, and debugging.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/agents/ - Logging should include:
- Classification method used (
deterministicvsllm) - Confidence scores (deterministic and LLM)
- Event type classified
- Time taken for classification
- LLM API call count (for cost tracking)
- Classification method used (
- Metrics should be accessible for reporting:
- Percentage of emails classified deterministically vs LLM
- Average confidence scores by method
- Classification accuracy (when ground truth available)
- Use Python's
loggingmodule with appropriate log levels
- All classification calls are logged with method, confidence, and event type
- LLM fallback calls are logged at INFO level with timing information
- A
RouterMetricsclass or module tracks aggregate statistics:- Total classifications
- Deterministic vs LLM counts
- Average confidence by method
- Metrics can be retrieved for reporting (e.g.,
get_router_metrics() -> dict) - Log output is clean and useful for debugging without being verbose
-
pytest tests/test_email_classification.py::test_router_loggingpasses
- Update:
src/agents/router.py(add logging and metrics) - Test with:
pytest tests/test_email_classification.py::test_router_logging -q
-
Read:
tests/conftest.pysrc/utils/email_parser.pysrc/agents/router.pyemails_manifest.json
-
Reference:
agent_docs/testing.md
Create tests/test_email_classification.py to validate that each email in the dataset is classified to the correct event type and that required metadata is extracted.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints for helper functions
- Must follow existing code conventions in
tests/ - Tests should iterate over all INPUT emails from
emails_manifest.json
- Tests confirm that each email’s
event_typefromclassify_emailmatches the manifest - Tests confirm that solicitor approval events contain the appointment phrase
-
pytest tests/test_email_classification.py -qpasses
- Create:
tests/test_email_classification.py - Test with:
pytest tests/test_email_classification.py -q
-
Read:
agent_docs/comparison.mdspec/MAS_Brief.mdground-truth/v1_mismatches.json
-
Reference:
docs/architecture.md(Auditor role)
Create src/agents/prompts/auditor_prompt.md that explains how to compare EOI vs contract data, identify mismatches, and classify severity/risk.
- Must use pattern-based logic (no hardcoded demo values)
- Must include clear mismatch schema (field, display name, eoi_value, contract_value, severity, rationale)
- Must follow existing markdown style in
src/agents/prompts/ - Prompt must describe how to compute
is_valid,mismatch_count, andrisk_score - Must emphasise not to mark documents as valid when any HIGH mismatches exist
- Prompt includes mismatch fields aligned with
v1_mismatches.json - Prompt defines severity levels (LOW/MEDIUM/HIGH) and guidance
- Prompt describes how to generate
amendment_recommendationandnext_action
- Create:
src/agents/prompts/auditor_prompt.md - Test with:
(Manual / later integration)
pytest tests/test_comparison.py -qafter Tasks 4.2–4.3
-
Read:
ground-truth/eoi_extracted.jsonground-truth/v1_extracted.jsonground-truth/v2_extracted.jsonground-truth/v1_mismatches.json
-
Reference:
agent_docs/comparison.mdsrc/agents/prompts/auditor_prompt.md
Implement compare_contract_to_eoi(eoi_data, contract_data) in src/agents/auditor.py that returns a structured comparison result including mismatches, severity, and overall validity.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints (consider a
@dataclassforComparisonResult) - Must follow existing code conventions in
src/agents/ - Logic must be driven by field mappings (e.g.
pricing.total_price,finance.is_subject_to_finance) rather than inline constants
- Comparing EOI vs V1 returns exactly 5 mismatches with fields and values matching
v1_mismatches.json - Comparing EOI vs V2 returns zero mismatches and
is_valid = True -
risk_scoreandnext_actionfields are populated consistently for both cases -
pytest tests/test_comparison.py::test_v1_and_v2_comparisonpasses
- Create / Update:
src/agents/auditor.py - Test with:
pytest tests/test_comparison.py::test_v1_and_v2_comparison -q
-
Read:
tests/conftest.pysrc/agents/auditor.pyground-truth/v1_mismatches.json
-
Reference:
agent_docs/testing.md
Create tests/test_comparison.py to validate contract vs EOI comparison, including mismatches, severity, risk score, and validity.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
tests/ - Tests should check fields and counts, not just booleans
- Tests assert that V1 produces 5 mismatches matching ground truth (field names and formatted values)
- Tests assert that V2 produces no mismatches and
is_validis True - Tests assert that V1 yields a HIGH
risk_scoreandnext_actionindicates a discrepancy alert -
pytest tests/test_comparison.py -qpasses
- Create:
tests/test_comparison.py - Test with:
pytest tests/test_comparison.py -q
-
Read:
agent_docs/emails.md(templates + triggers)spec/MAS_Brief.md(section C & D: required outbound emails and alerts)
-
Reference:
data/emails/templates/03_contract_to_solicitor.txtdata/emails/templates/05_vendor_release_request.txtground-truth/expected_outputs.json(for alert emails)
Create src/agents/prompts/comms_prompt.md that instructs the LLM on how to generate the four required email types given structured context.
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include clear templates and placeholder fields (e.g.
{lot_number},{purchaser_names}) -
Must follow existing markdown style in
src/agents/prompts/ -
Prompt must cover:
- Contract to solicitor
- Vendor DocuSign release request
- Internal discrepancy alert
- SLA overdue alert
- Prompt documents required headers (From/To/Subject) and body structure for all 4 emails
- Prompt describes how to include mismatch details and recommendations in discrepancy alerts
- Prompt describes SLA alert content (property, appointment, time elapsed, recommended action)
- Create:
src/agents/prompts/comms_prompt.md - Test with:
(Manual / later integration)
pytest tests/test_comms.py -qafter Tasks 5.2–5.3
-
Read:
spec/MAS_Brief.md(email requirements)data/emails/templates/*.txtground-truth/expected_outputs.json
-
Reference:
agent_docs/emails.mdsrc/agents/prompts/comms_prompt.md
Implement all four email builder functions in src/agents/comms.py to generate structured email objects from context dicts.
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include docstrings and type hints
-
Must follow existing code conventions in
src/agents/ -
Functions to implement (names can be refined, but must be clear and typed):
build_contract_to_solicitor_email(context)build_vendor_release_email(context)build_discrepancy_alert_email(context)build_sla_overdue_alert_email(context)
-
Builders must take structured context (deal, comparison result, appointment, SLA status) and return an email model (e.g. dataclass) that can be rendered to text
- Contract-to-solicitor builder produces an email matching the sample template for Lot 95 / Smith (subject, recipients, attachment name) when given equivalent context
- Vendor-release builder produces an email structurally matching the sample request with DocuSign language
- Discrepancy alert and SLA alert builders produce bodies containing all required fields as specified in
MAS_Brief.md/expected_outputs.json -
pytest tests/test_comms.py::test_email_builders_against_expected_outputs -qpasses
- Create / Update:
src/agents/comms.py - Test with:
pytest tests/test_comms.py::test_email_builders_against_expected_outputs -q
-
Read:
tests/conftest.pysrc/agents/comms.pyground-truth/expected_outputs.json
-
Reference:
agent_docs/testing.mdagent_docs/emails.md
Create tests/test_comms.py to validate all four email builders against canonical expected outputs.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
tests/ - Tests may allow minor whitespace differences but must enforce headers, key phrases, and dynamic fields (names, lot, prices, dates)
- Tests confirm that generated emails for the Smith / Lot 95 scenario match the structure and key content in
expected_outputs.json - Tests assert that discrepancy alerts list each mismatch with EOI and contract values
- Tests assert that SLA alerts include property, appointment, time elapsed, and a “Recommended Action” section
-
pytest tests/test_comms.py -qpasses
- Create:
tests/test_comms.py - Test with:
pytest tests/test_comms.py -q
-
Read:
agent_docs/state-machine.mdspec/MAS_Brief.mdground-truth/expected_outputs.json(workflow_stages & sla_test_scenario)
-
Reference:
docs/architecture.md
Implement src/orchestrator/state_machine.py defining a State enum, transition rules, and guard conditions for the Lot 95 workflow (and extensible for more deals).
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
src/orchestrator/ - State machine must support at least the states listed in
workflow_stages(EOI_RECEIVED → CONTRACT_V1_RECEIVED → … → EXECUTED) - Transitions must be driven by events (router events, comparison results, SLA events), not by direct function calls bypassing the model
- State enum covers all states in
expected_outputs.json.workflow_stages - Transition logic enforces versioning (V2 supersedes V1) and prevents invalid transitions
- A simple simulation using the manifest’s events reproduces the state sequence in
workflow_stages -
pytest tests/test_state_transitions.py::test_happy_path_transitions -qpasses
- Create:
src/orchestrator/state_machine.py - Test with:
pytest tests/test_state_transitions.py::test_happy_path_transitions -q
-
Read:
agent_docs/state-machine.md(deal data model)ground-truth/expected_outputs.json(deal_id, stages)
-
Reference:
docs/architecture.md(persistence layer expectations)
Create src/orchestrator/deal_store.py to persist deals, states, events, and SLA timers using SQLite.
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include docstrings and type hints
-
Must follow existing code conventions in
src/orchestrator/ -
Should provide a small API, e.g.:
upsert_deal(deal)record_event(deal_id, event)update_state(deal_id, new_state)get_deal(deal_id)get_pending_sla_checks(now)
-
Must use standard library
sqlite3(no heavy ORM required)
- Deals and events for
LOT95_FAKE_RISE_VIC_3336can be persisted and retrieved without data loss - Persisted state sequence matches the state machine transitions when replaying events
-
pytest tests/test_state_transitions.py::test_persistence_of_deal_state -qpasses
- Create:
src/orchestrator/deal_store.py - Test with:
pytest tests/test_state_transitions.py::test_persistence_of_deal_state -q
-
Read:
agent_docs/state-machine.md(SLA rules)emails_manifest.json(sla_rulessection)data/emails/incoming/04_solicitor_approved.txt
-
Reference:
src/utils/date_resolver.pyground-truth/expected_outputs.json.sla_test_scenario
Create src/orchestrator/sla_monitor.py to schedule and evaluate SLA deadlines (e.g. buyer signature due 2 days after appointment).
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include docstrings and type hints
-
Must follow existing code conventions in
src/orchestrator/ -
Must integrate with
deal_storeto:- Register SLA timers when solicitor appointment is set
- Cancel SLA timers when buyer signs
- Emit SLA overdue events when deadlines pass
- For the happy path (with buyer-signed email present), SLA timer is scheduled then cancelled before firing
- For the SLA test scenario (with buyer-signed email removed), SLA monitor emits an SLA overdue event at the expected time per
sla_test_scenario -
pytest tests/test_state_transitions.py::test_sla_overdue_scenario -qpasses
- Create:
src/orchestrator/sla_monitor.py - Test with:
pytest tests/test_state_transitions.py::test_sla_overdue_scenario -q
-
Read:
tests/conftest.pysrc/agents/router.pysrc/agents/auditor.pysrc/orchestrator/state_machine.pysrc/orchestrator/deal_store.pysrc/orchestrator/sla_monitor.pyemails_manifest.jsonground-truth/expected_outputs.json
-
Reference:
agent_docs/testing.mdagent_docs/state-machine.md
Create tests/test_state_transitions.py to validate the orchestrated workflow, including state sequence and SLA behaviour.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
tests/ - Tests should replay the events described in
workflow_stagesandsla_test_scenario
- Happy-path test reproduces the state sequence from
EOI_RECEIVEDthroughEXECUTEDas inworkflow_stages - SLA test removes the buyer-signed event and asserts that an SLA overdue alert is generated at the right time
- Guards prevent invalid transitions (e.g. cannot send to solicitor before a valid contract is confirmed)
-
pytest tests/test_state_transitions.py -qpasses
- Create:
tests/test_state_transitions.py - Test with:
pytest tests/test_state_transitions.py -q
-
Read:
docs/demo-script.mddocs/architecture.mdemails_manifest.jsonground-truth/expected_outputs.json
-
Reference:
spec/MAS_Brief.md(end-to-end demo requirements)
Implement src/main.py as a CLI entry point that runs the full demo workflow for the provided dataset and prints key agent interactions and state transitions.
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include docstrings and type hints
-
Must follow existing code conventions in
src/ -
CLI should:
- Ingest EOI and contracts
- Process emails in manifest order through Router, Extractor, Auditor, Orchestrator, SLA monitor, and Comms
- Log/print clear, concise messages suitable for a 3-minute demo
- Running
python -m src.main(orpython src/main.py) processes the Lot 95 dataset end-to-end without errors - Console output shows: V1 discrepancies → discrepancy alert → V2 validated → solicitor approval & appointment → vendor release request → DocuSign emails → executed state → SLA test (when configured)
-
pytest tests/test_end_to_end.py::test_demo_entrypoint -qpasses
- Create / Update:
src/main.py - Test with:
pytest tests/test_end_to_end.py::test_demo_entrypoint -q
-
Read:
- All key modules:
src/agents/*,src/orchestrator/*,src/utils/*,src/main.py emails_manifest.jsonground-truth/expected_outputs.json
- All key modules:
-
Reference:
agent_docs/testing.md
Create tests/test_end_to_end.py to validate the full workflow and generated emails against expected_outputs.json.
- Must use pattern-based logic (no hardcoded demo values)
- Must include docstrings and type hints
- Must follow existing code conventions in
tests/ - Test should orchestrate the same flow as
src/main.py, but in a programmable way (no reliance on stdout parsing)
- Test confirms that all expected outbound emails (solicitor, vendor release, discrepancy alert, SLA alert in the SLA scenario) are generated with correct structure and key content
- Test confirms that final deal state is
EXECUTEDfor the normal run - Test confirms that SLA overdue alert is only generated in the explicitly simulated SLA failure scenario
-
pytest tests/test_end_to_end.py -qpasses
- Create:
tests/test_end_to_end.py - Test with:
pytest tests/test_end_to_end.py -q
-
Read:
spec/MAS_Brief.mdspec/judging-criteria.mddocs/architecture.mdsrc/main.pybehaviour
-
Reference:
- Existing
docs/demo-script.md(if present) as a starting point
- Existing
Write or refine docs/demo-script.md into a tight 3-minute demo walkthrough aligned with how src/main.py runs and what the judges care about.
- Must use pattern-based logic (no hardcoded demo values)
- Must be clear, step-by-step, with specific CLI commands and which logs to point at
- Must emphasise multi-agent collaboration, error handling, and SLA logic per judging criteria
- Script fits into ~3 minutes when read aloud and executed
- Script clearly calls out each agent’s role (Router, Extractor, Auditor, Comms, Orchestrator/SLA)
- Script explicitly maps key moments to judging criteria (architecture, collaboration, safety, real-world value)
- Document is linked from
README.md
- Create / Update:
docs/demo-script.md - Test with:
Manual dry run following the script using
python -m src.main
-
Read:
docs/architecture.mdagent_docs/state-machine.mddocs/demo-script.md
-
Reference:
spec/judging-criteria.md(System Design & Collaboration sections)
Create an up-to-date architecture diagram (SVG or PNG) showing agents, data stores, and message flows, suitable for use in the README and demo.
- Must use pattern-based logic (no hardcoded demo values)
- Must clearly label agents, key tools, and state machine
- File should be reasonably small and stored under
assets/(e.g.assets/architecture.svg)
- Diagram reflects the actual implemented architecture (Router, Extractor, Auditor, Comms, Orchestrator, SLA monitor, Deal store)
- Diagram visually distinguishes data flow (EOI/contract, emails) and control flow (events, state transitions)
- Diagram is referenced in
docs/architecture.mdandREADME.md
- Create:
assets/architecture.svg(or.png) - Test with: Open file visually and confirm clarity; no automated test required
-
Read:
- Existing
README.md spec/MAS_Brief.mdspec/judging-criteria.mddocs/demo-script.md
- Existing
-
Reference:
docs/INDEX.md(doc index)
Update README.md with clear setup instructions, how to run the demo, how to run tests, and a concise explanation of the multi-agent architecture.
-
Must use pattern-based logic (no hardcoded demo values)
-
Must include:
- Quickstart section (install, run demo, run tests)
- Short agent overview
- Link to architecture diagram and demo script
- Note on safety / guardrails and limitations
-
Language should be concise and accessible to judges unfamiliar with the codebase
- Someone new to the repo can clone, install, run the demo, and run tests using only the README
- README references the architecture diagram and demo script
- README mentions how the system meets the MAS brief and judging criteria at a high level
- Update:
README.md - Test with: Manual copy-paste of commands from README into a clean environment
-
Read:
docs/demo-script.mdREADME.md
-
Reference:
spec/judging-criteria.md(Presentation & UX)
Provide a short guide (e.g. docs/demo-recording.md or README section) on how to record a terminal or screen demo of the system following the demo script.
- Must use pattern-based logic (no hardcoded demo values)
- Must be tool-agnostic where possible (e.g. suggest
asciinemaor common screen recorders, but not required) - Instructions should be optional and non‑blocking for running the code
- Document clearly explains how to run the demo commands while recording
- Document links back to
docs/demo-script.md - Judges or teammates could reproduce a similar recording by following the steps
- Create:
docs/demo-recording.md(or new “Recording the Demo” section inREADME.md) - Test with: Manual trial recording using the described steps (no automated test required)