Skip to content

[Enhancement] Live System Event Tracking: Fix or Remove Underdeveloped Feature #29

Description

@abdulmeLINK

Issue Description

The Live System Event tracking feature exists in the dashboard but is either unnecessary (duplicates other monitoring) or underdeveloped (limited functionality, unclear value proposition).

Evidence

Event tracking EXISTS but with limitations:

  1. Dashboard Event Page (dashboard/frontend/src/pages/EventsPage.tsx)

    • 738 lines of code
    • Shows events from Collector API
    • WebSocket support for real-time updates
    • Filtering by level, component, time
  2. Backend Event API (dashboard/backend/app/services/collector_client.py)

    • Line 438: get_system_events() method
    • Fetches events from Collector service
  3. Collector Event Storage (src/collector/event_monitor.py)

    • Line 464+: Collects events from FL Server, Policy Engine, SDN Controller
    • Topology snapshots stored as events
  4. Empty events.jsonl file:

    ls -l logs/events.jsonl
    # File exists but is empty (0 bytes)

What's NOT working:

  • logs/events.jsonl is never populated
  • Event collection seems limited to topology snapshots
  • No clear event schema or taxonomy
  • Unclear distinction from metrics/logs
  • Not integrated with critical system actions

Verified with:

# Check event file usage
grep -r "events.jsonl" src --include="*.py"
# Returns: 0 matches - file never accessed!

# Check EventMonitor usage
grep -r "EventMonitor\|event_monitor" src --include="*.py"
# Returns: Limited usage, mainly in collector

Problem Statement

Key questions without clear answers:

  1. What is an "event"? vs metric? vs log? No clear definition
  2. What events should be tracked? No comprehensive list
  3. Who needs events? Research? Debugging? Operations?
  4. Why use events? What problems do they solve that logs/metrics don't?

Current state analysis:

Option A: Events are UNNECESSARY

If events duplicate existing monitoring:

  • Logs already captured (docker logs, application logs)
  • Metrics already collected (Collector service, SQLite)
  • Dashboard shows FL round progress, policy decisions
  • Topology changes tracked by SDN controller

Evidence: Most "events" are just metric snapshots

Option B: Events are UNDERDEVELOPED

If events provide unique value, they need:

  • Clear event taxonomy
  • Comprehensive event generation
  • Proper storage (events.jsonl never used)
  • Research/debugging workflows
  • Integration with critical actions

Suggested Improvements

If Keeping Events (Option B - Recommended)

Define clear event taxonomy:

# Event types that provide unique value
class SystemEventType(Enum):
    # FL Training Events (actionable)
    FL_ROUND_STARTED = "fl_round_started"
    FL_ROUND_COMPLETED = "fl_round_completed"
    FL_CLIENT_CONNECTED = "fl_client_connected"
    FL_CLIENT_DISCONNECTED = "fl_client_disconnected"
    FL_CLIENT_FAILED = "fl_client_failed"
    FL_MODEL_AGGREGATED = "fl_model_aggregated"
    
    # Policy Events (compliance/audit)
    POLICY_EVALUATED = "policy_evaluated"
    POLICY_ALLOWED = "policy_allowed"
    POLICY_DENIED = "policy_denied"
    POLICY_VIOLATION = "policy_violation"
    
    # Network Events (troubleshooting)
    NODE_ADDED = "node_added"
    NODE_REMOVED = "node_removed"
    LINK_UP = "link_up"
    LINK_DOWN = "link_down"
    NETWORK_PARTITION = "network_partition"
    
    # System Events (operations)
    SCENARIO_STARTED = "scenario_started"
    SCENARIO_COMPLETED = "scenario_completed"
    SERVICE_STARTED = "service_started"
    SERVICE_CRASHED = "service_crashed"
    CONFIGURATION_CHANGED = "config_changed"

Implement proper event logging:

# src/core/events/event_logger.py
class EventLogger:
    """Centralized event logging to events.jsonl."""
    
    def __init__(self, log_file: str = "logs/events.jsonl"):
        self.log_file = Path(log_file)
        self.log_file.parent.mkdir(exist_ok=True)
    
    def log_event(
        self,
        event_type: SystemEventType,
        component: str,
        details: Dict[str, Any],
        severity: str = "INFO"
    ):
        """Log structured event to JSONL file."""
        event = {
            "timestamp": datetime.utcnow().isoformat(),
            "event_type": event_type.value,
            "component": component,
            "severity": severity,
            "details": details,
            "version": "1.0.0"
        }
        
        with open(self.log_file, 'a') as f:
            f.write(json.dumps(event) + '\n')

# Global instance
event_logger = EventLogger()

Integrate with all components:

# In FL Server
from src.core.events.event_logger import event_logger, SystemEventType

def on_round_start(round_num: int):
    event_logger.log_event(
        SystemEventType.FL_ROUND_STARTED,
        component="fl_server",
        details={"round": round_num, "clients_selected": client_count}
    )

# In Policy Engine
def evaluate_policy(policy_id: str, context: dict) -> dict:
    result = _evaluate(policy_id, context)
    
    event_logger.log_event(
        SystemEventType.POLICY_EVALUATED if result['allowed'] else SystemEventType.POLICY_DENIED,
        component="policy_engine",
        details={
            "policy_id": policy_id,
            "allowed": result['allowed'],
            "client_id": context.get('client_id')
        },
        severity="WARNING" if not result['allowed'] else "INFO"
    )
    
    return result

Create event analysis tools:

# scripts/analyze_events.py
def analyze_events(log_file: str):
    """Analyze event log for patterns and issues."""
    events = []
    with open(log_file) as f:
        for line in f:
            events.append(json.loads(line))
    
    # Count by type
    type_counts = Counter(e['event_type'] for e in events)
    
    # Find anomalies
    policy_denials = [e for e in events if e['event_type'] == 'policy_denied']
    client_failures = [e for e in events if e['event_type'] == 'fl_client_failed']
    
    # Timeline analysis
    round_starts = [e for e in events if e['event_type'] == 'fl_round_started']
    round_durations = _compute_durations(round_starts)
    
    return {
        'total_events': len(events),
        'by_type': type_counts,
        'anomalies': {
            'policy_denials': len(policy_denials),
            'client_failures': len(client_failures)
        },
        'round_durations': round_durations
    }

If Removing Events (Option A)

Simplify to use existing monitoring:

  1. Remove EventsPage.tsx from dashboard (738 lines)
  2. Remove event_monitor.py from collector
  3. Remove logs/events.jsonl references
  4. Use logs for debugging (already comprehensive)
  5. Use metrics for analysis (already collected)
  6. Use dashboard tabs for monitoring (already exists)

Benefits of removal:

  • Less code to maintain
  • Clearer separation: logs for debugging, metrics for analysis
  • One less storage format to manage
  • Simpler architecture

Recommendation

Option B: Improve event system for these reasons:

  1. Research value: Events provide timeline of FL training for papers
  2. Debugging: Event sequences help troubleshoot complex issues
  3. Audit trail: Policy decisions need immutable record
  4. Reproducibility: Event logs enable experiment replay

But needs work:

  • ⚠️ Define clear event taxonomy
  • ⚠️ Implement comprehensive event generation
  • ⚠️ Fix events.jsonl writing (currently broken)
  • ⚠️ Add event analysis tools
  • ⚠️ Document use cases for researchers

Implementation Plan

Phase 1: Fix Basic Infrastructure (Week 1)

  • Fix events.jsonl writing (currently never populated)
  • Define event schema and taxonomy
  • Implement EventLogger class
  • Add unit tests for event logging

Phase 2: Instrument Components (Week 2)

  • Add event logging to FL Server (round lifecycle)
  • Add event logging to Policy Engine (decisions)
  • Add event logging to Collector (network changes)
  • Add event logging to scenario runner (lifecycle)

Phase 3: Analysis Tools (Week 3)

  • Create event analysis scripts
  • Add event export (CSV, JSON)
  • Timeline visualization
  • Event-based alerts

Phase 4: Dashboard Integration (Week 4)

  • Improve EventsPage with better filtering
  • Add event timeline visualization
  • Add event search
  • Add event export UI

Acceptance Criteria (If Improving)

  • logs/events.jsonl populated with events
  • Events generated for all critical FL actions
  • Event schema documented
  • Analysis scripts work on real event logs
  • Dashboard EventsPage shows useful information
  • Events help with actual debugging/research

Acceptance Criteria (If Removing)

  • EventsPage removed from dashboard
  • Event collection code removed
  • No references to events.jsonl
  • Documentation updated to use logs/metrics
  • No loss of debugging capability

Related Issues

Priority

Recommended: Medium

Current state is confusing. Either fix properly or remove entirely.

Target Version: v1.1.0

Additional Notes

Current state: ⚠️ Half-implemented feature that provides limited value

Decision needed: Invest in making events useful OR simplify by removing them

Recommendation: Fix it - event logs have genuine research value for FL systems, but implementation needs work.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dashboardDashboard frontend/backend issuesenhancementNew feature or requestmonitoringMonitoring and metrics issues

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions