Issue Description
The Live System Event tracking feature exists in the dashboard but is either unnecessary (duplicates other monitoring) or underdeveloped (limited functionality, unclear value proposition).
Evidence
Event tracking EXISTS but with limitations:
-
Dashboard Event Page (dashboard/frontend/src/pages/EventsPage.tsx)
- 738 lines of code
- Shows events from Collector API
- WebSocket support for real-time updates
- Filtering by level, component, time
-
Backend Event API (dashboard/backend/app/services/collector_client.py)
- Line 438:
get_system_events() method
- Fetches events from Collector service
-
Collector Event Storage (src/collector/event_monitor.py)
- Line 464+: Collects events from FL Server, Policy Engine, SDN Controller
- Topology snapshots stored as events
-
Empty events.jsonl file:
ls -l logs/events.jsonl
# File exists but is empty (0 bytes)
What's NOT working:
logs/events.jsonl is never populated
- Event collection seems limited to topology snapshots
- No clear event schema or taxonomy
- Unclear distinction from metrics/logs
- Not integrated with critical system actions
Verified with:
# Check event file usage
grep -r "events.jsonl" src --include="*.py"
# Returns: 0 matches - file never accessed!
# Check EventMonitor usage
grep -r "EventMonitor\|event_monitor" src --include="*.py"
# Returns: Limited usage, mainly in collector
Problem Statement
Key questions without clear answers:
- What is an "event"? vs metric? vs log? No clear definition
- What events should be tracked? No comprehensive list
- Who needs events? Research? Debugging? Operations?
- Why use events? What problems do they solve that logs/metrics don't?
Current state analysis:
Option A: Events are UNNECESSARY
If events duplicate existing monitoring:
- Logs already captured (
docker logs, application logs)
- Metrics already collected (Collector service, SQLite)
- Dashboard shows FL round progress, policy decisions
- Topology changes tracked by SDN controller
Evidence: Most "events" are just metric snapshots
Option B: Events are UNDERDEVELOPED
If events provide unique value, they need:
- Clear event taxonomy
- Comprehensive event generation
- Proper storage (events.jsonl never used)
- Research/debugging workflows
- Integration with critical actions
Suggested Improvements
If Keeping Events (Option B - Recommended)
Define clear event taxonomy:
# Event types that provide unique value
class SystemEventType(Enum):
# FL Training Events (actionable)
FL_ROUND_STARTED = "fl_round_started"
FL_ROUND_COMPLETED = "fl_round_completed"
FL_CLIENT_CONNECTED = "fl_client_connected"
FL_CLIENT_DISCONNECTED = "fl_client_disconnected"
FL_CLIENT_FAILED = "fl_client_failed"
FL_MODEL_AGGREGATED = "fl_model_aggregated"
# Policy Events (compliance/audit)
POLICY_EVALUATED = "policy_evaluated"
POLICY_ALLOWED = "policy_allowed"
POLICY_DENIED = "policy_denied"
POLICY_VIOLATION = "policy_violation"
# Network Events (troubleshooting)
NODE_ADDED = "node_added"
NODE_REMOVED = "node_removed"
LINK_UP = "link_up"
LINK_DOWN = "link_down"
NETWORK_PARTITION = "network_partition"
# System Events (operations)
SCENARIO_STARTED = "scenario_started"
SCENARIO_COMPLETED = "scenario_completed"
SERVICE_STARTED = "service_started"
SERVICE_CRASHED = "service_crashed"
CONFIGURATION_CHANGED = "config_changed"
Implement proper event logging:
# src/core/events/event_logger.py
class EventLogger:
"""Centralized event logging to events.jsonl."""
def __init__(self, log_file: str = "logs/events.jsonl"):
self.log_file = Path(log_file)
self.log_file.parent.mkdir(exist_ok=True)
def log_event(
self,
event_type: SystemEventType,
component: str,
details: Dict[str, Any],
severity: str = "INFO"
):
"""Log structured event to JSONL file."""
event = {
"timestamp": datetime.utcnow().isoformat(),
"event_type": event_type.value,
"component": component,
"severity": severity,
"details": details,
"version": "1.0.0"
}
with open(self.log_file, 'a') as f:
f.write(json.dumps(event) + '\n')
# Global instance
event_logger = EventLogger()
Integrate with all components:
# In FL Server
from src.core.events.event_logger import event_logger, SystemEventType
def on_round_start(round_num: int):
event_logger.log_event(
SystemEventType.FL_ROUND_STARTED,
component="fl_server",
details={"round": round_num, "clients_selected": client_count}
)
# In Policy Engine
def evaluate_policy(policy_id: str, context: dict) -> dict:
result = _evaluate(policy_id, context)
event_logger.log_event(
SystemEventType.POLICY_EVALUATED if result['allowed'] else SystemEventType.POLICY_DENIED,
component="policy_engine",
details={
"policy_id": policy_id,
"allowed": result['allowed'],
"client_id": context.get('client_id')
},
severity="WARNING" if not result['allowed'] else "INFO"
)
return result
Create event analysis tools:
# scripts/analyze_events.py
def analyze_events(log_file: str):
"""Analyze event log for patterns and issues."""
events = []
with open(log_file) as f:
for line in f:
events.append(json.loads(line))
# Count by type
type_counts = Counter(e['event_type'] for e in events)
# Find anomalies
policy_denials = [e for e in events if e['event_type'] == 'policy_denied']
client_failures = [e for e in events if e['event_type'] == 'fl_client_failed']
# Timeline analysis
round_starts = [e for e in events if e['event_type'] == 'fl_round_started']
round_durations = _compute_durations(round_starts)
return {
'total_events': len(events),
'by_type': type_counts,
'anomalies': {
'policy_denials': len(policy_denials),
'client_failures': len(client_failures)
},
'round_durations': round_durations
}
If Removing Events (Option A)
Simplify to use existing monitoring:
- Remove
EventsPage.tsx from dashboard (738 lines)
- Remove
event_monitor.py from collector
- Remove
logs/events.jsonl references
- Use logs for debugging (already comprehensive)
- Use metrics for analysis (already collected)
- Use dashboard tabs for monitoring (already exists)
Benefits of removal:
- Less code to maintain
- Clearer separation: logs for debugging, metrics for analysis
- One less storage format to manage
- Simpler architecture
Recommendation
Option B: Improve event system for these reasons:
- Research value: Events provide timeline of FL training for papers
- Debugging: Event sequences help troubleshoot complex issues
- Audit trail: Policy decisions need immutable record
- Reproducibility: Event logs enable experiment replay
But needs work:
- ⚠️ Define clear event taxonomy
- ⚠️ Implement comprehensive event generation
- ⚠️ Fix
events.jsonl writing (currently broken)
- ⚠️ Add event analysis tools
- ⚠️ Document use cases for researchers
Implementation Plan
Phase 1: Fix Basic Infrastructure (Week 1)
Phase 2: Instrument Components (Week 2)
Phase 3: Analysis Tools (Week 3)
Phase 4: Dashboard Integration (Week 4)
Acceptance Criteria (If Improving)
Acceptance Criteria (If Removing)
Related Issues
Priority
Recommended: Medium
Current state is confusing. Either fix properly or remove entirely.
Target Version: v1.1.0
Additional Notes
Current state: ⚠️ Half-implemented feature that provides limited value
Decision needed: Invest in making events useful OR simplify by removing them
Recommendation: Fix it - event logs have genuine research value for FL systems, but implementation needs work.
Issue Description
The Live System Event tracking feature exists in the dashboard but is either unnecessary (duplicates other monitoring) or underdeveloped (limited functionality, unclear value proposition).
Evidence
Event tracking EXISTS but with limitations:
Dashboard Event Page (
dashboard/frontend/src/pages/EventsPage.tsx)Backend Event API (
dashboard/backend/app/services/collector_client.py)get_system_events()methodCollector Event Storage (
src/collector/event_monitor.py)Empty events.jsonl file:
ls -l logs/events.jsonl # File exists but is empty (0 bytes)What's NOT working:
logs/events.jsonlis never populatedVerified with:
Problem Statement
Key questions without clear answers:
Current state analysis:
Option A: Events are UNNECESSARY
If events duplicate existing monitoring:
docker logs, application logs)Evidence: Most "events" are just metric snapshots
Option B: Events are UNDERDEVELOPED
If events provide unique value, they need:
Suggested Improvements
If Keeping Events (Option B - Recommended)
Define clear event taxonomy:
Implement proper event logging:
Integrate with all components:
Create event analysis tools:
If Removing Events (Option A)
Simplify to use existing monitoring:
EventsPage.tsxfrom dashboard (738 lines)event_monitor.pyfrom collectorlogs/events.jsonlreferencesBenefits of removal:
Recommendation
Option B: Improve event system for these reasons:
But needs work:
events.jsonlwriting (currently broken)Implementation Plan
Phase 1: Fix Basic Infrastructure (Week 1)
events.jsonlwriting (currently never populated)Phase 2: Instrument Components (Week 2)
Phase 3: Analysis Tools (Week 3)
Phase 4: Dashboard Integration (Week 4)
Acceptance Criteria (If Improving)
logs/events.jsonlpopulated with eventsAcceptance Criteria (If Removing)
events.jsonlRelated Issues
Priority
Recommended: Medium
Current state is confusing. Either fix properly or remove entirely.
Target Version: v1.1.0
Additional Notes
Current state:⚠️ Half-implemented feature that provides limited value
Decision needed: Invest in making events useful OR simplify by removing them
Recommendation: Fix it - event logs have genuine research value for FL systems, but implementation needs work.