An end-to-end data analytics project analysing electricity reliability, fault frequency, and outage duration using Nigerian electricity data. The project combines structured monthly electricity data with recorded fault/outage events and transforms them into an interactive Power BI reliability intelligence dashboard.
Electricity reliability isn't only about how much electricity is recorded.
A period can have a relatively high electricity value and still experience significant reliability problems if faults are frequent or outages take longer to restore.
This project was built to examine that relationship by bringing together:
- Monthly NBS electricity data
- Mendeley fault/outage records
- Fault frequency
- Outage duration
- Monthly reliability metrics
- A reliability risk score
- Interactive Power BI analysis
The objective was to move beyond isolated charts and create a model that makes it easier to identify periods of higher reliability risk and understand the underlying outage patterns.
The analysis was designed to:
- Examine electricity trends over time.
- Analyse recorded fault frequency.
- Measure outage duration.
- Compare total, average, and maximum outage duration.
- Identify periods with higher reliability risk.
- Create a consistent monthly analytical dataset.
- Allow reliability patterns to be explored by year.
- Build an interactive Power BI dashboard for decision-oriented analysis.
The NBS dataset covers:
2015–2025
The dataset contains 12 monthly observations for each year, giving:
132 monthly observations
The NBS data was cleaned and structured into a consistent monthly date format before being incorporated into the analytical model.
The Mendeley dataset contains recorded electricity fault/outage events covering:
2017–2024
The cleaned dataset contains:
2,324 fault records
The fault data includes reported and cleared/restoration information used to prepare the outage-duration analysis.
The available fault records were converted into usable datetime fields where the underlying date/time information supported the calculation.
The project was prepared using Python and Pandas before being brought into Power BI.
The workflow included:
Raw Data
↓
Data Inspection
↓
Data Cleaning
↓
Date Validation
↓
NBS Monthly Structure
↓
Mendeley Fault Preparation
↓
Reported / Restored Datetime Preparation
↓
Outage Duration Calculation
↓
Monthly Aggregation
↓
Reliability Metrics
↓
Power BI Data Model
↓
Interactive Dashboard
The NBS data was inspected for:
- Sheet structure
- Date format
- Monthly coverage
- Yearly coverage
- Missing values
- Duplicate records
The final NBS date structure was validated to ensure that each year from 2015 through 2025 contained 12 monthly observations.
The Mendeley fault data required additional preparation because reported and restoration information was stored in separate date and time fields.
The preparation included:
- Standardising date fields.
- Preparing reported datetime values.
- Preparing restored datetime values.
- Calculating outage duration where valid datetime information was available.
- Checking for negative outage durations.
- Validating the resulting fault-level data before aggregation.
The final monthly analytical structure contains the following fields:
| Field | Description |
|---|---|
Date |
Monthly analysis date |
NBS_Total_Value |
Monthly NBS value used in the analysis |
Fault_Count |
Number of recorded faults for the period |
Total_Outage_Hours |
Combined outage duration for the period |
Average_Outage_Hours |
Average outage duration for the period |
Maximum_Outage_Hours |
Longest recorded outage duration for the period |
Reliability Risk Score |
Risk indicator used in the Power BI analysis |
The project also uses Year and Month fields for time-based analysis and filtering.
Rather than treating fault count as the only indicator of reliability, the project examines multiple dimensions of outage performance.
Measures how many recorded fault events occurred during a period.
Measures the combined duration of recorded outages within a period.
Measures the average duration of recorded outages.
Identifies the longest recorded outage duration within a period.
A risk indicator was incorporated into the Power BI model to support the identification and comparison of higher-risk periods.
The score is used as an analytical indicator rather than being presented as an official industry reliability standard.
The final Power BI report contains five focused sections.
The Executive Overview provides the high-level view of the electricity reliability analysis.
It is designed to answer:
- What does the overall reliability picture look like?
- How do electricity values and reliability indicators change over time?
- Which high-level metrics require attention?
This page moves from the high-level overview into the relationship between faults and outages.
The analysis focuses on:
- Fault frequency
- Outage duration
- Reliability trends
- Changes across the analysis period
The purpose is to make the underlying reliability behaviour easier to explore rather than relying only on headline KPIs.
This page focuses on analytical interpretation and risk.
A dynamic Key Reliability Finding responds to the selected year and identifies the highest-risk period within the selected context.
The dashboard can surface:
- The highest-risk period
- Recorded fault count
- Total outage hours
- Reliability risk score
This allows the finding to change dynamically when the year selection changes rather than relying on a static written observation.
The Executive Summary is designed for users who need the most important reliability information quickly.
It prioritises the major indicators and provides a concise view of the analysis without requiring users to navigate through the detailed fault-level information.
The Fault Details page provides access to the underlying fault-level analysis through the report's drill-through functionality.
This keeps the main analytical pages clean while allowing users to move from an identified reliability issue into the supporting fault information.
The Power BI report includes interactive analysis through:
- Year-based filtering
- Date filtering where applicable
- Dynamic DAX measures
- Dynamic key reliability findings
- Drill-through analysis
- Cross-filtering between visuals
- Reliability risk analysis
The dynamic finding is particularly useful because selecting a different year changes the reliability finding instead of displaying a fixed conclusion.
- Python
- Pandas
- Jupyter Notebook
- Python
- Pandas
- DAX
- Microsoft Power BI
- Power Query
- DAX
- NBS electricity data
- Mendeley fault/outage dataset
electricity-reliability-analysis/
│
├── README.md
│
├── data/
│ └── README.md
│
├── notebooks/
│ └── electricity_reliability_analysis.ipynb
│
├── powerbi/
│ └── Electricity_Reliability_Dashboard.pbix
│
├── screenshots/
│ ├── executive-overview.png
│ ├── reliability-outage-analysis.png
│ ├── key-insights-risk-analysis.png
│ ├── executive-summary.png
│ └── fault-details.png
│
└── docs/
└── methodology.md
Raw source datasets are not included in this repository unless redistribution rights permit their publication.
Data validation was performed before the datasets were used for dashboard analysis.
The validation process included:
- Checking missing values.
- Checking duplicate records.
- Validating date ranges.
- Confirming NBS monthly coverage.
- Checking Mendeley date coverage.
- Validating reported/restored datetime fields.
- Checking outage duration calculations.
- Checking for negative outage durations.
- Validating the monthly aggregation before importing the analytical dataset into Power BI.
The NBS dataset was confirmed to contain:
12 months per year from 2015 through 2025.
The Mendeley records used in the analysis cover:
2017 through 2024.
The NBS and Mendeley datasets represent different aspects of electricity performance and do not constitute a single identical measurement system.
The analysis therefore does not claim that the NBS values directly represent outage performance.
Instead, the project uses the available datasets to examine electricity values alongside recorded fault and outage behaviour within a common time-based analytical model.
The Mendeley dataset also contains periods where complete reported/restored time information is unavailable. Outage-duration metrics therefore depend on records with usable datetime information.
The reliability risk score is an analytical construct used within this project and should not be interpreted as an official utility reliability index.
The central idea behind the project is simple:
Electricity reliability cannot be understood from a single number.
Fault frequency tells one part of the story.
Outage duration tells another.
Looking at them together provides a more useful way to identify periods that deserve further investigation.
The purpose of the dashboard is therefore not simply to display electricity data, but to turn multiple reliability indicators into an interactive analytical view.
The completed project transforms raw electricity and fault information into a structured reliability analysis containing:
- 11 years of monthly NBS data
- 132 monthly NBS observations
- 2,324 recorded Mendeley fault events
- Monthly fault and outage metrics
- Reliability risk analysis
- Dynamic Power BI findings
- Interactive year-based analysis
- Drill-through fault analysis
- A five-page Power BI dashboard
The project demonstrates an end-to-end workflow from raw data → cleaning → validation → analytical modelling → DAX → Power BI → reliability intelligence.
Akpan Daniel
Data Analyst | Power BI | Python | Pandas | Data Analytics