Skip to content

Repository files navigation

Nigerian Commodity Price Volatility Forecasting

Live Dashboard: http://forecasting2.vercel.app/

This project contains an institutional-grade data engineering pipeline and machine learning forecasting model designed to predict the realized volatility of key Nigerian commodities (Maize, Rice, PMS/Gasoline, and Diesel) driven by local and global macroeconomic indicators.

Project Structure

  • datapreparation.ipynb: The end-to-end data pipeline. It orchestrates API pulls from FRED and FEWS NET, integrates local CSV/Excel drops for Nigerian macro indicators, cleans outliers, handles missing gaps with PCHIP/Linear interpolation, and computes the realized volatility targets.
  • model_building.ipynb: The modeling and analysis engine. It handles feature scaling, Principal Component Analysis (PCA) for dimensionality reduction, creates autoregressive time lags, and trains both linear (Elastic Net) and non-linear (Random Forest) models.
  • output/master_dataset.csv: The finalized, mathematically clean dataset containing all features and targets from 2018 onwards (generated by the datapreparation notebook).
  • mlruns/: The MLflow tracking directory storing every model iteration, hyperparameter, and diagnostic metric.

Key Data Sources

  1. FEWS NET FDW API: Wholesale local agricultural prices (Maize, Rice).
  2. FRED API: Global macroeconomic factors (Brent Crude, US 10T Treasury Yield) and CBN Official M2/FX metrics.
  3. NBS (National Bureau of Statistics): Official state-level and aggregated retail energy prices (PMS, Diesel).
  4. Kaggle: Nigerian Parallel Market operations to calculate the critical FX Premium.

Setup & How to Run

Prerequisites

You will need a .env file containing your FRED API Key:

FRED_API_KEY=your_key_here

Libraries

Ensure your environment (or Google Colab) has the following installed:

pip install pandas numpy fredapi python-dotenv scikit-learn matplotlib seaborn plotly mlflow

Execution Order

  1. Run datapreparation.ipynb: Execute all cells top to bottom. It will query the APIs, ingest your local Excel data, scrub the mathematical anomalies, and output the clean master_dataset.csv.
  2. Run model_building.ipynb: Execute all cells. It will load the master dataset, perform PCA/Feature Engineering, start the MLflow tracking server, train both Elastic Net and Random Forests across all 4 commodities, and generate the Forecast and Feature Importance plots.

Viewing MLOps Diagnostics

Because this project utilizes mlflow, you can view the local tracking UI to compare models:

  1. Open a terminal in the project directory.
  2. Run mlflow ui.
  3. Open http://localhost:5000 in your browser.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages