This project implements an end-to-end Optical Character Recognition (OCR) system that combines:
- YOLO for text line detection in images and videos
- OCR model for text recognition from the detected text regions
The system follows a two-stage pipeline:
- Detect text regions using YOLO
- Crop detected regions and recognize text using an OCR model
This design allows the detection and recognition components to be trained, optimized, and replaced independently, making the system flexible and scalable for real-world OCR applications.
Input Image / Video
│
▼
YOLO (Text Line Detection)
│
▼
Crop Detected Text Regions
│
▼
OCR Model (Text Recognition)
│
▼
Recognized Text Output
.
├── Create_OCR_data.py # Generate OCR training dataset
├── Create_Yolo_data.py # Generate YOLO training dataset
├── create_ocr_data.yaml # OCR data generation config
├── create_yolo_data_config.yaml # YOLO data generation config
├── Train_ocr_from_scratch.py # Train OCR model from scratch
├── Train_yolo.py # Train YOLO model
├── train_ocr_config.yaml # OCR training config
├── train_yolo_data.yaml # YOLO training config
├── System_Inference.py # End-to-end inference pipeline
│
├── model/ # Trained models
├── ocr_dataset/ # OCR dataset
├── runs/ # Training logs and outputs
├── SceneTrialTrain/ # YOLO dataset
├── video/ # Demo videos
└── README.md
First, clone this repository to your local machine:
git clone https://github.com/VyDat-1702/OCR-YOLOv11.git
cd OCR-YOLOv11/- Python ≥ 3.10
- Conda (recommended)
- NVIDIA GPU with CUDA support (optional but recommended)
Create and activate a conda environment:
conda create -n pytorch_env python=3.12
conda activate pytorch_envInstall dependencies:
pip install -r requirements.txtpython Create_Yolo_data.py --config create_yolo_data_config.yamlThis script prepares the dataset for training the YOLO text detection model.
python Create_OCR_data.py --config create_ocr_data.yamlThis script generates cropped text images and corresponding labels for OCR training.
python Train_yolo.py --config train_yolo_data.yamlTraining results (weights, logs) are saved in: runs/
python Train_ocr_from_scratch.py --config train_ocr_config.yamlThe trained OCR model is saved in the model/ directory.
Run the complete OCR pipeline on images or videos:
python System_Inference.pyOutput:
- Images or videos with detected text regions
- Recognized text results for each detected region
- It is not recommended to commit large trained model files (
.pt) directly to the repository. - Use Git LFS or download trained models from releases or external storage if needed.
- Ensure you have sufficient disk space for datasets and training outputs.
