A lightweight Node.js OCR microservice that extracts text from images using Tesseract.js.
The service is designed to process images generated from PDF pages and return both page-level text and combined document text. It includes image preprocessing with Sharp, JWT authentication, service-to-service authentication, Swagger documentation, and Docker support.
- OCR-based text extraction from images
- Multiple image/page processing
- Page-level extracted text
- Combined full-text response
- Image preprocessing using Sharp
- Grayscale, normalization, sharpening, and thresholding
- Tesseract.js OCR engine
- English language trained data
- JWT Bearer authentication
- Service-to-service authentication using
X-SERVICE-KEY - Health check endpoint
- Swagger/OpenAPI documentation
- Docker support
- In-memory multipart file processing
- Per-page file size validation
- Temporary processing without storing uploaded images on disk
- Node.js 20
- Express.js
- JavaScript
- Tesseract.js
- Sharp
- Multer
- JWT
- Swagger / OpenAPI
- Docker
PDF Pages / Images
|
v
+----------------------------+
| Text Extraction API |
| Express.js |
+-------------+--------------+
|
v
+----------------------------+
| Image Preprocessing |
| Sharp |
| |
| Resize / Grayscale |
| Normalize / Sharpen |
| Threshold |
+-------------+--------------+
|
v
+----------------------------+
| Tesseract.js |
| OCR |
+-------------+--------------+
|
v
+----------------------------+
| Page Text + Full Text |
+----------------------------+
Text-Extraction-Service/
│
├── src/
│ ├── config/
│ │ └── swagger.js
│ │
│ ├── controllers/
│ │ └── ocr.controller.js
│ │
│ ├── docs/
│ │ └── ocr.swagger.js
│ │
│ ├── middlewares/
│ │ └── auth.middleware.js
│ │
│ ├── routes/
│ │ └── ocr.routes.js
│ │
│ ├── services/
│ │ └── ocr.service.js
│ │
│ └── server.js
│
├── tessdata/
│ └── eng.traineddata
│
├── Dockerfile
├── package.json
├── package-lock.json
└── .dockerignore
POST /api/text-extraction/extractThe endpoint accepts multiple images using multipart/form-data.
Authorization: Bearer <JWT_TOKEN>
X-SERVICE-KEY: <SERVICE_KEY>The form field name is:
images
Example using cURL:
curl -X POST \
http://localhost:3000/api/text-extraction/extract \
-H "Authorization: Bearer <JWT_TOKEN>" \
-H "X-SERVICE-KEY: <SERVICE_KEY>" \
-F "images=@page-1.png" \
-F "images=@page-2.png"Multiple images can be uploaded in the same request.
A successful request returns:
{
"success": true,
"pages": [
{
"page": 1,
"text": "Text extracted from page one..."
},
{
"page": 2,
"text": "Text extracted from page two..."
}
],
"fullText": "Text extracted from page one...\n\nText extracted from page two..."
}Field Description
success Indicates whether OCR processing completed successfully
pages Contains OCR results for each uploaded image
page Page/image sequence number
text Text extracted from the individual page
fullText All page text combined in document order
Images
|
v
Validate Files
|
v
Image Preprocessing
|
+--> Resize
+--> Grayscale
+--> Normalize
+--> Sharpen
+--> Threshold
|
v
Tesseract.js
|
v
Extract Text
|
v
Page-Level Results
|
v
Combined Full Text
For images larger than 500 KB, the service applies preprocessing using Sharp:
Resize to maximum width of 1000px
↓
Grayscale
↓
Normalize contrast
↓
Sharpen
↓
Threshold
This preprocessing is intended to improve OCR quality while reducing unnecessary image size.
Images smaller than 500 KB are passed directly to the OCR engine.
The service uses Tesseract.js with the English trained data:
tessdata/eng.traineddata
OCR configuration includes:
Page Segmentation Mode: 3
OCR Engine Mode: 1
Preserve Interword Spaces: enabled
User Defined DPI: 300
The service initializes the OCR worker before starting the HTTP server.
Each uploaded image is validated before OCR processing.
- Maximum image size: 3 MB per page
- Empty or very small files are skipped
- Images are processed in batches
- Current worker count: 1
- Current batch size: 1
These values can be adjusted in:
src/services/ocr.service.js
Protected requests require a valid JWT:
Authorization: Bearer <JWT_TOKEN>The token is verified using the configured JWT_SECRET.
The service also requires:
X-SERVICE-KEY: <SERVICE_KEY>The service key is validated before JWT authentication and OCR processing.
This provides an additional security layer for communication between backend services or through an API Gateway.
The service exposes:
GET /healthExample response:
{
"status": "OK"
}This endpoint can be used by deployment platforms and monitoring systems to verify service availability.
Swagger documentation is configured in the application.
When running locally, open the Swagger UI URL configured by the application, typically:
http://localhost:3000/api-docs
Swagger configuration is located in:
src/config/swagger.js
API documentation:
src/docs/ocr.swagger.js
The service uses environment variables for runtime configuration.
PORT
JWT_SECRET
SERVICE_KEY
CORS_ALLOWED_ORIGINS
Example:
PORT=3000
JWT_SECRET=your-jwt-secret
SERVICE_KEY=your-service-key
CORS_ALLOWED_ORIGINS=http://localhost:4200Multiple CORS origins can be provided as a comma-separated list:
CORS_ALLOWED_ORIGINS=http://localhost:4200,https://example.comDo not commit production secrets to source control.
Install:
- Node.js 20+
- npm
git clone https://github.com/tejaspatil-web/Text-Extraction-Service.git
cd Text-Extraction-Servicenpm installCreate a .env file:
PORT=3000
JWT_SECRET=your-jwt-secret
SERVICE_KEY=your-service-key
CORS_ALLOWED_ORIGINS=http://localhost:4200npm startFor development:
npm run devThe API will be available at:
http://localhost:3000
Health check:
http://localhost:3000/health
The repository includes a Dockerfile based on Node.js 20 Slim.
docker build -t text-extraction-service .docker run -p 3000:3000 \
-e PORT=3000 \
-e JWT_SECRET="your-jwt-secret" \
-e SERVICE_KEY="your-service-key" \
-e CORS_ALLOWED_ORIGINS="http://localhost:4200" \
text-extraction-serviceThe service will be available at:
http://localhost:3000
This service can be used as part of a document-processing pipeline.
For example:
PDF
|
v
PDF-to-PNG Service
|
v
PNG Pages
|
v
Text Extraction Service
|
v
OCR Text
|
v
RAG / LLM Processing
This separation allows PDF conversion and OCR processing to be independently maintained and deployed.
400 Bad Request{
"error": "No images uploaded"
}401 Unauthorized{
"message": "Service key missing"
}403 Forbidden{
"message": "Invalid service key"
}401 Unauthorized500 Internal Server Error{
"error": "OCR failed",
"details": "..."
}This service can be used for:
- PDF text extraction
- OCR processing
- Document digitization
- Resume processing
- AI document processing
- RAG document ingestion
- Scanned document processing
- Image-to-text conversion
- Document analysis pipelines
The service uses in-memory file uploads through Multer, so uploaded images do not need to be persisted to disk before OCR processing.
Tesseract workers are initialized during application startup. The server starts listening only after OCR initialization completes successfully.
The OCR service is implemented as an asynchronous generator so batch progress can be yielded internally while processing multiple pages.
This project is available for personal and educational use.
If you plan to distribute or reuse this project, add an appropriate open-source license such as MIT.