Production-Oriented QLoRA Fine-Tuning & LLMOps Pipeline
Fine-tuned TinyLlama using QLoRA, PEFT, TRL, and the HuggingFace ecosystem
with a deployment-oriented inference architecture.
Agent-Forge is a production-oriented GenAI engineering project that implements the complete lifecycle of modern LLM adaptation and deployment β from raw dataset to containerized inference server.
π― Goal: Deeply understand and implement end-to-end LLM fine-tuning & deployment infrastructure.
|
π§ ML Engineering
|
βοΈ LLMOps & Infra
|
HuggingFace Dataset
β
βΌ
Conversational Formatting
β
βΌ
Tokenizer (TinyLlama)
β
βΌ
4-bit NF4 Quantization βββββ BitsAndBytes
β
βΌ
QLoRA + PEFT Adapter Injection
β
βΌ
SFT Training (SFTTrainer / TRL)
β
βΌ
Inference Evaluation (Before vs After)
β
βΌ
LoRA Adapter Saving
β
βΌ
FastAPI Inference Server
β
βΌ
Redis + Qdrant + Docker Deployment
| Property | Value |
|---|---|
| π€ Model | TinyLlama/TinyLlama-1.1B-Chat-v1.0 |
| π¦ Parameters | 1.1 Billion |
| β‘ Quantization | 4-bit NF4 |
| π― Fine-Tuning Method | QLoRA + PEFT |
- β Lightweight and experimentation-friendly
- β QLoRA compatible out of the box
- β Runs on free Colab GPU tiers
- β Fast iteration cycles
- β Ideal for mastering modern fine-tuning pipelines
| Setting | Value |
|---|---|
| Quantization | 4-bit NF4 |
| Double Quantization | β Enabled |
| Compute Dtype | float16 |
| Backend | BitsAndBytes |
Benefits:
- π Dramatically lower VRAM usage
- β‘ Faster model loading
- π§© Memory-efficient training
- β Colab-compatible fine-tuning
| Parameter | Value |
|---|---|
Rank r |
8 |
Alpha Ξ± |
32 |
| Dropout | 0.05 |
| Precision | float16 |
| Quantization | 4-bit NF4 |
| Target Modules | q_proj, k_proj, v_proj, o_proj |
ββββββββββββββββββββββββββββββββββββββββββββ
β Parameter Efficiency β
β βββββββββββββββββββββββββββββββββββββββββββ£
β Trainable Params : 2,252,800 β
β Total Params : 1,102,301,184 β
β Trainable % : 0.2044% β
ββββββββββββββββββββββββββββββββββββββββββββ
π‘ Key Insight: 99.8% of the base model remained completely frozen. Only lightweight LoRA adapters were trained β maximum efficiency, minimal compute.
bitext/Bitext-customer-support-llm-chatbot-training-dataset
Each training sample contains:
| Field | Description |
|---|---|
instruction |
The user query / task |
category |
Domain category |
intent |
Intent classification |
response |
Target model response |
Samples were transformed into instruction-tuned conversational format aligned with TinyLlama's chat template.
Dataset
β
βββΊ Formatting (conversational template)
β
βββΊ Tokenization
β
βββΊ QLoRA Adapter Injection
β
βββΊ SFT Training (via TRL SFTTrainer)
β
βββΊ Inference Evaluation (Before vs After)
β
βββΊ LoRA Adapter Saving
The same prompts were run in two passes to directly compare behavior:
Base TinyLlama QLoRA Fine-Tuned TinyLlama
βββββββββββββββββ ββββββββββββββββββββββββββββ
Generic responses vs. Domain-adapted responses
No task grounding vs. Customer support grounded
Prompt-agnostic vs. Instruction-tuned behavior
Endpoint:
POST /generate
Content-Type: application/json
{
"prompt": "How do I reset my password?"
}Inference Flow:
Request β Tokenizer β TinyLlama + LoRA β Generation β JSON Response
Planned caching layer for:
- β‘ Repeated prompt acceleration
- π Latency reduction
- π Throughput optimization
- πΎ Inference result storage
Planned migration toward a Kubernetes Troubleshooting Assistant:
User Query
β
βΌ
Embedding Model
β
βΌ
Qdrant Vector Retrieval
β
βΌ
Context Injection into Prompt
β
βΌ
Fine-Tuned LLM β Response
| Component | Technology |
|---|---|
| π API Server | FastAPI |
| π¦ Containerization | Docker |
| ποΈ Caching | Redis |
| π Vector Store | Qdrant |
| π€ Model Hosting | HuggingFace Spaces |
| π CI/CD | GitHub Actions |
π§ ML Concepts
- Tokenization & chat templates
- 4-bit NF4 quantization mechanics
- QLoRA parameter efficiency
- PEFT & LoRA adapter training
- Transformer fine-tuning workflows
- SFTTrainer from TRL
βοΈ Engineering Concepts
- GPU memory optimization strategies
- HuggingFace ecosystem integration
- Dependency & version conflict debugging
- Inference evaluation pipelines
- Production LLMOps design patterns
- π‘ vLLM Inference Serving
- π Weights & Biases Experiment Tracking
- π§ͺ MLflow Logging Integration
- π Streaming Responses via SSE
- π§ Multi-turn Conversation Memory
- βΈοΈ Kubernetes-Specialized Fine-Tuning
- π Fully Automated CI/CD Deployment
Made with π₯ by Vivek Tyagi β AI Engineer
If you found this useful, drop a β β it means a lot!