Skip to content

Latest commit

Β 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Typing SVG

Production-Oriented QLoRA Fine-Tuning & LLMOps Pipeline
Fine-tuned TinyLlama using QLoRA, PEFT, TRL, and the HuggingFace ecosystem
with a deployment-oriented inference architecture.


Python PyTorch HuggingFace FastAPI Docker Redis Qdrant


Stars License: MIT


πŸš€ Overview

Agent-Forge is a production-oriented GenAI engineering project that implements the complete lifecycle of modern LLM adaptation and deployment β€” from raw dataset to containerized inference server.

🎯 Goal: Deeply understand and implement end-to-end LLM fine-tuning & deployment infrastructure.

🧠 ML Engineering

  • QLoRA Fine-Tuning
  • 4-bit NF4 Quantization
  • PEFT / LoRA Adapters
  • Supervised Fine-Tuning (SFT)
  • Conversational Dataset Engineering

βš™οΈ LLMOps & Infra

  • FastAPI Inference Serving
  • Redis Caching Architecture
  • Qdrant RAG Integration
  • Docker Containerization
  • Deployment-Ready Pipelines

πŸ—οΈ Architecture

  HuggingFace Dataset
          β”‚
          β–Ό
  Conversational Formatting
          β”‚
          β–Ό
  Tokenizer (TinyLlama)
          β”‚
          β–Ό
  4-bit NF4 Quantization  ◄──── BitsAndBytes
          β”‚
          β–Ό
  QLoRA + PEFT Adapter Injection
          β”‚
          β–Ό
  SFT Training  (SFTTrainer / TRL)
          β”‚
          β–Ό
  Inference Evaluation (Before vs After)
          β”‚
          β–Ό
  LoRA Adapter Saving
          β”‚
          β–Ό
  FastAPI Inference Server
          β”‚
          β–Ό
  Redis + Qdrant + Docker Deployment

🧠 Base Model

Property Value
πŸ€– Model TinyLlama/TinyLlama-1.1B-Chat-v1.0
πŸ“¦ Parameters 1.1 Billion
⚑ Quantization 4-bit NF4
🎯 Fine-Tuning Method QLoRA + PEFT

Why TinyLlama?

  • βœ… Lightweight and experimentation-friendly
  • βœ… QLoRA compatible out of the box
  • βœ… Runs on free Colab GPU tiers
  • βœ… Fast iteration cycles
  • βœ… Ideal for mastering modern fine-tuning pipelines

⚑ QLoRA Configuration

Setting Value
Quantization 4-bit NF4
Double Quantization βœ… Enabled
Compute Dtype float16
Backend BitsAndBytes

Benefits:

  • πŸ“‰ Dramatically lower VRAM usage
  • ⚑ Faster model loading
  • 🧩 Memory-efficient training
  • βœ… Colab-compatible fine-tuning

🧩 PEFT / LoRA Setup

Parameter Value
Rank r 8
Alpha Ξ± 32
Dropout 0.05
Precision float16
Quantization 4-bit NF4
Target Modules q_proj, k_proj, v_proj, o_proj

πŸ“Š Training Metrics

╔══════════════════════════════════════════╗
β•‘         Parameter Efficiency             β•‘
╠══════════════════════════════════════════╣
β•‘  Trainable Params   :     2,252,800      β•‘
β•‘  Total Params       : 1,102,301,184      β•‘
β•‘  Trainable %        :        0.2044%     β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•

πŸ’‘ Key Insight: 99.8% of the base model remained completely frozen. Only lightweight LoRA adapters were trained β€” maximum efficiency, minimal compute.


πŸ“š Dataset

bitext/Bitext-customer-support-llm-chatbot-training-dataset

Each training sample contains:

Field Description
instruction The user query / task
category Domain category
intent Intent classification
response Target model response

Samples were transformed into instruction-tuned conversational format aligned with TinyLlama's chat template.


πŸ”₯ Training Pipeline

Dataset
  β”‚
  β”œβ”€β–Ί Formatting  (conversational template)
  β”‚
  β”œβ”€β–Ί Tokenization
  β”‚
  β”œβ”€β–Ί QLoRA Adapter Injection
  β”‚
  β”œβ”€β–Ί SFT Training  (via TRL SFTTrainer)
  β”‚
  β”œβ”€β–Ί Inference Evaluation  (Before vs After)
  β”‚
  └─► LoRA Adapter Saving

πŸ§ͺ Evaluation Strategy

The same prompts were run in two passes to directly compare behavior:

 Base TinyLlama                 QLoRA Fine-Tuned TinyLlama
 ─────────────────              ────────────────────────────
 Generic responses       vs.    Domain-adapted responses
 No task grounding       vs.    Customer support grounded
 Prompt-agnostic         vs.    Instruction-tuned behavior

βš™οΈ Inference Server

Endpoint:

POST /generate
Content-Type: application/json

{
  "prompt": "How do I reset my password?"
}

Inference Flow:

Request β†’ Tokenizer β†’ TinyLlama + LoRA β†’ Generation β†’ JSON Response

πŸ—‚οΈ Redis Caching Layer

Planned caching layer for:

  • ⚑ Repeated prompt acceleration
  • πŸ“‰ Latency reduction
  • πŸ“ˆ Throughput optimization
  • πŸ’Ύ Inference result storage

πŸ”Ž Qdrant RAG Integration

Planned migration toward a Kubernetes Troubleshooting Assistant:

User Query
    β”‚
    β–Ό
Embedding Model
    β”‚
    β–Ό
Qdrant Vector Retrieval
    β”‚
    β–Ό
Context Injection into Prompt
    β”‚
    β–Ό
Fine-Tuned LLM β†’ Response

🐳 Deployment Stack

Component Technology
🌐 API Server FastAPI
πŸ“¦ Containerization Docker
πŸ—‚οΈ Caching Redis
πŸ” Vector Store Qdrant
πŸ€— Model Hosting HuggingFace Spaces
πŸ”„ CI/CD GitHub Actions

πŸ“ˆ Key Learnings

🧠 ML Concepts
  • Tokenization & chat templates
  • 4-bit NF4 quantization mechanics
  • QLoRA parameter efficiency
  • PEFT & LoRA adapter training
  • Transformer fine-tuning workflows
  • SFTTrainer from TRL
βš™οΈ Engineering Concepts
  • GPU memory optimization strategies
  • HuggingFace ecosystem integration
  • Dependency & version conflict debugging
  • Inference evaluation pipelines
  • Production LLMOps design patterns

πŸ›£οΈ Roadmap

  • πŸ“‘ vLLM Inference Serving
  • πŸ“Š Weights & Biases Experiment Tracking
  • πŸ§ͺ MLflow Logging Integration
  • πŸ”„ Streaming Responses via SSE
  • 🧠 Multi-turn Conversation Memory
  • ☸️ Kubernetes-Specialized Fine-Tuning
  • πŸš€ Fully Automated CI/CD Deployment

πŸ› οΈ Tech Stack

Python PyTorch HuggingFace TRL PEFT BitsAndBytes FastAPI Docker Redis Qdrant GitHub Actions


Made with πŸ”₯ by Vivek Tyagi β€” AI Engineer

If you found this useful, drop a ⭐ β€” it means a lot!

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages