NexusSupport AI is an end-to-end AI engineering project that brings together classical machine learning, transformer-based deep learning, Retrieval-Augmented Generation (RAG), large language models, agentic decision-making, API serving, monitoring and containerization.
The goal is simple:
Build a support system that can understand a request, find trusted information, produce a grounded response and know when it should stop and ask for human help.
Unlike a simple chatbot, NexusSupport AI is designed as a complete AI application pipeline.
Customer Request
β
βΌ
ββββββββββββββββββββββββ
β Text Preprocessing β
ββββββββββββ¬ββββββββββββ
β
βΌ
ββββββββββββββββββββββββ
β Ticket Classificationβ
β TF-IDF + Logistic Regβ
ββββββββββββ¬ββββββββββββ
β
βΌ
Confidence Check
/ \
Low Good
β β
βΌ βΌ
Escalate RAG Retrieval
β
βΌ
Evidence Check
/ \
Weak Strong
β β
βΌ βΌ
Escalate LLM
β
βΌ
Grounded Answer
β
βΌ
Metrics
A real AI product needs more than a model.
It needs a way to:
- understand incoming data,
- select an appropriate prediction model,
- retrieve reliable information,
- use an LLM safely,
- make decisions based on intermediate results,
- expose the system as a service,
- monitor system behavior,
- and package the application for deployment.
NexusSupport AI combines these responsibilities in one modular project.
It is especially useful as a portfolio project because it demonstrates the full path from data β model β GenAI β API β monitoring β deployment.
Imagine a customer sends:
"I was charged twice for my Pro plan. Can I get a refund?"
NexusSupport AI can:
| Stage | What happens |
|---|---|
| π§Ή Preprocess | Cleans and prepares the message |
| π§ Classify | Predicts the support category |
| π― Evaluate confidence | Checks whether the prediction is reliable |
| π Retrieve | Searches the FAQ knowledge base |
| π Ground | Supplies relevant evidence to the LLM |
| βοΈ Generate | Produces a natural-language response |
| π‘οΈ Escalate | Avoids guessing when confidence/evidence is weak |
| π Monitor | Records request and latency metrics |
The system currently supports four demonstration ticket categories:
billing
technical
account
general
The baseline classifier uses:
Raw Ticket
β
Text Cleaning
β
TF-IDF
β
Logistic Regression
β
Category + Confidence
A strong ML engineering workflow does not begin with the most complicated model.
The classical model provides a fast and interpretable reference point. The project can then compare it against a transformer model.
Evaluation includes:
- Accuracy
- Precision
- Recall
- F1-score
- Confusion matrix
- Class distribution
The trained pipeline is saved to:
models/ticket_classifier.joblib
The project also fine-tunes DistilBERT for the same classification problem.
Support Ticket
β
Tokenizer
β
DistilBERT
β
Classification Head
β
Predicted Category
This creates a useful comparison:
| Approach | Main idea |
|---|---|
| TF-IDF + Logistic Regression | Fast classical NLP baseline |
| DistilBERT | Context-aware transformer model |
The purpose is not simply to use a transformer because it is larger.
The purpose is to evaluate whether its additional complexity provides a meaningful improvement.
The RAG system gives the LLM access to the project's knowledge base.
FAQ Documents
β
Chunking
β
Sentence Embeddings
β
FAISS Index
β
Semantic Retrieval
β
Relevant Context
β
LLM
β
Grounded Answer
Knowledge-base documents are stored under:
data/knowledge_base/
βββ account_faq.txt
βββ billing_faq.txt
βββ general_faq.txt
βββ technical_faq.txt
The current pipeline uses:
- Sentence Transformers
all-MiniLM-L6-v2- FAISS
- Anthropic Claude
Instead of asking the LLM to answer from memory alone:
Question β LLM β Answer
NexusSupport AI uses:
Question
β
Search trusted documents
β
Relevant evidence
β
LLM
β
Grounded answer
This also means the knowledge base can be updated without retraining the LLM.
The agent is the orchestration layer.
It connects classification, retrieval, decision-making and generation.
ββββββββββββββββββββ
β Incoming Request β
ββββββββββ¬ββββββββββ
βΌ
ββββββββββββββββββββ
β Classify β
ββββββββββ¬ββββββββββ
βΌ
ββββββββββββββββββββ
β Confidence OK? β
βββββββββ¬βββββββββββ
No / \ Yes
/ \
βΌ βΌ
Escalate Retrieve
β
βΌ
βββββββββββββββββ
β Evidence OK? β
βββββββββ¬ββββββββ
No / \ Yes
/ \
βΌ βΌ
Escalate LLM
β
βΌ
Answer
The agent uses explicit decision points rather than blindly generating a response.
This is an important difference between a simple chatbot and an agentic workflow.
βββββββββββββββββββββββ
β USER β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β FastAPI β
β app.py β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β SupportAgent β
β agent.py β
ββββββββββββ¬βββββββββββ
β
βββββββββββββββββββββββΌββββββββββββββββββββββ
β β β
βΌ βΌ βΌ
βββββββββββββββββ βββββββββββββββββ βββββββββββββββββ
β Classical ML β β RAG β β LLM β
β TF-IDF + LR β β Embeddings β β Claude β
βββββββββββββββββ β + FAISS β βββββββββββββββββ
β βββββββββββββββββ β
ββββββββββββββββββββββββ¬ββββββββββββββββββββββ
βΌ
Final Response
β
βΌ
βββββββββββββββββββββ
β Streamlit Monitor β
βββββββββββββββββββββ
ββββββββββββββββββββββββββββ
β DistilBERT Comparison β
β deep_learning.py β
ββββββββββββββββββββββββββββ
Docker β AWS / GCP
nexussupport-ai/
β
βββ data/
β βββ generate_sample_data.py
β βββ support_tickets.csv
β βββ knowledge_base/
β βββ account_faq.txt
β βββ billing_faq.txt
β βββ general_faq.txt
β βββ technical_faq.txt
β
βββ src/
β βββ preprocessing.py
β βββ classical_ml.py
β βββ deep_learning.py
β βββ rag_pipeline.py
β βββ agent.py
β βββ app.py
β
βββ dashboard/
β βββ dashboard.py
β
βββ deploy/
β βββ README.md
β
βββ models/
β βββ ticket_classifier.joblib
β
βββ assets/
β βββ nexussupport-ai-logo.png
β
βββ Dockerfile
βββ requirements.txt
βββ README.md
| Layer | Technology |
|---|---|
| Language | π Python |
| Data | Pandas Β· NumPy |
| Classical ML | Scikit-learn |
| NLP | TF-IDF |
| Deep Learning | PyTorch |
| Transformer | DistilBERT |
| Embeddings | Sentence Transformers |
| Vector Search | FAISS |
| Generative AI | Anthropic Claude |
| Agent | Custom Agent Workflow |
| API | FastAPI + Uvicorn |
| Dashboard | Streamlit |
| Visualization | Plotly |
| Model Storage | Joblib |
| Container | Docker |
| Cloud | AWS / GCP |
Recommended environment:
- Python 3.11
- Git
- Internet connection for model downloads
- Anthropic API key for LLM-based features
- Docker (optional)
git clone https://github.com/DewmikaSenarathna/NexusSupport-AI
cd nexussupport-aipython -m venv venv
.\venv\Scripts\Activate.ps1python -m venv venv
venv\Scripts\activatepython3 -m venv venv
source venv/bin/activatepython -m pip install --upgrade pip
pip install -r requirements.txt$env:GOOGLE_API_KEY="YOUR_API_KEY"export GOOGLE_API_KEY="YOUR_API_KEY"For the cleanest learning experience, run the components in this order.
python data/generate_sample_data.pyThis creates the synthetic support-ticket dataset and the FAQ knowledge base.
python src/preprocessing.pyThis checks:
- text cleaning,
- data splitting,
- class distribution,
- knowledge-base chunking.
python src/classical_ml.pyThis trains:
TF-IDF + Logistic Regression
and saves:
models/ticket_classifier.joblib
python src/rag_pipeline.pyThis:
- loads FAQ documents,
- creates chunks,
- creates embeddings,
- searches with FAISS,
- sends relevant context to the LLM,
- generates an answer.
The first run may download the embedding model.
python src/agent.pyThis runs the complete:
Classify
β
Confidence Check
β
Retrieve
β
Evidence Check
β
Generate / Escalate
workflow.
uvicorn src.app:app --reload --port 8000Open:
http://localhost:8000/docs
FastAPI provides interactive documentation for the available endpoints.
Keep FastAPI running and open another terminal.
streamlit run dashboard/dashboard.pyThe dashboard normally opens at:
http://localhost:8501
When the main system is working:
python src/deep_learning.pyThis fine-tunes DistilBERT and evaluates it against the classical ML approach.
A GPU is recommended for faster training.
| Method | Endpoint | Purpose |
|---|---|---|
GET |
/health |
Check API status |
POST |
/classify |
Classify a support ticket |
POST |
/ask |
Run RAG + LLM answering |
POST |
/agent |
Run the complete agent workflow |
GET |
/metrics |
View API metrics |
{
"text": "I was charged twice for my subscription"
}Possible response:
{
"category": "billing",
"confidence": 0.94,
"all_scores": {
"account": 0.01,
"billing": 0.94,
"general": 0.02,
"technical": 0.03
}
}The exact values depend on the trained model.
The Streamlit dashboard provides two main views.
βββββββββββββββββββββββββββββββββββββββ
β MODEL PERFORMANCE β
βββββββββββββββββββββββββββββββββββββββ€
β β
β Accuracy β
β Confusion Matrix β
β Class Distribution β
β β
βββββββββββββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββ
β LIVE API METRICS β
βββββββββββββββββββββββββββββββββββββββ€
β Total Requests β
β Average Latency β
β Requests by Endpoint β
β Latency by Endpoint β
βββββββββββββββββββββββββββββββββββββββ
These metrics are currently stored in application memory and are intended for demonstration and learning.
Build:
docker build -t nexussupport-ai .Run:
docker run -p 8000:8000 \
-e GOOGLE_API_KEY=YOUR_API_KEY \
nexussupport-aiThen open:
http://localhost:8000/docs
The project also contains deployment guidance under:
deploy/README.md
One of the useful experiments in this project is comparing two different approaches to the same classification task.
SUPPORT TICKET
β
ββββββββββ΄βββββββββ
β β
βΌ βΌ
Classical ML Deep Learning
β β
TF-IDF + LR DistilBERT
β β
ββββββββββ¬βββββββββ
βΌ
Compare Results
Compare:
- Accuracy
- Precision
- Recall
- F1-score
- Training time
- Inference behavior
- Model complexity
This helps demonstrate model selection based on evidence, rather than choosing a model simply because it is more advanced.
A language model by itself follows:
Question β LLM β Answer
NexusSupport AI adds a knowledge layer:
Question
β
Semantic Search
β
Trusted Context
β
LLM
β
Grounded Answer
This is useful when the answer depends on private, changing or domain-specific documents.
A fixed chatbot normally follows one path.
An agent can make decisions.
"Do I understand this?"
β
ββββββββ΄βββββββ
β β
No Yes
β β
βΌ βΌ
Escalate Retrieve
β
"Is evidence good?"
β
ββββββ΄βββββ
β β
No Yes
β β
βΌ βΌ
Escalate LLM
β
βΌ
Answer
This makes the workflow more controlled and explainable.
NexusSupport AI uses confidence and retrieval thresholds to reduce unsupported answers.
If:
classification confidence < threshold
the system can escalate.
If:
retrieval relevance < threshold
the system can also escalate.
The current values are demonstration settings and should be tuned and validated using real evaluation data before production use.
Each major responsibility has its own module.
Preprocessing
ML
Deep Learning
RAG
Agent
API
Dashboard
Deployment
A simple model is established before comparing it with a transformer.
The LLM receives retrieved context rather than relying only on its pretrained knowledge.
Low confidence can lead to human escalation instead of an unsupported answer.
The API is designed to load trained models once and reuse them for requests.
The project records agent traces and basic API metrics.
NexusSupport AI is a portfolio and learning-oriented AI engineering prototype.
- Support-ticket data is synthetic.
- FAQ documents are demonstration documents.
- API metrics are stored in memory.
- Authentication is not implemented.
- Long-term monitoring storage is not implemented.
- Automated model-drift detection is not implemented.
- Agent thresholds are fixed demonstration values.
- Production-grade security controls are not included.
These limitations provide clear paths for future development.
- Replace synthetic tickets with a real, properly licensed dataset
- Add more ticket categories
- Add multilingual support
- Add cross-validation
- Add hyperparameter optimization
- Add confidence calibration
- Add persistent vector storage
- Add metadata filtering
- Improve document chunking
- Add retrieval evaluation
- Add citation validation
- Build automated document ingestion
- Add ticket summarization
- Add human approval workflow
- Add conversation memory
- Add additional tools
- Improve decision policies
- Add experiment tracking
- Add model versioning
- Add model drift detection
- Store metrics in a database
- Add automated evaluation
- Add CI/CD
- Add API authentication
- Add rate limiting
- Use a secret manager
- Add stronger request validation
- Add production logging and audit trails
