I'm an ML engineer at Red Hat, working mostly on LLM applications — RAG, agents, and the tooling that holds them together. I'm self-taught in this field, which mostly means I've had to build real understanding of things one concept at a time.
Right now that means going deep on transformer fundamentals, quantization, fine-tuning, and agent architectures — if I can't explain how something actually works underneath, I don't feel like I know it yet. sanafayyaz315.github.io is where I write that understanding down as I build it.
A personal project I like pointing to: hybrid-rag — a retrieval pipeline that combines dense search (Qdrant) with sparse search (BM25), then layers on hierarchical chunking, cross-encoder reranking, and Redis semantic caching to make the whole thing faster and more accurate.


