AI Engineer specialized in AI Security & GenAI Governance, with hands-on experience protecting production AI ecosystems in regulated, large-scale environments (financial services, telecom, enterprise ERP). Specialist in LLM red teaming (OWASP Top 10 for LLMs, Microsoft PyRIT, Promptfoo), continuous evaluation of agents in production, and multi-agent architecture — combining senior software engineering (Python, Cloud, CI/CD) with rigorous security controls. Track record of 32 GenAI use cases evaluated and secured before reaching production.
Currently focused on transforming legacy automation into autonomous AI agents, while pursuing a postgraduate degree in Applied AI Engineering (UNIPDS) and an MBA in Machine Learning Engineering (FIAP).
Autonomous Agents & Process Automation
- Led the conversion of legacy RPA flows into autonomous AI agents combining LLMs, RAG, and computer vision, cutting execution time per process by 40% and achieving 95%+ task success rate in production enterprise environments.
- Implemented security-by-design guardrails and validation from conception, blocking 8 vulnerabilities pre-production through CI/CD pipelines.
GenAI Evaluation & Observability
- Structured continuous evaluation for 8+ GenAI use cases in production using golden sets, regression testing, and groundedness/hallucination/safety metrics (DeepEval, G-Eval, ROUGE/BLEU/METEOR), reducing quality incidents by 35%.
- Monitored production agents on task success rate, tool-calling accuracy, latency, and cost per execution, cutting drift/degradation detection time by 50% via thresholds and alerting.
- Combined automated evaluation (LLM-as-a-judge) with human review in a continuous loop, sustaining 90% adherence to defined quality criteria.
AI Security & Governance
- Ran red teaming covering OWASP Top 10 for LLMs risks (prompt injection, jailbreak, data leakage, tool abuse) across 12 use cases, blocking 20+ vulnerabilities before production.
- Automated security testing pipelines with Microsoft PyRIT, Promptfoo, and DeepTeam, cutting evaluation time per use case by 60%, producing evidence, reports, and playbooks aligned with NIST AI RMF.
- Grew from sub-lead to squad lead within 8 months, coordinating a team of 3 and aligning technical risk with cross-functional stakeholders.
Full-Stack AI Solutions
- Built a financial document entity-extraction platform from scratch (Python, Django, LangChain), processing 12,000 documents/month at 97% average accuracy, eliminating manual data entry.
- Developed a multi-agent marketing automation platform (briefing, content, review, validation), reducing campaign creation cycle time by 45%.
- Delivered full-stack improvements on high-traffic platforms serving 150,000+ users, optimizing SQL queries and reducing operational rework by 30%.
- Internal PDF RAG Knowledge Chat — Full-stack RAG pipeline (LangChain, ChromaDB, FastAPI, Docker) with cited answers and confidence scoring in under 120ms per query.
- Risk Scoring ML Pipeline — Risk-scoring pipeline with feature engineering and supervised models achieving AUC ≈ 0.91 (scikit-learn, PyTorch, structured data validation).
- Financial Document Entity Extraction — LLM-based structured entity extraction from financial documents (Django, LangChain, REST APIs).
- Postgraduate, Applied AI Engineering — UNIPDS (May 2026 – May 2027, in progress) LLMs, RAG, MCP, autonomous agents, fine-tuning (LoRA/PEFT), AI security & governance.
- MBA, Machine Learning Engineering — FIAP (Jan 2026 – Oct 2026, in progress) MLflow, DVC, Docker, Kubernetes, CI/CD, model monitoring, data drift, LLMOps.
- B.Tech, Systems Analysis and Development — Estácio de Sá (Jul 2022 – Jan 2025, completed)
📫 Email: marcospaulomaio2607@gmail.com 💼 LinkedIn: linkedin.com/in/marcos-maio

