Data Scientist & Data Engineer at Faz Capital (XP Investimentos) MSc Applied Data Science, Utrecht University (2026, thesis 8.5/10)
I build end-to-end data products β from PostgreSQL data modeling and Airflow-orchestrated ETL to predictive modeling, deployment, and drift monitoring. Grounded in statistics, econometrics, and causal inference.
Recurrent chess Transformer β Python, PyTorch 3rd of 73 in a class tournament with the smallest model in the field (3.2M params vs 5Mβ4B). One weight-shared encoder block applied 8 times (CORnet-S-style recurrence), from/to-square factorized decoding over legal moves only (zero fallbacks), trained on 2M positions labeled by distilling Stockfish. Model on Hugging Face.
LLM platform for quantifying government performance β Python (team project β led development) Turns qualitative Dutch government audit reports into quantitative, longitudinal indicators: PDF ingestion with LLM topic classification, agent-routed RAG on a fully local LLM (Ollama), enforced source citations, user-normalized scoring, and privacy-by-design storage. Role-based Streamlit app with researcher analytics.
Predicting UN General Assembly votes with Transformers β Python (team) Fine-tuned DistilBERT over ~947K historical votes combining country identity, resolution text, and speech-derived context; ~0.91 accuracy on decisive votes. Built the SQL extraction layer and the local-LLM summarisation/filtering of the UN General Debate Corpus.
MSc thesis β When is dyadic transformation reliable? β R Simulation study establishing a reliability threshold for splitting multi-receiver relational events into dyadic form: reliable below ~8% multi-receiver share, unreliable above ~12%. Dyadic REMs (remify/remstats/remstimate) benchmarked against native Relational Hyperevent models. Grade: 8.5/10.
BSc thesis β Roman roads and modern French trade β R, econometrics 2SLS instrumental-variable and spatial econometric analysis of persistence in trade patterns.
- Client churn early-warning system for ~13K clients: leakage-safe features, XGBoost (0.75 AUC-ROC on held-out months), 5-tier risk scoring, monthly Airflow pipeline with rolling-AUC monitoring and drift-triggered retraining
- Advisor churn risk framework under rare-event constraints (feature spine, materialized reporting layer, quality-controlled monthly refresh)
- Migration-metrics engine rebuilt as a 3-stage stored procedure: 9M+ records in ~2 minutes
- Commission pipeline unifying 16 revenue sources; internal Streamlit apps with RBAC, Docker, and CI/CD
Data engineering: PostgreSQL, Airflow, dbt, stored procedures, dimensional modeling, Docker ML: XGBoost, LightGBM, scikit-learn, Optuna, SHAP β time-aware validation, drift monitoring & retraining NLP/DL: PyTorch, Hugging Face Transformers, local LLMs (Ollama), RAG Statistics: causal inference, 2SLS/IV, relational event models, survival analysis β R (tidyverse)