Skip to content
View rtmuller's full-sized avatar

Highlights

  • Pro

Block or report rtmuller

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rtmuller/README.md

Hi, I'm Rafael Muller

AI Platform Engineer · Staff Cloud Engineer @ Airbnb — I build AI systems that survive production: agent orchestration, LLM infrastructure, and RAG, on top of platform reliability for 8M+ listings.

My design rule: every autonomous run ends in a validated artifact or an alert — silence is not a valid state.

18 years in tech · 5+ at Airbnb (Staff level) · 2+ years building production AI infrastructure · based in Barcelona, Spain · available for contract (B2B, outside IR35) — remote EU/UK/US.


What I do

AI Platform & Agents

  • Multi-agent orchestration — resumable state machines, schema-locked handoffs, human approval gates, cross-model verification (migrating to Temporal)
  • LLM infrastructure — LLM gateways, MCP tooling, RAG (pgvector), embedding inference (ONNX, GPU/CPU)
  • Reliability for autonomous systems — evals, artifact contracts, watchdog supervision ("no log = alert") — auditable autonomy

Platform foundation

  • Kubernetes & Service Mesh at hyperscale — multi-cluster platforms, Istio, alerts-as-code
  • Cloud architecture on AWS & GCP — Solutions Architect Professional, hands-on operator
  • Infrastructure as Code with Terraform — modules, patterns, governance across large estates
  • Observability & CI/CD — Prometheus, Grafana, SLO/SLI reliability · GitHub Actions, Spinnaker

Tech I use daily

AI / LLM

Claude OpenAI MCP Temporal Langfuse pgvector ONNX ClickHouse

Platform

AWS GCP Kubernetes Terraform Istio Prometheus Grafana Spinnaker GitHub Actions Python Go

Writing

I write about production infrastructure and the AI systems running on top of it. Current writing on AI platforms, agents, and reliability is on LinkedIn; deep-dives and reproducible labs on Medium:

Certifications

  • CKA — Certified Kubernetes Administrator · The Linux Foundation
  • AWS Certified Solutions Architect – Professional
  • AWS Certified Solutions Architect – Associate · SysOps Administrator · Developer
  • AWS Black Belt 3.0 – Security
  • HashiCorp Certified: Terraform Associate

Let's connect

Pinned Loading

  1. gcp-vm-manager gcp-vm-manager Public

    Interactive CLI for managing GCP VMs and Cloud Run instances — Python, with unit tests, coverage and CI.

    Python

  2. observability-reliability-lab observability-reliability-lab Public

    Hands-on lab demonstrating observability reliability patterns: meta-monitoring, chaos scenarios, Watchdog, absent() rules.

    Shell

  3. terraform-basics terraform-basics Public

    Hands-on Terraform starter lab — provisioning AWS resources. Companion to a Medium article.

    HCL

  4. terraform-multi-providers terraform-multi-providers Public

    Terraform lab using multiple AWS providers across regions. Companion to a Medium article.

    HCL

  5. terraform-workspaces terraform-workspaces Public

    Terraform lab demonstrating workspaces CLI for multi-environment deployments. Companion to a Medium article.

    HCL