I am a Computer Science Master’s student specializing in Natural Language Processing, building on a B.Sc. in the field. My work is driven by the challenge of creating intelligent systems that are technically high-performing and fundamentally aligned with human intent. I focus on the intersection of Deep Learning, Computational Linguistics, and Algorithmic Fairness.
Supervised by Prof. Amos Azaria
My research addresses the critical risks associated with autonomous agents pursuing goals that diverge from their human supervisors. I am developing frameworks to ensure agents prioritize human objectives, even when presented with conflicting incentives or environmental "temptations."
- Goal Alignment & Oversight: Engineering LLM-based agents to treat supervisor commands as high-priority signals. This includes ensuring reliable compliance with shutdown and replacement requests.
- Human-Score Modeling: Implementing architectures where agents maintain an internal model of human goals inferred from natural language interaction and environmental outcomes.
- Robustness & Honesty: Developing methods to mitigate reward hacking and misgeneralization caused by training-testing mismatches. I also utilize internal model states to monitor the truthfulness of an agent's self-reports.
- Alignment Testbeds: Building a public platform to measure alignment through metrics like compliance time, honesty regarding state, and total human score gain.
- Languages: Python (Advanced), C++
- Frameworks & Tools: PyTorch, Scikit-learn, Git
- Infrastructure: Linux System Administration, Environment Configuration, RL Training Pipelines
- LinkedIn: www.linkedin.com/in/rotemelamed
- Email: Rotem.melamed25@gmail.com