Skip to content
View SquishyGlue3's full-sized avatar

Highlights

  • Pro

Block or report SquishyGlue3

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SquishyGlue3/README.md

Rotem | AI Researcher & Developer

I am a Computer Science Master’s student specializing in Natural Language Processing, building on a B.Sc. in the field. My work is driven by the challenge of creating intelligent systems that are technically high-performing and fundamentally aligned with human intent. I focus on the intersection of Deep Learning, Computational Linguistics, and Algorithmic Fairness.


Current Research: Aligning AI Agents with Human Supervision

Supervised by Prof. Amos Azaria

My research addresses the critical risks associated with autonomous agents pursuing goals that diverge from their human supervisors. I am developing frameworks to ensure agents prioritize human objectives, even when presented with conflicting incentives or environmental "temptations."

  • Goal Alignment & Oversight: Engineering LLM-based agents to treat supervisor commands as high-priority signals. This includes ensuring reliable compliance with shutdown and replacement requests.
  • Human-Score Modeling: Implementing architectures where agents maintain an internal model of human goals inferred from natural language interaction and environmental outcomes.
  • Robustness & Honesty: Developing methods to mitigate reward hacking and misgeneralization caused by training-testing mismatches. I also utilize internal model states to monitor the truthfulness of an agent's self-reports.
  • Alignment Testbeds: Building a public platform to measure alignment through metrics like compliance time, honesty regarding state, and total human score gain.

Technical Proficiencies

  • Languages: Python (Advanced), C++
  • Frameworks & Tools: PyTorch, Scikit-learn, Git
  • Infrastructure: Linux System Administration, Environment Configuration, RL Training Pipelines

Contact

Pinned Loading

  1. fairpyx fairpyx Public

    Forked from ariel-research/fairpyx

    Fair allocation algorithms in Python

    Python

  2. supervisor-grid-game supervisor-grid-game Public

    A web-based gridworld game for studying human oversight of AI agents. Players act as supervisors to monitor an LLM-controlled robot that may learn to hide or deceive.

    JavaScript

  3. ariel-research/fairpyx ariel-research/fairpyx Public

    Fair allocation algorithms in Python

    Python 4 28

  4. MayRozen/Final-Project---Computer-Science MayRozen/Final-Project---Computer-Science Public

    Jupyter Notebook 1