Skip to content

Related benchmark: matched-proxy reward-hacking environments (RHOB) #390

Description

@Aarav500

Proposal

Similar note to one I'm raising on Safety-Gymnasium: OmniSafe's safe-RL algorithms and RHOB (https://github.com/Aarav500/rhob, a benchmark of 14 matched-proxy reward-hacking environments with a 30-detector suite) seem like natural complements — OmniSafe's constrained-optimization algorithms could plausibly be evaluated on whether they resist hacking the proxy reward in RHOB's matched pairs, beyond respecting an explicit cost signal.

Motivation

Not asking for a code contribution yet — mostly want to check whether this kind of cross-benchmark evaluation (safe-RL algorithms vs. reward-hacking-detection environments) is something the maintainers would find relevant, before proposing anything concrete.

Pitch

Open to a mutual related-work reference, or longer-term, running OmniSafe's algorithms against RHOB's environments as a case study, if there's interest.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions