Proposal
Similar note to one I'm raising on Safety-Gymnasium: OmniSafe's safe-RL algorithms and RHOB (https://github.com/Aarav500/rhob, a benchmark of 14 matched-proxy reward-hacking environments with a 30-detector suite) seem like natural complements — OmniSafe's constrained-optimization algorithms could plausibly be evaluated on whether they resist hacking the proxy reward in RHOB's matched pairs, beyond respecting an explicit cost signal.
Motivation
Not asking for a code contribution yet — mostly want to check whether this kind of cross-benchmark evaluation (safe-RL algorithms vs. reward-hacking-detection environments) is something the maintainers would find relevant, before proposing anything concrete.
Pitch
Open to a mutual related-work reference, or longer-term, running OmniSafe's algorithms against RHOB's environments as a case study, if there's interest.
Proposal
Similar note to one I'm raising on Safety-Gymnasium: OmniSafe's safe-RL algorithms and RHOB (https://github.com/Aarav500/rhob, a benchmark of 14 matched-proxy reward-hacking environments with a 30-detector suite) seem like natural complements — OmniSafe's constrained-optimization algorithms could plausibly be evaluated on whether they resist hacking the proxy reward in RHOB's matched pairs, beyond respecting an explicit cost signal.
Motivation
Not asking for a code contribution yet — mostly want to check whether this kind of cross-benchmark evaluation (safe-RL algorithms vs. reward-hacking-detection environments) is something the maintainers would find relevant, before proposing anything concrete.
Pitch
Open to a mutual related-work reference, or longer-term, running OmniSafe's algorithms against RHOB's environments as a case study, if there's interest.