HuggingEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
-
Updated
Sep 9, 2026 - Python
HuggingEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
An OpenEnv benchmark testing the ability of AI agents to act as Site Reliability Engineers (SREs) by diagnosing and filtering raw production failure logs.
Reproducible OpenEnv-compatible reinforcement-learning environments for social-media integrity moderation.
A family of long-horizon software-engineering environments for OpenEnv, adapted from https://github.com/Proximal-Labs/frontier-swe
A rigorously formulated, non-stationary Partially Observable Markov Decision Process (POMDP) environment evaluating LLM crisis triage under real-world FEMA/ICS resource constraints. Features mathematical trajectory-aware reward shaping, NHPP disaster spawning, and 422→200 hallucination recovery.
OpenMedRL is an open-source reinforcement learning environment for benchmarking LLM-powered medical agents in emergency care. It simulates triage, dynamic patient progression, resource constraints, and uncertainty-aware clinical decision-making.
Enterprise-grade OpenEnv environment for AI-driven customer support triage. Features stochastic noise, multi-turn dialogue reasoning, and deterministic reward scoring for benchmarking AI agents.
Enterprise AIOps Omni-Environment: A production-grade OpenEnv sandbox for evaluating RL agents on CloudOps, FinOps, and Data Governance tasks.
Universal evaluation layer for standard RL environments. Measures what an agent learned - not just how much reward it accumulated.
🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environments (OpenEnv, verifiers-spec, Gymnasium).
we're addicted to solve some real issues ~ Team ComputeXor
Deterministic evaluation environment for AI code reviewers covering bugs, security (OWASP), and architecture via FastAPI + OpenEnv.
Gymnasium RL environment for AI-powered customer support triage — classify, prioritize, assign, and respond to emails under SLA pressure. Built for the OpenEnv spec.
AI-powered system for low-exposure route optimization using AQI, simulation, and intelligent decision-making
An agent must triage incoming support/compliance emails using respond, escalate, or archive while minimizing risk and maximizing completion quality.
An OpenEnv RL environment where an LLM agent plays the buyer and negotiates against an LLM-powered seller over real marketplace listings.
An OpenEnv-compliant reinforcement learning environment designed to train and evaluate AI agents on real-world SQL debugging, performance tuning, and schema design.
📧 Intelligent Agentic Workflow for Autonomous Enterprise Email Triage. Built with OpenEnv, featuring Chain-of-Thought reasoning and Self-Correcting agent logic for high-stakes corporate routing.
CyberRange is an advanced, self-improving simulated environment designed to train and benchmark autonomous security agents in complex enterprise incident response.
To associate your repository with the openenv topic, visit your repo's landing page and select "manage topics."