diff --git a/README.md b/README.md index 6dba02a7..210b161f 100644 --- a/README.md +++ b/README.md @@ -6,9 +6,14 @@ Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems +

+ NeurIPS 2026
+ MALGAI @ ICLR 2026 +

+

-Project Page +Project Website arXiv Paper @@ -25,22 +30,31 @@ Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems License - Repo stars

-The project-page source and visual previews are available in [website/](./website/README.md). +**News:** ๐ŸŽ‰ Dr. MAS has been accepted to **NeurIPS 2026!** + +Dr. MAS is designed for **end-to-end post-training** of **Multi-Agent LLM Systems** via **Reinforcement Learning (RL)**, enabling multiple LLM-based agents to collaborate on complex reasoning and decision-making tasks. -`Dr.MAS` is designed for **end-to-end post-training** of **Multi-Agent LLM Systems** via **Reinforcement Learning (RL)**, enabling multiple LLM-based agents to collaborate on complex reasoning and decision-making tasks. +

+ + Figure 1: Dr. MAS uses agent-wise advantage normalization for stable reinforcement learning and supports co-training multiple LLM actors. + +
+ Figure 1. Agent-wise normalization for stable multi-LLM RL. +

This framework features **flexible agent registry**, **customizable multi-agent orchestration**, **LLM sharing/non-sharing (e.g., heterogeneous LLMs)**, **per-agent configuration**, and **shared resource pooling**, making it well suited for training multi-agent LLM systems with RL.

- framework + + Figure 2: The Dr. MAS framework connects environment workers, multi-agent orchestration, agent-model assignment, shared GPU scheduling, and per-worker-group optimization. + +
+ Figure 2. The Dr. MAS framework for end-to-end multi-agent LLM RL.

# Feature Summary @@ -54,7 +68,7 @@ This framework features **flexible agent registry**, **customizable multi-agent | **Shared Resource Pooling** | โœ… Shared GPU pool across multiple LLM worker groups for efficient hardware utilization
โœ… Gradient updates applied independently for each worker group during optimization | | **Environments** | โœ… Math
โœ… Search | | **Model Support** | โœ… Qwen2.5
โœ… Qwen3
โœ… LLaMA3.2
and more | -| **RL Algorithms** | โœ… Dr.MAS
โœ… GRPO
๐Ÿงช GiGPO (experimental)
๐Ÿงช DAPO (experimental)
๐Ÿงช RLOO (experimental)
๐Ÿงช PPO (experimental)
and more | +| **RL Algorithms** | โœ… Dr. MAS
โœ… GRPO
โœ… GiGPO
โœ… DAPO
โœ… RLOO
โœ… PPO
and more | # Table of Contents diff --git a/docs/drmas/drmas_framework.png b/docs/drmas/drmas_framework.png index 0b1a7cce..78556d48 100644 Binary files a/docs/drmas/drmas_framework.png and b/docs/drmas/drmas_framework.png differ diff --git a/docs/drmas/drmas_overview.png b/docs/drmas/drmas_overview.png new file mode 100644 index 00000000..4cc8bce1 Binary files /dev/null and b/docs/drmas/drmas_overview.png differ diff --git a/website/index.html b/website/index.html index 3b4c0876..d38d0dc5 100644 --- a/website/index.html +++ b/website/index.html @@ -50,7 +50,10 @@
Multi-agent reinforcement learning

Dr. MASStable Reinforcement Learning
for Multi-Agent LLM Systems

-

NeurIPS 2026

+

+ NeurIPS 2026 + MALGAI @ ICLR 2026 +

Lang Feng Longtao Zheng diff --git a/website/src/styles.css b/website/src/styles.css index e26d8152..5eccbbbb 100644 --- a/website/src/styles.css +++ b/website/src/styles.css @@ -98,12 +98,13 @@ sup { font-size: .55em; } .hero::before { content: ''; position: absolute; z-index: -1; width: 900px; max-width: 100%; height: 550px; top: 0; left: 50%; transform: translateX(-50%); background: radial-gradient(ellipse at 50% 0, #ebdfef9c, transparent 70%); pointer-events: none; } .hero-content { text-align: center; } .hero-topline { display: flex; flex-wrap: wrap; align-items: center; justify-content: center; gap: 12px; } -.venue-badge { display: inline-flex; align-items: center; color: var(--accent-strong); border: 1px solid #cdb9d6; background: #f3edf6; padding: 5px 10px; font: 14px var(--mono); border-radius: 4px; letter-spacing: .035em; } +.venue-badge { display: inline-flex; align-items: center; color: var(--accent-strong); border: 1px solid #cdb9d6; background: #f3edf6; padding: 5px 10px; font: 14px var(--mono); border-radius: 4px; letter-spacing: .035em; white-space: nowrap; } +.venue-workshop { color: var(--muted); border-color: var(--line); background: transparent; } .hero-eyebrow { display: inline-flex; gap: 10px; align-items: center; color: #7d6686; font-size: 10px; letter-spacing: .11em; } .hero h1 { margin-top: 18px; } .hero-name { display: block; font-size: clamp(48px, 5vw, 64px); font-weight: 600; letter-spacing: -.055em; line-height: 1.1; color: var(--accent); } .hero-title { display: block; font-size: clamp(26px, 3.2vw, 42px); letter-spacing: -.04em; line-height: 1.22; margin-top: 12px; font-weight: 500; } -.hero-description { color: var(--muted); font-size: 15px; margin-top: 18px; } +.hero-description { display: flex; flex-wrap: wrap; align-items: center; justify-content: center; gap: 8px; color: var(--muted); font-size: 15px; margin-top: 18px; } .authors { display: flex; align-items: center; justify-content: center; flex-wrap: wrap; gap: 8px 23px; margin-top: 23px; font-size: 15px; font-weight: 500; } .authors a { text-decoration-color: var(--line); text-decoration-line: underline; text-underline-offset: 4px; } .authors a:hover { color: var(--accent); text-decoration-color: var(--accent); }