GitHub - NVIDIA-NeMo/ProRL-Agent-Server: Agentic RL on Any Harness at Scale

Polar is a RL rollout framework for real-world agent harnesses.

Harness as Environment. Bring your agent harnesses as RL-ready environments without code change.
Smart Rollout Pipeline. Save GPU hours with Polar's parallel rollout staging & runtime prewarm.
Rollout as a Service. Server mode by design -- scaling Async RL with any training frameworks.

Architecture Overview

The Rollout Server manages and dispatches client requests into distributed Gateway Nodes, which asynchronously prepare runtime, execute agents, build trajectories and evaluate them. Agent harnesses are listened by a proxy that sits between agnostic agent execution processes and local inference servers.

Installation

Install the Rollout Server (Polar):

uv venv
uv pip install -e .

Install the Inference Server (SGLang):

uv pip install --prerelease=allow sglang==0.5.10
bash scripts/patch/patch_sglang.sh

The patch applies necessary TITO and prompt token id emission on pinned sglang version. We'll remove this once upstream supports go through. vllm integration is on the way.

Polar is trainer agnostic. So choice of Trainer and Training Backend are highly flexible given Polar's server boundaries.

Currently, we provide a demo-purpose Slime integration in Slime bridge installation guide.

(Optional) For SWE-bench official evaluation harness:

uv pip install -e ".[swebench]"

(Optional) To enable polar dashboard UI, build the frontend once.

cd web && npm install && npm run build

Usage Guide

⭐ Choose your Agent Harness: pick a built-in harness, or use the generic shell harness with wrapped agents.
🚀 Trajectory Construction and Eval: See builder and evaluator guides for registered strategies.
🔧 Deployment Topology: configure the Polar service.
▶️ Request for Rollout: client side task submission via rollout API.

CLI Interface

A typical local run uses five commands. Each takes the same topology.yaml.

polar serve_rollout   -c topology.yaml                            # central orchestrator (port 8080)
polar serve_gateway   -c topology.yaml --node-id <node>           # one per gateway node (port 8100+)
polar dashboard       -c topology.yaml [--port 8090]              # observability & monitoring dashboard
polar submit          <task.json|yaml> -c topology.yaml           # submit a task and tail it
polar status          -c topology.yaml                            # one-shot health / topology check

Examples

Calculator: minimal smoke test.
Count Stars: minimal test for VLM.
SWE-bench Verified: benchmark-style evaluation on SWE-bench Verified tasks.
SWE-Gym Slime GRPO: training path that connects Polar rollouts to Slime.

This project is under active development. We are adding new examples for different tasks / models on diverse hardware setups. Contributions are welcome!

Roadmap

Our development goal for Polar is low-intrusion and neutral, finding the lowest common ancestor to cover and support diverse training and inference frameworks.

Initial release & tech report.
Slime bridge & RL example.
CUA (VLM / VLA) Support.
More built-in evaluators (eg. self distillation with textual feedback).
vLLM dual inference support.
More trainer bridges (NemoRL, VERL, etc.).

📖 Reference

Important

If you find it useful, please consider citing our work:

@article{zhang2026prorl,
  title={ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents},
  author={Zhang, Hao and Liu, Mingjie and Zhang, Shaokun and Han, Songyang and Hu, Jian and Jin, Zhenghui and Zhang, Yuchi and Diao, Shizhe and Lu, Ximing and Xu, Binfeng and others},
  journal={arXiv preprint arXiv:2603.18815},
  year={2026}
}

Name		Name	Last commit message	Last commit date
Latest commit History 4,402 Commits
assets		assets
examples		examples
scripts		scripts
src		src
tests		tests
web		web
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
pyproject.toml		pyproject.toml
uv.lock		uv.lock

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Architecture Overview

Installation

Install the Rollout Server (Polar):

Install the Inference Server (SGLang):

Polar is trainer agnostic. So choice of Trainer and Training Backend are highly flexible given Polar's server boundaries.

(Optional) For SWE-bench official evaluation harness:

(Optional) To enable polar dashboard UI, build the frontend once.

Usage Guide

CLI Interface

Examples

Roadmap

📖 Reference

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

Architecture Overview

Installation

Install the Rollout Server (Polar):

Install the Inference Server (SGLang):

Polar is trainer agnostic. So choice of Trainer and Training Backend are highly flexible given Polar's server boundaries.

(Optional) For SWE-bench official evaluation harness:

(Optional) To enable polar dashboard UI, build the frontend once.

Usage Guide

CLI Interface

Examples

Roadmap

📖 Reference

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages