Researchers at Microsoft Research Asia have formalized a training paradigm called Harnessed Agentic RL and released Agent Lightning v1.0 as an open-source implementation. The key idea is to let the same agent harness used in deployment participate directly in reinforcement learning, avoiding the need to reimplement the agent loop inside the training framework.
Key features and goals
- Lightweight and transparent: the Agent Lightning v1.0 codebase is roughly 3,500 lines of code, designed for readability and extensibility.
- Training on real harnesses: the system places an LLM proxy between the agent and the model so the deployed harness can remain unchanged while the training framework observes its model calls.
- Native Kubernetes support: agents run as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes, or local infrastructure, removing dependence on paid sandbox services.
- Complete coding-agent example: an end-to-end pipeline increased Qwen3.5-9B Pass@1 on SWE-bench Verified from 41.8% to 56.4% (an absolute gain of 14.6 percentage points) using about 6,000 open-source training samples.
Why this departs from traditional agentic RL
Traditional agentic RL typically assumes the training framework controls the interaction loop with the environment, so a rollout maps to one continuous token trajectory and systems often reimplement the agent loop in the trainer. Many modern harnesses — e.g., mini-SWE-agent, OpenHands, OpenCode, Claude Code, Codex — contain their own context management, tool protocols, and execution logic, which makes reimplementation expensive and potentially behavior-changing.
Agent Lightning addresses this by inserting an LLM proxy: the agent continues to run unchanged, and pointing its model endpoint at the Agent Lightning proxy allows the training framework to observe and record model calls. Agent Lightning v1.0 formalizes this approach as Harnessed Agentic RL: the deployed harness participates directly in RL training.
Four training challenges with real harnesses
Because the harness manages the interaction loop, the trainer only observes LLM request-response pairs and a single rollout can split into a variable number of training samples. This introduces four challenges:
- Retokenization and sample merging: harnesses keep context as text, but RL needs the token IDs sampled during rollout; re-tokenizing can change token boundaries and prevent merging adjacent calls into one sample.
- Advantage calculation: splitting a rollout into multiple samples means computing advantages at sample level can overcount rollouts that generate more samples, distorting rollout-level statistics.
- Loss normalization: averaging loss by sample count gives more weight to rollouts that produce more samples, so normalization must avoid distortion from harness behavior.
- Training backend scheduling: sample counts and lengths are known only after harness completion, while GPU counts and parallel configurations are typically fixed — the backend must map variable workloads to fixed resources.
Architecture: a complete agent RL control plane in about 3,500 lines of code
Agent Lightning v1.0 implements simplicity-first system design with three core components:
- API Gateway: an OpenAI-compatible LLM proxy that stores rollouts, models, and events; it links each model call from the harness to its rollout and records prompts, responses, and log probabilities needed for training.
- Rollout Controller: starts and manages agent execution as local processes or standard Kubernetes jobs, keeping agent execution separate from the trainer.
- Customized Trainer: built on verl, it creates rollouts, waits for completion, collects samples, and assembles final training samples via a sample adapter.
For many existing harnesses, directing the model endpoint to the Agent Lightning proxy is enough to connect the harness to RL training quickly.
Collocated Async RL: improved utilization with fewer GPUs
Rollout durations vary significantly across agents. Synchronous RL wastes GPU time waiting for the slowest agent, while fully asynchronous RL increases utilization but requires separate GPU pools for rollout and training. Agent Lightning v1.0 introduces Collocated Async RL, which lets rollouts and model updates share the same GPUs.
When enough rollouts are collected, the API Gateway pauses new requests, waits for in-flight requests to finish, performs the model update, and then resumes rollouts. This transition is transparent to the external harness. In experiments, Collocated Async RL delivered roughly a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL.
Running agents on Kubernetes: scalable and cost-effective rollouts
Collecting large numbers of rollouts requires substantial CPU, memory, and compute. Other frameworks often use commercial sandbox services (e.g., Modal Sandbox or E2B), which can become costly at scale. Agent Lightning v1.0 runs agents as standard Kubernetes jobs, reusing self-managed clusters, cloud Kubernetes, or local infra. This improves resource efficiency, lowers rollout cost, and keeps the pipeline open source and reproducible.
Experimental pipeline and results: 6,000 samples, 14.6-point gain
The researchers built a full pipeline using SWE-smith, mini-SWE-agent, and Qwen3.5-9B, covering data cleaning, environment construction, reward-hacking safeguards, and RL training. The training set contained about 6,000 samples and did not require large-scale compute. RL training alone increased Qwen3.5-9B Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute improvement of 14.6 percentage points.
The experiments also validated the earlier technical analysis: rollout-level advantage computation combined with rollout-level normalization outperformed sample-level handling, producing higher validation reward and more stable policy entropy during training.
Conclusion
Agent Lightning v1.0 and the Harnessed Agentic RL paradigm prioritize training on the actual deployed harness, delivering a compact, Kubernetes-native, and reproducible agent RL pipeline. The demonstrated gains in the coding-agent example suggest this approach can be an effective and cost-efficient alternative to traditional agentic RL frameworks.
Further materials
The technical report and the Agent Lightning v1.0 GitHub project are available from Microsoft Research.



