What is new

Microsoft Research Asia has released Agent Lightning v1.0, an open-source framework for training AI agents with reinforcement learning (RL). The headline idea is a training approach the team calls Harnessed Agentic RL: instead of rebuilding an agent's logic inside a separate RL training system, the exact agent harness used in production takes part directly in training. The framework itself is deliberately small, about 3,500 lines of code, and it runs agents as standard Kubernetes jobs rather than relying on paid commercial sandbox services such as Modal Sandbox or E2B.

In a worked example, the team trained a coding agent built on the open Qwen3.5-9B model using about 6,000 training samples. Reinforcement learning alone raised the model's Pass@1 score on SWE-bench Verified, a benchmark that measures whether an AI agent can correctly fix real software bugs, from 41.8% to 56.4%, a gain of 14.6 percentage points.

Background: why agent training is hard

An AI agent is not just a language model. It is usually a full system: a model, plus a harness that manages tools, memory, context and the back-and-forth with its environment. Reinforcement learning is a training method where a system improves by trying actions and receiving rewards or penalties based on the outcome. In the agent world, RL typically works by having the model propose an action, observing what happens, feeding that observation back into the context, and repeating, a loop often called ReAct-style interaction.

Most existing RL training frameworks, such as verl, AReaL and slime, assume they own this loop completely, treating a whole run as one continuous sequence of tokens. That assumption breaks down for modern coding and general-purpose agents, such as mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex, each of which manages context, tool calls and dependencies in its own way. To train these agents under the old assumption, developers had to reimplement the agent's loop inside the RL framework, an expensive step that also risks training a version of the agent that behaves differently from the one actually shipped.

How it works

Agent Lightning avoids rebuilding the agent by inserting itself between the harness and the model as an LLM proxy. Developers simply point the endpoint the harness already uses to call the model toward Agent Lightning instead, leaving the rest of the harness untouched. The training system then observes and records the model calls flowing through that proxy, which is enough to drive RL training without ever touching the harness's internal code.

Because the harness, not the training framework, now manages the loop with the environment, a single run can be split unpredictably into many separate model calls. The researchers identify four technical problems this creates: text passed back through a tokenizer can shift token boundaries, making it hard to merge related calls into one training sample; computing reward baselines at the sample level double-counts rollouts that happen to produce more samples; averaging the training loss by sample count can overweight those same rollouts; and GPU scheduling has to cope with workloads whose size is only known after the agent has already run.

Agent Lightning v1.0 addresses these issues with a system built from three parts: an API Gateway that logs every model call and links it to its run, a Rollout Controller that launches and manages agent runs either as local processes or Kubernetes jobs, and a Customized Trainer, built on the existing verl framework, that reassembles the logged calls into coherent training samples.

The framework also introduces what it calls Collocated Async RL, a way of running agent rollouts and model updates on the same GPUs rather than on separate pools. Once enough rollouts are collected, the gateway briefly pauses new requests, lets in-flight ones finish, applies the model update, and resumes, a transition the researchers say is invisible to the agent harness itself. According to the team, this produced roughly a 2x end-to-end speedup over fully synchronous RL while using fewer GPUs than a fully asynchronous setup.

Why it matters

For practitioners, the practical appeal is twofold. First, because the production harness is trained directly rather than a reimplemented stand-in, there is less risk of a gap between the agent that gets evaluated during training and the one that gets deployed. Second, running agents as ordinary Kubernetes jobs rather than through commercial sandbox providers removes a cost barrier that scales with the number of rollouts, which matters for teams that want to run large volumes of agent training without a growing cloud bill.

The SWE-bench Verified result is also notable for its data efficiency. A 14.6 percentage point improvement from around 6,000 training samples, built on an existing open dataset (SWE-smith) and an existing open-source agent (mini-SWE-agent), suggests that meaningful gains do not necessarily require enormous training sets, at least for this coding task.

Limits and open questions

The published results compare two ways of computing training signal, rollout-level versus sample-level advantage and loss normalization, and report that the rollout-level approach yields higher validation reward and more stable policy entropy during training. The materials do not say how Agent Lightning performs on agent types beyond coding, such as web browsing or multi-tool business agents, nor do they report training time, hardware cost in absolute terms, or how the 14.6 point gain compares against other RL methods tested under the same conditions. The team also includes safeguards against reward hacking, where an agent finds a shortcut that scores well without actually solving the task, but the specifics of those safeguards are not detailed in the material reviewed here.

How to try it

Agent Lightning v1.0 has been open-sourced by Microsoft Research Asia, along with the coding agent training example built on SWE-smith, mini-SWE-agent and Qwen3.5-9B. The core integration step described by the researchers is to point an existing agent harness's model endpoint at the Agent Lightning proxy, after which the framework can begin logging calls and feeding them into RL training through its Kubernetes-based rollout controller.

Sources