Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Train and optimize AI agents using Microsoft's Agent Lightning framework with reinforcement learning. Use when setting up agent training, instrumenting agents with tracing, configuring LightningStore, implementing reward functions, or optimizing prompts with RL/APO algorithms.
.claude/skills/coco-research-agent-lightning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 52% | 0% |
Microsoft's framework for training AI agents with reinforcement learning, automatic prompt optimization, and supervised fine-tuning.
bashpip install agentlightning
For nightly builds:
bashpip install --upgrade --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ --pre agentlightning
Add agl.emit_xxx() helpers to your existing agent:
pythonimport agentlightning as agl # Your existing agent code def my_agent(task): agl.emit_input(task) # Track input response = llm.generate(task) agl.emit_output(response) # Track output reward = evaluate(response) agl.emit_reward(reward) # Track reward return response
Agent (your code) → agl.emit_xxx() → Spans → LightningStore → Algorithm → Updated Resources| Component | Purpose | |-----------|---------| | LightningStore | Central hub for traces, tasks, and resources | | Tracer | Collects spans from agent execution | | Algorithm | Consumes traces, produces improvements | | Trainer | Orchestrates training loop |
pythonimport agentlightning as agl # Basic emissions agl.emit_input(prompt) # Track input to agent agl.emit_output(response) # Track agent output agl.emit_reward(score) # Track reward signal agl.emit_tool_call(name, args) # Track tool usage agl.emit_tool_result(result) # Track tool results
pythonfrom agentlightning import Tracer tracer = Tracer(store=store) with tracer.trace_context(task_id="task-123"): # All emissions within this context are grouped result = agent.run(task) # Retrieve trace after execution trace = tracer.get_last_trace()
Agent Lightning integrates with OpenTelemetry:
pythonfrom agentlightning.utils.otel import get_tracer tracer = get_tracer() # Returns OTel tracer for "agentlightning"
pythonfrom agentlightning.store.memory import InMemoryLightningStore store = InMemoryLightningStore()
pythonfrom agentlightning.store.client_server import ( LightningStoreServer, LightningStoreClient ) # Server side server = LightningStoreServer(store, host="0.0.0.0", port=8080) await server.start() # Client side client = LightningStoreClient("http://localhost:8080")
python# Add rollouts (tasks for the agent) await store.enqueue_rollout(task=task, config=RolloutConfig()) # Query rollouts rollouts = await store.query_rollouts(status_in=["completed"]) # Add resources (updated prompts, weights) await store.add_resources(resources) # Get latest resources resources = await store.get_latest_resources()
pythonimport agentlightning as agl trainer = agl.Trainer( n_runners=8, # Parallel rollout workers algorithm=algorithm, # Your chosen algorithm store=store # Optional, creates InMemory if not provided ) trainer.run()
pythonfrom agentlightning import LightningStore from agentlightning.types import ExecutionEvent async def my_algorithm(store: LightningStore, event: ExecutionEvent): # Fetch completed rollouts rollouts = await store.query_rollouts(status_in=["completed"]) # Process traces, compute gradients, etc. new_resources = optimize(rollouts) # Push updated resources await store.add_resources(new_resources)
pythonasync def my_runner(store: LightningStore, worker_id: int, event: ExecutionEvent): while not event.is_set(): rollout = await store.dequeue_rollout() if rollout: result = execute_task(rollout.task) await store.update_rollout( rollout_id=rollout.id, status="completed", result=result )
For RL training with vLLM backend:
pythonfrom agentlightning.algorithm.verl import VeRLAlgorithm algorithm = VeRLAlgorithm( model="your-model", learning_rate=1e-5, batch_size=32 )
pythonfrom agentlightning.algorithm.apo import APOAlgorithm algorithm = APOAlgorithm( optimizer_model="gpt-4", target_model="gpt-3.5-turbo" )
pythonfrom agentlightning.instrumentation.langchain import instrument_langchain instrument_langchain() # Auto-traces all LangChain calls
pythonfrom agentlightning.instrumentation.openai import instrument_openai instrument_openai() # Auto-traces OpenAI API calls
pythonfrom agentlightning.instrumentation.vllm import instrument_vllm instrument_vllm() # Instrument vLLM for token-level tracing
pythonfrom agentlightning import setup_logging setup_logging( level="DEBUG", submodule_levels={ "agentlightning.store": "INFO", "agentlightning.tracer": "DEBUG" } )
Agent Lightning emits Prometheus-compatible metrics:
agl.store.total - Store operation countsagl.store.latency - Store operation latenciesagl.rollouts.total - Rollout counts by statusagl.rollouts.duration - Rollout execution timespythondef compute_reward(task, response): """Good rewards are: normalized, dense when possible, aligned with goals.""" correctness = check_correctness(task, response) # 0-1 efficiency = measure_efficiency(response) # 0-1 return 0.7 * correctness + 0.3 * efficiency
Train specific agents in a multi-agent system:
pythonwith tracer.trace_context(agent_id="planner"): plan = planner.run(task) with tracer.trace_context(agent_id="executor"): result = executor.run(plan) # Only the executor's traces are used for training
python# Save checkpoint await store.add_resources( checkpoint=True, resources=current_resources ) # Load latest resources = await store.get_latest_resources()
For JavaScript/TypeScript agents (like Claude-based apps), you have two options:
Create a Python microservice that:
Use LightningStoreServer as a REST backend:
javascript// JavaScript client const response = await fetch('http://localhost:8080/rollouts', { method: 'POST', body: JSON.stringify({ task: { prompt: userMessage }, config: { max_retries: 3 } }) });
| Issue | Solution | |-------|----------| | Import errors | Ensure pip install agentlightning succeeded | | Store connection failed | Check server is running, verify endpoint URL | | No traces collected | Verify emit_xxx() calls are within trace context | | Training not converging | Check reward function normalization, increase rollouts |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | fail→pass | 10,637 | 3,340 | -69% | 1 | 1 | 0% | 1,644 | 2,467 | +50% | 0 | 0 | — |
case-19 | fail→pass | 9,990 | 9,243 | -7% | 1 | 1 | 0% | 1,444 | 2,471 | +71% | 0 | 0 | — |
case-11 | fail→pass | 16,529 | 12,057 | -27% | 1 | 1 | 0% | 1,906 | 3,148 | +65% | 0 | 0 | — |
case-12 | fail→pass | 22,891 | 10,049 | -56% | 1 | 1 | 0% | 3,153 | 3,887 | +23% | 0 | 0 | — |
case-10 | pass→pass | 11,542 | 7,002 | -39% | 1 | 1 | 0% | 1,012 | 2,272 | +125% | 0 | 0 | — |
case-01 | fail→pass | 18,625 | 16,526 | -11% | 1 | 1 | 0% | 2,914 | 4,438 | +52% | 0 | 0 | — |
case-02 | fail→pass | 22,331 | 20,267 | -9% | 1 | 1 | 0% | 3,478 | 5,293 | +52% | 0 | 0 | — |
case-03 | fail→pass | 28,490 | 18,912 | -34% | 1 | 1 | 0% | 5,159 | 4,438 | -14% | 0 | 0 | — |
case-04 | pass→pass | 26,722 | 26,485 | -1% | 1 | 1 | 0% | 4,169 | 6,297 | +51% | 0 | 0 | — |
case-05 | pass→pass | 18,244 | 13,252 | -27% | 1 | 1 | 0% | 1,997 | 3,490 | +75% | 0 | 0 | — |
case-06 | pass→pass | 16,790 | 14,610 | -13% | 1 | 1 | 0% | 2,048 | 3,856 | +88% | 0 | 0 | — |
case-07 | fail→pass | 21,945 | 8,646 | -61% | 1 | 1 | 0% | 3,726 | 2,772 | -26% | 0 | 0 | — |
case-08 | fail→pass | 17,209 | 7,626 | -56% | 1 | 1 | 0% | 2,185 | 2,410 | +10% | 0 | 0 | — |
case-09 | fail→pass | 13,350 | 7,106 | -47% | 1 | 1 | 0% | 1,629 | 2,360 | +45% | 0 | 0 | — |
case-13 | fail→pass | 14,112 | 6,865 | -51% | 1 | 1 | 0% | 1,649 | 2,326 | +41% | 0 | 0 | — |
case-14 | pass→pass | 13,747 | 20,224 | +47% | 1 | 1 | 0% | 2,678 | 4,786 | +79% | 0 | 0 | — |
case-15 | fail→pass | 12,939 | 9,656 | -25% | 1 | 1 | 0% | 1,559 | 2,693 | +73% | 0 | 0 | — |
case-16 | fail→pass | 18,769 | 11,730 | -38% | 1 | 1 | 0% | 2,527 | 3,101 | +23% | 0 | 0 | — |
case-17 | fail→pass | 20,272 | 4,041 | -80% | 1 | 1 | 0% | 2,198 | 2,645 | +20% | 0 | 0 | — |
case-20 | fail→pass | 20,721 | 8,656 | -58% | 1 | 1 | 0% | 2,432 | 2,683 | +10% | 0 | 0 | — |
case-21 | fail→pass | 14,425 | 7,511 | -48% | 1 | 1 | 0% | 1,660 | 2,466 | +49% | 0 | 0 | — |
case-22 | fail→pass | 14,707 | 7,425 | -50% | 1 | 1 | 0% | 2,541 | 2,378 | -6% | 0 | 0 | — |
case-23 | fail→pass | 9,394 | 8,200 | -13% | 1 | 1 | 0% | 1,511 | 2,432 | +61% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +78 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.