Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run an Anthropic Claude Managed Agent — a cloud agent harness (container + filesystem + tools), the cloud counterpart of the local wasm-agent runtime
.claude/skills/ruvnet-managed-agent/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -10% | 0% |
ruflo-agent has two agent runtimes behind one mental model:
| Runtime | Tools | Use it when | |---|---|---| | WASM (local, rvagent) | wasm_agent_* / wasm_gallery_* | fast, free, ephemeral, offline, untrusted code in a sandbox | | Managed (Anthropic cloud) | managed_agent_* (this skill) | long-running / async work (minutes–hours), a real cloud container with pre-installed packages + network, persistent filesystem + transcript across turns |
This skill drives the managed runtime — Anthropic's Claude Managed Agents (beta). The model: Agent (model + system + tools + MCP servers + skills) → Environment (container template) → Session (running instance) → Events (turns / tool-use / status, persisted server-side). See docs/adr/0001-wasm-contract.md and project ADR-115.
ANTHROPIC_API_KEY (or CLAUDE_API_KEY) in the environment, with Claude Managed Agents beta access.managed_agent_* tool returns a structured "use wasm_agent_create for a local no-key runtime" error — fall back to the WASM skill.mcp__plugin_ruflo-core_ruflo__managed_agent_create{ model?, system?, name?, networking?, packages?, initScript?, mcpServers?, skills? } → { sessionId, agentId, environmentId, status }. Provisions Agent + Environment + Session. Save the three ids.
mcpServers: [{type:"url", url, name, authorization_token?}] — the cloud agent must be able to reach the URL. A local ruflo mcp start is not reachable from Anthropic's cloud; deploy/tunnel an HTTP ruflo MCP server first if you want the cloud agent to have ruflo's tools.packages: {pip?:[], npm?:[], apt?:[], cargo?:[], gem?:[], go?:[]} — installed in the container.mcp__plugin_ruflo-core_ruflo__managed_agent_prompt{ sessionId, message, maxWaitMs? } → sends a user turn, polls the event log until the session goes idle (default 180s, capped 600s) → { finished, status, stopReason, assistantText, toolUses[], eventCount }. For very long tasks, raise maxWaitMs or follow up with managed_agent_events.
mcp__plugin_ruflo-core_ruflo__managed_agent_status { sessionId } (idle/running/error) · mcp__plugin_ruflo-core_ruflo__managed_agent_events { sessionId, raw? } (full transcript: user turns, agent thinking, tool_use, tool_result, status — the cloud counterpart of wasm_agent_files).mcp__plugin_ruflo-core_ruflo__managed_agent_list { limit? } — every session on the org (so you can see which are still running / billing).mcp__plugin_ruflo-core_ruflo__managed_agent_terminate { sessionId, environmentId? } — always do this when done: a cloud session keeps billing container time + tokens until deleted. Pass environmentId to also delete the environment ruflo created.cost-tracking namespace.managed_agent_list then managed_agent_terminate anything stale.managed-agents-2026-04-01); multiagent / define-outcomes on the agent config are research preview.managed_agent_create { "model": "claude-haiku-4-5-20251001", "system": "Terse. Do exactly what is asked.", "name": "scratch" }
→ { sessionId: "sesn_…", agentId: "agent_…", environmentId: "env_…", status: "idle" }
managed_agent_prompt { "sessionId": "sesn_…", "message": "echo hello > /tmp/x && cat /tmp/x — then stop." , "maxWaitMs": 60000 }
→ { finished: true, status: "idle", stopReason: "end_turn", assistantText: "Done.", toolUses: [{name:"bash", input:{command:"echo hello > /tmp/x && cat /tmp/x"}}] }
managed_agent_terminate { "sessionId": "sesn_…", "environmentId": "env_…" }
→ { sessionDeleted: true, environmentDeleted: true }| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 10,310 | 6,222 | -40% | 1 | 1 | 0% | 1,956 | 1,727 | -12% | 0 | 0 | — |
case-02 | fail→fail | 10,461 | 12,730 | +22% | 1 | 1 | 0% | 1,643 | 1,354 | -18% | 0 | 0 | — |
case-03 | fail→fail | 12,038 | 6,691 | -44% | 1 | 1 | 0% | 2,373 | 1,668 | -30% | 0 | 0 | — |
case-04 | fail→pass | 8,265 | 2,564 | -69% | 1 | 1 | 0% | 1,379 | 1,685 | +22% | 0 | 0 | — |
case-05 | fail→pass | 7,575 | 1,401 | -82% | 1 | 1 | 0% | 1,455 | 1,454 | -0% | 0 | 0 | — |
case-06 | pass→pass | 9,520 | 5,327 | -44% | 1 | 1 | 0% | 1,904 | 2,234 | +17% | 0 | 0 | — |
case-07 | pass→pass | 12,071 | 3,506 | -71% | 1 | 1 | 0% | 2,128 | 1,837 | -14% | 0 | 0 | — |
case-08 | fail→pass | 11,408 | 4,888 | -57% | 1 | 1 | 0% | 2,068 | 2,185 | +6% | 0 | 0 | — |
case-09 | pass→pass | 11,008 | 5,966 | -46% | 1 | 1 | 0% | 1,810 | 1,649 | -9% | 0 | 0 | — |
case-10 | fail→pass | 8,140 | 2,351 | -71% | 1 | 1 | 0% | 1,397 | 1,656 | +19% | 0 | 0 | — |
case-11 | fail→pass | 10,415 | 3,088 | -70% | 1 | 1 | 0% | 1,809 | 1,627 | -10% | 0 | 0 | — |
case-12 | fail→pass | 4,710 | 2,126 | -55% | 1 | 1 | 0% | 834 | 1,604 | +92% | 0 | 0 | — |
case-13 | pass→pass | 12,666 | 5,675 | -55% | 1 | 1 | 0% | 2,266 | 2,351 | +4% | 0 | 0 | — |
case-14 | fail→pass | 11,219 | 1,907 | -83% | 1 | 1 | 0% | 2,211 | 1,491 | -33% | 0 | 0 | — |
case-15 | pass→pass | 7,349 | 2,244 | -69% | 1 | 1 | 0% | 1,199 | 1,647 | +37% | 0 | 0 | — |
case-16 | fail→pass | 11,626 | 2,924 | -75% | 1 | 1 | 0% | 1,915 | 1,705 | -11% | 0 | 0 | — |
case-17 | pass→pass | 14,346 | 2,962 | -79% | 1 | 1 | 0% | 2,366 | 1,671 | -29% | 0 | 0 | — |
case-18 | pass→pass | 10,089 | 1,299 | -87% | 1 | 1 | 0% | 1,615 | 1,363 | -16% | 0 | 0 | — |
case-19 | pass→pass | 4,252 | 1,819 | -57% | 1 | 1 | 0% | 730 | 1,481 | +103% | 0 | 0 | — |
case-20 | pass→pass | 9,471 | 6,688 | -29% | 1 | 1 | 0% | 2,042 | 2,604 | +28% | 0 | 0 | — |
case-21 | pass→pass | 16,820 | 9,762 | -42% | 1 | 1 | 0% | 3,231 | 3,234 | +0% | 0 | 0 | — |
case-22 | pass→pass | 6,290 | 4,481 | -29% | 1 | 1 | 0% | 1,172 | 2,035 | +74% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.