Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run the Level 2 dummy agent integration test suite and produce a detailed HTML report with per-test input → outcome analysis.
.claude/skills/aden-hive-integration-test-reporting-skill/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 114% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 154% | 0% |
Run the Level 2 dummy agent integration test suite and produce a detailed HTML report with per-test input → outcome analysis.
User wants to run integration tests and see results:
/test-reporting/test-reporting test_component_queen_live.py/test-reporting --allIf the user provides a specific test file or pattern, use it. Otherwise run the full suite.
bash# Full suite cd core && echo "1" | uv run python tests/dummy_agents/run_all.py --interactive 2>&1 # Specific file (requires manual provider setup) cd core && uv run python -c " import sys sys.path.insert(0, '.') from tests.dummy_agents.run_all import detect_available from tests.dummy_agents.conftest import set_llm_selection avail = detect_available() claude = [p for p in avail if 'Claude Code' in p['name']] if not claude: avail_names = [p['name'] for p in avail] raise RuntimeError(f'No Claude Code subscription. Available: {avail_names}') provider = claude[0] set_llm_selection( model=provider['model'], api_key=provider['api_key'], extra_headers=provider.get('extra_headers'), api_base=provider.get('api_base'), ) import pytest sys.exit(pytest.main([ 'tests/dummy_agents/TEST_FILE_HERE', '-v', '--override-ini=asyncio_mode=auto', '--no-header', '--tb=long', '--log-cli-level=WARNING', '--junitxml=/tmp/hive_test_results.xml', ])) "
After the test run completes, collect:
--junitxml output (if available)run_all.py output (the Unicode table)Write the report to /tmp/hive_integration_test_report.html.
The report MUST include these sections:
For EVERY test (not just failures), include a row with:
| Column | Description | |--------|-------------| | Component | Test file grouping (e.g., component_queen_live) | | Test Name | Function name (e.g., test_queen_starts_in_planning_without_worker) | | Status | PASS / FAIL / SKIP / ERROR with color badge | | Duration | Wall-clock seconds | | What | One-line description of what the test verifies | | How | How it works (setup → action → assertion) | | Why | Why this test matters (what bug/behavior it catches) | | Input | The input data or configuration (graph spec, initial prompt, phase, etc.) | | Expected Outcome | What the test asserts | | Actual Outcome | What actually happened (PASS: matches expected / FAIL: actual vs expected) | | Failure Detail | For failures only: full traceback + diagnosis |
These MUST be derived from the test function's docstring and code. Read each test file to extract:
Use these mappings for the component test files:
test_component_llm.py → "LLM Provider" — streaming, tool calling, tokens
test_component_tools.py → "Tool Registry + MCP" — connection, execution
test_component_event_loop.py → "EventLoopNode" — iteration, output, stall
test_component_edges.py → "Edge Evaluation" — conditional, priority
test_component_conversation.py → "Conversation Persistence" — storage, cursor
test_component_escalation.py → "Escalation Flow" — worker→queen signaling
test_component_continuous.py → "Continuous Mode" — conversation threading
test_component_queen.py → "Queen Phase (Unit)" — phase state, tools, events
test_component_queen_live.py → "Queen Phase (Live)" — real queen, real LLM
test_component_queen_state_machine.py → "Queen State Machine" — edge cases, races
test_component_worker_comms.py → "Worker Communication" — events, data flow
test_component_strict_outcomes.py → "Strict Outcomes" — exact path, output, qualityUse this structure:
html<!DOCTYPE html> <html lang="en"> <head> <meta charset="utf-8"> <title>Hive Integration Test Report — {timestamp}</title> <style> :root { --pass: #22c55e; --fail: #ef4444; --skip: #f59e0b; --bg: #0f172a; --surface: #1e293b; --text: #e2e8f0; --muted: #94a3b8; --border: #334155; } * { box-sizing: border-box; margin: 0; padding: 0; } body { font-family: 'SF Mono', 'Fira Code', monospace; background: var(--bg); color: var(--text); padding: 2rem; line-height: 1.6; } h1, h2, h3 { font-weight: 600; } h1 { font-size: 1.5rem; margin-bottom: 1rem; } h2 { font-size: 1.2rem; margin: 2rem 0 1rem; border-bottom: 1px solid var(--border); padding-bottom: 0.5rem; } .summary { display: grid; grid-template-columns: repeat(auto-fit, minmax(150px, 1fr)); gap: 1rem; margin-bottom: 2rem; } .card { background: var(--surface); padding: 1rem; border-radius: 8px; border: 1px solid var(--border); } .card .label { color: var(--muted); font-size: 0.75rem; text-transform: uppercase; } .card .value { font-size: 1.5rem; font-weight: 700; margin-top: 0.25rem; } .card .value.pass { color: var(--pass); } .card .value.fail { color: var(--fail); } table { width: 100%; border-collapse: collapse; font-size: 0.8rem; } th { background: var(--surface); position: sticky; top: 0; text-align: left; padding: 0.5rem; border-bottom: 2px solid var(--border); color: var(--muted); text-transform: uppercase; font-size: 0.7rem; } td { padding: 0.5rem; border-bottom: 1px solid var(--border); vertical-align: top; } tr:hover { background: rgba(255,255,255,0.03); } .badge { display: inline-block; padding: 2px 8px; border-radius: 4px; font-size: 0.7rem; font-weight: 700; } .badge.pass { background: rgba(34,197,94,0.2); color: var(--pass); } .badge.fail { background: rgba(239,68,68,0.2); color: var(--fail); } .badge.skip { background: rgba(245,158,11,0.2); color: var(--skip); } .detail { background: #1a1a2e; padding: 0.75rem; border-radius: 4px; margin-top: 0.5rem; font-size: 0.75rem; white-space: pre-wrap; overflow-x: auto; max-height: 200px; overflow-y: auto; } .component-header { background: var(--surface); padding: 0.75rem 0.5rem; font-weight: 600; font-size: 0.85rem; } .meta { color: var(--muted); font-size: 0.75rem; } </style> </head> <body> <h1>Hive Integration Test Report</h1> <p class="meta">Generated: {timestamp} | Provider: {provider} | Duration: {duration}s</p> <div class="summary"> <div class="card"><div class="label">Total</div><div class="value">{total}</div></div> <div class="card"><div class="label">Passed</div><div class="value pass">{passed}</div></div> <div class="card"><div class="label">Failed</div><div class="value fail">{failed}</div></div> <div class="card"><div class="label">Verdict</div><div class="value {verdict_class}">{verdict}</div></div> </div> <h2>Test Results</h2> <table> <thead> <tr> <th>Component</th> <th>Test</th> <th>Status</th> <th>Time</th> <th>What</th> <th>Input → Expected → Actual</th> </tr> </thead> <tbody> <!-- For each test: --> <tr> <td>{component}</td> <td>{test_name}</td> <td><span class="badge {status_class}">{status}</span></td> <td>{duration}s</td> <td>{what_description}</td> <td> <strong>Input:</strong> {input_description}<br> <strong>Expected:</strong> {expected_outcome}<br> <strong>Actual:</strong> {actual_outcome} <!-- If failed: --> <div class="detail">{failure_traceback}</div> </td> </tr> </tbody> </table> <h2>Failure Analysis</h2> <!-- Only if there are failures --> <p>For each failure, provide:</p> <ul> <li><strong>Root cause:</strong> Why it failed</li> <li><strong>Impact:</strong> What this means for the system</li> <li><strong>Suggested fix:</strong> How to address it</li> </ul> </body> </html>
/tmp/hive_integration_test_report.html Test Report: /tmp/hive_integration_test_report.html Result: 74/76 PASSED (2 failures) Failures:
--junitxml when running pytest to get structured results<div class="detail">| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | pass→pass | 8,002 | 4,758 | -41% | 1 | 1 | 0% | 1,611 | 3,829 | +138% | 0 | 0 | — |
case-01 | fail→fail | 5,928 | 4,356 | -27% | 1 | 1 | 0% | 468 | 3,045 | +551% | 0 | 0 | — |
case-02 | fail→fail | 7,604 | 8,835 | +16% | 1 | 1 | 0% | 380 | 3,372 | +787% | 0 | 0 | — |
case-03 | fail→fail | 12,730 | 5,814 | -54% | 1 | 1 | 0% | 2,653 | 3,058 | +15% | 0 | 0 | — |
case-04 | fail→fail | 14,207 | 6,313 | -56% | 1 | 1 | 0% | 2,645 | 4,220 | +60% | 0 | 0 | — |
case-05 | fail→pass | 13,211 | 6,690 | -49% | 1 | 1 | 0% | 2,538 | 4,024 | +59% | 0 | 0 | — |
case-06 | pass→fail | 4,521 | 4,062 | -10% | 1 | 1 | 0% | 940 | 2,950 | +214% | 0 | 0 | — |
case-07 | pass→fail | 7,721 | 17,984 | +133% | 1 | 1 | 0% | 1,322 | 5,734 | +334% | 0 | 0 | — |
case-09 | fail→pass | 9,137 | 3,290 | -64% | 1 | 1 | 0% | 1,603 | 3,435 | +114% | 0 | 0 | — |
case-10 | fail→pass | 11,387 | 1,988 | -83% | 1 | 1 | 0% | 1,839 | 3,096 | +68% | 0 | 0 | — |
case-11 | pass→pass | 10,749 | 3,350 | -69% | 1 | 1 | 0% | 1,819 | 3,364 | +85% | 0 | 0 | — |
case-12 | fail→pass | 17,028 | 8,006 | -53% | 1 | 1 | 0% | 2,829 | 4,175 | +48% | 0 | 0 | — |
case-13 | fail→pass | 7,493 | 2,230 | -70% | 1 | 1 | 0% | 1,232 | 3,133 | +154% | 0 | 0 | — |
case-14 | fail→pass | 11,515 | 1,326 | -88% | 1 | 1 | 0% | 1,971 | 2,972 | +51% | 0 | 0 | — |
case-15 | fail→pass | 6,654 | 1,730 | -74% | 1 | 1 | 0% | 1,149 | 3,110 | +171% | 0 | 0 | — |
case-16 | pass→pass | 15,345 | 5,937 | -61% | 1 | 1 | 0% | 3,032 | 3,934 | +30% | 0 | 0 | — |
case-17 | fail→pass | 7,077 | 3,152 | -55% | 1 | 1 | 0% | 1,234 | 3,450 | +180% | 0 | 0 | — |
case-18 | pass→pass | 6,737 | 2,718 | -60% | 1 | 1 | 0% | 1,049 | 3,032 | +189% | 0 | 0 | — |
case-19 | pass→pass | 13,735 | 7,902 | -42% | 1 | 1 | 0% | 2,161 | 4,150 | +92% | 0 | 0 | — |
case-20 | fail→pass | 13,684 | 4,438 | -68% | 1 | 1 | 0% | 2,673 | 3,551 | +33% | 0 | 0 | — |
case-21 | fail→fail | 7,753 | 1,535 | -80% | 1 | 1 | 0% | 1,243 | 3,078 | +148% | 0 | 0 | — |
case-22 | pass→fail | 4,678 | 2,310 | -51% | 1 | 1 | 0% | 765 | 3,175 | +315% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 18 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.