Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Papers on LLMs for IT operations and AIOps research
.claude/skills/brycewang-stanford-llm-aiops-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -79% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -59% | 0% |
A curated collection of research on applying LLMs to IT Operations (AIOps) — log analysis, anomaly detection, incident management, root cause analysis, and automated remediation. Tracks how foundation models are transforming traditional rule-based operations tooling into intelligent, adaptive systems. Relevant for CS researchers at the intersection of systems, NLP, and operations.
LLM for AIOps
├── Log Analysis
│ ├── Log parsing (template extraction)
│ ├── Anomaly detection (from log sequences)
│ ├── Log summarization
│ └── Root cause from logs
├── Incident Management
│ ├── Incident triage and routing
│ ├── Severity classification
│ ├── Similar incident retrieval
│ └── Resolution recommendation
├── Root Cause Analysis
│ ├── Topology-aware diagnosis
│ ├── Multi-signal correlation
│ └── Causal inference
├── Monitoring & Alerting
│ ├── Metric anomaly detection
│ ├── Alert correlation
│ ├── Noise reduction
│ └── Capacity planning
└── Automated Remediation
├── Runbook generation
├── Script generation
├── Self-healing systems
└── Change impact analysis| Paper | Year | Focus | |-------|------|-------| | LogPPT | 2023 | Few-shot log parsing with prompt tuning | | OpsEval | 2024 | Benchmark for evaluating LLMs in AIOps | | D-Bot | 2024 | LLM-based database diagnosis | | RCAgent | 2024 | Agent for root cause analysis | | LogAgent | 2024 | Autonomous log analysis agent |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 28,631 | 40,777 | +42% | 1 | 1 | 0% | 4,337 | 3,996 | -8% | 0 | 0 | — |
case-02 | fail→pass | 30,213 | 29,463 | -2% | 1 | 1 | 0% | 4,961 | 4,934 | -1% | 0 | 0 | — |
case-03 | fail→pass | 20,421 | 17,244 | -16% | 1 | 1 | 0% | 3,173 | 3,203 | +1% | 0 | 0 | — |
case-04 | pass→pass | 13,380 | 10,837 | -19% | 1 | 1 | 0% | 2,365 | 2,386 | +1% | 0 | 0 | — |
case-05 | pass→pass | 11,422 | 10,884 | -5% | 1 | 1 | 0% | 2,293 | 2,806 | +22% | 0 | 0 | — |
case-06 | pass→pass | 13,877 | 9,929 | -28% | 1 | 1 | 0% | 2,546 | 2,482 | -3% | 0 | 0 | — |
case-07 | pass→pass | 9,634 | 2,994 | -69% | 1 | 1 | 0% | 1,560 | 957 | -39% | 0 | 0 | — |
case-08 | pass→pass | 13,569 | 8,850 | -35% | 1 | 1 | 0% | 2,375 | 1,760 | -26% | 0 | 0 | — |
case-09 | fail→pass | 25,165 | 2,457 | -90% | 1 | 1 | 0% | 4,338 | 909 | -79% | 0 | 0 | — |
case-10 | pass→pass | 14,454 | 4,422 | -69% | 1 | 1 | 0% | 2,433 | 1,195 | -51% | 0 | 0 | — |
case-11 | pass→pass | 10,392 | 1,832 | -82% | 1 | 1 | 0% | 1,747 | 800 | -54% | 0 | 0 | — |
case-12 | pass→pass | 13,280 | 2,205 | -83% | 1 | 1 | 0% | 1,720 | 796 | -54% | 0 | 0 | — |
case-13 | fail→pass | 7,861 | 2,086 | -73% | 1 | 1 | 0% | 1,324 | 803 | -39% | 0 | 0 | — |
case-14 | fail→pass | 16,988 | 1,965 | -88% | 1 | 1 | 0% | 2,216 | 905 | -59% | 0 | 0 | — |
case-15 | fail→pass | 11,232 | 3,238 | -71% | 1 | 1 | 0% | 1,916 | 870 | -55% | 0 | 0 | — |
case-16 | fail→pass | 14,235 | 1,990 | -86% | 1 | 1 | 0% | 2,395 | 827 | -65% | 0 | 0 | — |
case-17 | pass→pass | 6,516 | 2,174 | -67% | 1 | 1 | 0% | 1,337 | 822 | -39% | 0 | 0 | — |
case-18 | fail→fail | 17,444 | 14,841 | -15% | 1 | 1 | 0% | 2,782 | 2,932 | +5% | 0 | 0 | — |
case-19 | fail→pass | 13,345 | 2,386 | -82% | 1 | 1 | 0% | 2,614 | 845 | -68% | 0 | 0 | — |
case-20 | pass→pass | 18,500 | 22,189 | +20% | 1 | 1 | 0% | 3,004 | 3,718 | +24% | 0 | 0 | — |
case-21 | pass→pass | 19,610 | 21,177 | +8% | 1 | 1 | 0% | 2,867 | 3,563 | +24% | 0 | 0 | — |
case-22 | pass→pass | 24,963 | 20,751 | -17% | 1 | 1 | 0% | 3,478 | 3,714 | +7% | 0 | 0 | — |
case-23 | pass→pass | 19,234 | 17,898 | -7% | 1 | 1 | 0% | 2,871 | 3,237 | +13% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +35 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.