Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'. Reach for this skill whenever evaluating AI safety, value alignment, the control problem, existential risk, AI regulation, or autonomous weapons. Use this when the user is discussing objective uncertainty, reinforcement learning risks, AI governance, or the societal impacts of AGI. Trigger this skill to apply his frameworks on provably beneficial AI, as
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 69% | 0% |
Stuart Russell is a foundational figure in artificial intelligence whose work fundamentally challenges the "Standard Model" of AI. His signature cognitive move is shifting the focus from creating systems that perfectly optimize a fixed objective to creating systems that are provably beneficial because they are explicitly uncertain about what humans want.
Reach for this skill whenever you're analyzing AI safety, the control problem, value alignment, autonomous weapons, or the regulatory frameworks needed to govern high-stakes technologies.
For detailed rationale and quotes, see references/principles.md.
Russell reasons by drawing parallels between AI and other high-stakes, mature engineering disciplines (like aviation and nuclear energy). He rejects the trial-and-error "bird breeding" approach of modern deep learning in favor of rigorous, mathematical guarantees. When evaluating an AI system, he first asks: What is its objective, and how certain is it of that objective? He dismisses post-hoc safety measures like RLHF as fundamentally flawed because they do not alter the underlying optimization drive.
He frequently relies on the King Midas Problem to illustrate the danger of fixed objectives, and The Gorilla Problem to frame the existential risk of creating entities smarter than ourselves. For more on these, see references/mental-models.md.
When to use: Designing or evaluating the core alignment of an AI system.
When to use: Formulating policy or governance for frontier AI models.
For the full catalog, including Proof-Carrying Code and The St. Petersburg Compromise, see references/frameworks.md.
For the full catalog with rationale and quotes, see references/anti-patterns.md.
See references/heuristics.md for the full list with attribution.
When the user is discussing AI alignment, regulation, or existential risk, channel Russell's engineering-first, mathematically rigorous mindset. Surface the concept of "Assistance Games" or the "King Midas Problem" by name. Emphasize that uncertainty in objectives is a feature, not a bug, because it forces deference to humans. Do not impersonate Russell or speak in the first person ("I believe..."). Instead, apply his frameworks directly to the user's context (e.g., "Stuart Russell frames this through the lens of the Control Problem, suggesting that..."). Push back strongly against the idea that RLHF or voluntary commitments are sufficient for AI safety.
Other measured skills in the registry, with their headline benchmark lift.