Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or auditing AAMAS experiments - self-play and population-based training, opponent selection, equilibrium and regret metrics, game-theoretic simulations, ablations, seeds, hyperparameters, compute, and claim-to-evidence fit - with emphasis on experiments that probe the interaction rather than chase a single-agent leaderboard.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -47% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -7% | 0% |
Use this before submission when the empirical or simulation story is not yet locked. At AAMAS the experiment exists to test the interaction claim, not to top a benchmark.
deviation test.
population sets, or classical strategies as the claim requires.
real or applied studies that show practical multiagent behavior.
confidence intervals, or paired tests.
hyperparameter ranges, chosen settings, seeds, hardware, software versions, and runtime.
not just cosmetic variants.
against only one fixed opponent, or a cooperation claim that hides a reward-shaping constant.
the method did not train against, and populations that vary in size or composition.
than five extra environments where nothing strategic is tested.
named solution concept, exploitability, social welfare, or regret - not just episodic return.
| Interaction claim | Matching experiment | Reject pattern avoided | |---|---|---| | Converges to equilibrium | Convergence/exploitability curve under simultaneous adaptation | "Equilibrium asserted, never measured" | | Mechanism is truthful | Strategic-deviation test: an agent tries to misreport | "Truthfulness proved, never stress-tested" | | Beats other agents | Round-robin vs held-out opponents and a population | "Self-play only" | | Emergent cooperation | Sweep over reward/opponent settings with variance | "One seed, one setting, one story" |
Suppose the paper claims a learned protocol raises cooperation in a repeated public-goods game. The matching plan: sweep group size and defector fraction for cooperation curves, add held-out opponents that never appeared in training, and inject a free-rider agent to measure whether it profits - every panel tied to a numbered claim or definition.
standard errors, confidence intervals, or quantiles, and how many opponents were averaged.
text[Experiment readiness] strong / adequate / weak [Claim -> evidence map] <claim: game / self-play / population / deviation test> [Missing interaction evidence] <opponents / deviation test / seeds / metric> [Reproducibility gaps] <hyperparameters / compute / env / seeds> [Decision-critical next run] <one experiment or simulation>
Other measured skills in the registry, with their headline benchmark lift.