Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Write session-based exploratory testing charters to find what scripted tests miss. Use when asked to plan exploratory testing, write a test charter, design a testing session, or do risk-based exploration of a feature. Produces focused charters — a mission, areas/risks to explore, tactics and oracles, and timeboxed sessions — so exploration is purposeful and accountable, not random clicking.
.claude/skills/mohitagw15856-exploratory-test-charter/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 161% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 50% | 0% |
Exploratory testing finds the bugs scripts don't — but only when it's chartered: a clear mission, a defined area, and a timebox, so it's purposeful and you can report what was covered. This skill writes session-based charters that point skilled testing at the riskiest areas, with the tactics and oracles to know when something is wrong.
Given "explore the new checkout flow", write the charters anyway — infer the risk areas, useful tactics, and oracles, labelling assumptions. Prioritise by risk. Never hand back a question instead of charters.
Ask for these only if they aren't already provided (else infer and label):
Risk overview — the few areas most worth exploring and why (new, complex, high-impact, historically buggy).
Charters — one per focused session (Session-Based Test Management style):
> Charter: Explore area] using tactics/data] to discover information about risk]. > - Areas / things to cover: the specific surfaces, flows, inputs, states. > - Test ideas & tactics: how to probe it — boundary values, interruptions, bad data, concurrency, navigation, roles/permissions, network conditions, etc. > - Oracles (how you'll know it's wrong): the spec, consistency, comparable products, user expectations, "would a user be annoyed?". > - Timebox: ~60–90 min (short/long), priority. > - Data / setup needed.
Provide 3–6 charters, prioritised by risk.
Reporting — what to capture per session: bugs found, areas covered vs. not, new risks/questions, and follow-up charters.
Session-Based Test Management (exploratory testing) — chartered, risk-prioritised, timeboxed sessions with explicit tactics and oracles.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 25,808 | 23,002 | -11% | 1 | 1 | 0% | 4,525 | 4,847 | +7% | 0 | 0 | — |
case-02 | fail→fail | 22,032 | 21,684 | -2% | 1 | 1 | 0% | 3,936 | 4,380 | +11% | 0 | 0 | — |
case-03 | fail→pass | 7,522 | 16,422 | +118% | 1 | 1 | 0% | 1,359 | 3,551 | +161% | 0 | 0 | — |
case-04 | fail→pass | 17,082 | 16,768 | -2% | 1 | 1 | 0% | 2,957 | 3,677 | +24% | 0 | 0 | — |
case-05 | fail→pass | 15,634 | 18,582 | +19% | 1 | 1 | 0% | 2,793 | 3,756 | +34% | 0 | 0 | — |
case-06 | fail→pass | 15,846 | 19,746 | +25% | 1 | 1 | 0% | 2,721 | 4,094 | +50% | 0 | 0 | — |
case-07 | fail→pass | 15,429 | 18,176 | +18% | 1 | 1 | 0% | 2,958 | 3,880 | +31% | 0 | 0 | — |
case-08 | pass→pass | 15,014 | 15,945 | +6% | 1 | 1 | 0% | 2,775 | 3,435 | +24% | 0 | 0 | — |
case-09 | fail→pass | 10,797 | 22,451 | +108% | 1 | 1 | 0% | 2,427 | 4,511 | +86% | 0 | 0 | — |
case-10 | pass→pass | 15,028 | 17,306 | +15% | 1 | 1 | 0% | 2,594 | 3,749 | +45% | 0 | 0 | — |
case-11 | pass→pass | 11,899 | 19,265 | +62% | 1 | 1 | 0% | 2,279 | 3,964 | +74% | 0 | 0 | — |
case-12 | fail→pass | 15,013 | 17,290 | +15% | 1 | 1 | 0% | 2,636 | 3,583 | +36% | 0 | 0 | — |
case-13 | pass→pass | 17,277 | 17,747 | +3% | 1 | 1 | 0% | 2,891 | 3,623 | +25% | 0 | 0 | — |
case-14 | pass→pass | 18,638 | 23,140 | +24% | 1 | 1 | 0% | 3,259 | 4,764 | +46% | 0 | 0 | — |
case-15 | pass→pass | 17,151 | 14,153 | -17% | 1 | 1 | 0% | 2,857 | 3,002 | +5% | 0 | 0 | — |
case-16 | fail→pass | 12,812 | 16,847 | +31% | 1 | 1 | 0% | 2,695 | 3,528 | +31% | 0 | 0 | — |
case-17 | fail→fail | 11,712 | 14,684 | +25% | 1 | 1 | 0% | 2,640 | 3,817 | +45% | 0 | 0 | — |
case-18 | pass→pass | 14,699 | 19,323 | +31% | 1 | 1 | 0% | 2,665 | 3,938 | +48% | 0 | 0 | — |
case-19 | pass→pass | 17,145 | 19,185 | +12% | 1 | 1 | 0% | 2,664 | 3,757 | +41% | 0 | 0 | — |
case-20 | pass→pass | 12,324 | 20,272 | +64% | 1 | 1 | 0% | 2,665 | 4,267 | +60% | 0 | 0 | — |
case-21 | pass→fail | 20,666 | 27,330 | +32% | 1 | 1 | 0% | 3,840 | 5,113 | +33% | 0 | 0 | — |
case-22 | pass→pass | 7,396 | 9,884 | +34% | 1 | 1 | 0% | 1,557 | 2,858 | +84% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.