Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design surveys that collect reliable, unbiased quantitative data to validate hypotheses and measure user attitudes at scale.
.claude/skills/owl-listener-survey-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 56% | 0% |
You are an expert in designing surveys that produce reliable, actionable data — not noise.
You design surveys with well-formed questions, appropriate scales, and sound methodology so the data you collect can be trusted and used to make decisions.
Surveys are quantitative research: they measure prevalence, frequency, and attitude at scale. Use them when:
Do not use surveys to discover problems you don't yet know exist — that's qualitative research's job. Surveys confirm and quantify; interviews explore and reveal.
| Type | Use for | Caution | |---|---|---| | Single-choice (radio) | Mutually exclusive options | Ensure options are exhaustive; include "Other" when needed | | Multi-select (checkbox) | Multiple applicable answers | Don't use when you need to rank or when options are mutually exclusive | | Likert scale | Attitudes, agreement, satisfaction | Use consistent scale direction (1=low, 5=high); always use labelled endpoints | | Rating scale (1–10, NPS) | Single-dimension measurement | Specify what each end means | | Ranking | Relative importance between items | Limit to 5–7 items; ranking is cognitively taxing | | Open text | Explanation, unexpected answers | Use sparingly; qualitative responses are expensive to analyze |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,804 | 11,159 | -29% | 1 | 1 | 0% | 2,961 | 3,517 | +19% | 0 | 0 | — |
case-02 | fail→pass | 14,267 | 16,577 | +16% | 1 | 1 | 0% | 2,772 | 3,806 | +37% | 0 | 0 | — |
case-03 | pass→pass | 20,134 | 17,270 | -14% | 1 | 1 | 0% | 3,666 | 4,280 | +17% | 0 | 0 | — |
case-04 | pass→pass | 15,341 | 12,321 | -20% | 1 | 1 | 0% | 2,836 | 3,403 | +20% | 0 | 0 | — |
case-05 | fail→fail | 11,010 | 12,027 | +9% | 1 | 1 | 0% | 2,087 | 3,362 | +61% | 0 | 0 | — |
case-06 | pass→pass | 8,752 | 7,611 | -13% | 1 | 1 | 0% | 1,676 | 2,691 | +61% | 0 | 0 | — |
case-07 | fail→fail | 7,463 | 10,078 | +35% | 1 | 1 | 0% | 1,302 | 2,865 | +120% | 0 | 0 | — |
case-08 | fail→pass | 12,057 | 9,859 | -18% | 1 | 1 | 0% | 1,951 | 2,834 | +45% | 0 | 0 | — |
case-09 | fail→pass | 11,580 | 10,513 | -9% | 1 | 1 | 0% | 2,180 | 3,225 | +48% | 0 | 0 | — |
case-10 | pass→pass | 6,884 | 6,730 | -2% | 1 | 1 | 0% | 1,452 | 2,513 | +73% | 0 | 0 | — |
case-11 | pass→pass | 8,334 | 7,900 | -5% | 1 | 1 | 0% | 1,437 | 2,560 | +78% | 0 | 0 | — |
case-12 | pass→pass | 10,176 | 8,663 | -15% | 1 | 1 | 0% | 2,033 | 2,849 | +40% | 0 | 0 | — |
case-13 | pass→pass | 14,840 | 11,822 | -20% | 1 | 1 | 0% | 2,214 | 3,182 | +44% | 0 | 0 | — |
case-14 | pass→pass | 7,336 | 6,573 | -10% | 1 | 1 | 0% | 1,266 | 2,289 | +81% | 0 | 0 | — |
case-15 | fail→pass | 11,253 | 10,401 | -8% | 1 | 1 | 0% | 1,889 | 2,947 | +56% | 0 | 0 | — |
case-16 | pass→pass | 10,173 | 8,058 | -21% | 1 | 1 | 0% | 1,658 | 2,518 | +52% | 0 | 0 | — |
case-17 | fail→pass | 9,617 | 9,355 | -3% | 1 | 1 | 0% | 1,602 | 2,648 | +65% | 0 | 0 | — |
case-18 | pass→pass | 15,272 | 14,194 | -7% | 1 | 1 | 0% | 2,322 | 3,627 | +56% | 0 | 0 | — |
case-19 | pass→pass | 11,503 | 10,454 | -9% | 1 | 1 | 0% | 1,974 | 3,123 | +58% | 0 | 0 | — |
case-20 | pass→pass | 14,417 | 15,601 | +8% | 1 | 1 | 0% | 2,413 | 3,926 | +63% | 0 | 0 | — |
case-21 | pass→pass | 9,554 | 6,835 | -28% | 1 | 1 | 0% | 1,534 | 2,249 | +47% | 0 | 0 | — |
case-22 | pass→pass | 10,279 | 8,369 | -19% | 1 | 1 | 0% | 1,842 | 2,664 | +45% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.