Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when doing upstream market-research methodology — sizing a market as TAM/SAM/SOM computed BOTH top-down and bottoms-up (never a single unsourced number), planning a survey sample size with finite-population correction and per-segment minimums, or scoring candidate market segments against Kotler's measurable/substantial/accessible/differentiable/actionable criteria. Outputs always show the method and the assumptions. For market-research analysts and product-marketing at the sizing/survey/segm
.claude/skills/alirezarezvani-market-research/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 14 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 82% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 142% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 175% | 0% |
Upstream market-research methodology: market sizing, survey/sampling design, and segmentation. The discipline here is method + assumptions: a TAM is never a single number, a survey is never powered only in aggregate, and a segment is never a demographic slice.
Market-research analysts, product marketers, and strategy teams need rigorous evidence before anyone optimizes a campaign or sets a strategy. This skill structures three methodology decisions:
Three deterministic tools:
market_sizer.py — Computes TAM/SAM/SOM by both top-down and bottoms-up methods side-by-side, reports the divergence, and flags failed triangulation. Never returns a single number.sample_size_planner.py — Survey sample size from confidence, margin of error, and expected proportion, with the finite-population correction and per-segment minimums (a survey powered overall is not powered per reported segment).segmentation_scorer.py — Scores candidate segments against Kotler's five criteria and enforces a substantiality + accessibility gate; a slice that is too small or unreachable is dropped.Invoke this skill when:
Do NOT use this skill to: measure a live campaign (attribution, ROAS, CPA → marketing-skill/campaign-analytics), build demand-gen / paid-media plans (marketing-skill/marketing-demand-acquisition), set positioning / GTM strategy (marketing-skill/marketing-strategy-pmm), or set pricing (commercial/pricing-strategist).
assets/market_research_brief_template.md (objective, the decision this informs, sizing approach, sampling plan, assumptions register).market_sizer.py --input market.json --method both --profile {b2b-saas|consumer|enterprise|marketplace|hardware|services}. Reconcile the top-down/bottoms-up delta before quoting anything.sample_size_planner.py --input survey.json. Fund the per-segment floors, not just the overall n.segmentation_scorer.py --input segments.json --profile <same>. Drop segments failing the substantiality/accessibility gate.| Script | Purpose | Profiles | |---|---|---| | scripts/market_sizer.py | TAM/SAM/SOM top-down AND bottoms-up + triangulation flag | b2b-saas, consumer, enterprise, marketplace, hardware, services | | scripts/sample_size_planner.py | Survey n + FPC + per-segment minima | n/a (parameter-driven) | | scripts/segmentation_scorer.py | Kotler 5-criteria scoring + gate | b2b-saas, consumer, enterprise, marketplace, hardware, services |
All three: stdlib-only, --help, --sample, --output {human,json}.
Run the onboarding questionnaire once before you start — it captures your defaults so every tool in this skill is pre-configured. Customization is the point: the answers actually change tool behavior.
bashpython3 scripts/onboard.py # interactive (also: --defaults, --set key=value, --reset) python3 scripts/onboard.py --show # see the questions + current effective config
Answers are saved to ~/.config/research-ops/market-research.json (global) or ./.research-ops/market-research.json (--scope project) and are read automatically by config_loader.py. They set the default market profile, the default survey confidence and margin of error, and the default sizing method. CLI flags always override saved config; RESEARCH_OPS_NO_CONFIG=1 ignores it.
The four questions: market profile · survey confidence · margin of error · sizing method.
This skill ships an isolated, opt-in bridge to engineering/autoresearch-agent. Only when you ask to "optimize" / "reconcile the sizing" / "run a loop" does an autoresearch experiment iteratively reconcile your market model so top-down and bottoms-up triangulate. scripts/ar_evaluator.py is the ground-truth evaluator; it prints tam_divergence: <fraction> (lower is better).
bash/ar:setup --domain custom --name tam-triangulation \ --target market.json \ --eval "python3 ar_evaluator.py --target market.json" \ --metric tam_divergence --direction lower /ar:loop custom/tam-triangulation
Isolated: no hard dependency — autoresearch runs only on demand, and the loop edits market.json, never the evaluator.
references/market_sizing_canon.md — TAM/SAM/SOM frameworks (Bessemer, a16z); top-down vs bottoms-up; Fermi estimation; market-model conventions; common sizing fallacies.references/survey_methodology.md — Cochran Sampling Techniques; Dillman Tailored Design Method; Groves Survey Methodology; question-wording bias (Schuman & Presser); AAPOR standards.references/segmentation_and_ci.md — Kotler segmentation criteria; needs-based vs firmographic; Porter Five Forces; SCIP ethics; Christensen JTBD; conjoint/MaxDiff primer.| Neighbor | Scope | Difference | |---|---|---| | marketing-skill/campaign-analytics | Attribution, ROAS, CPA, funnel of a live campaign | That measures spend deployed; this is upstream methodology | | marketing-skill/marketing-demand-acquisition | Demand-gen, paid media, channel mix | That runs acquisition; this builds the evidence | | marketing-skill/marketing-strategy-pmm | Positioning, GTM, category | That sets strategy; this sizes and segments the market | | commercial/pricing-strategist | Pricing model + WTP + packaging | That sets price; this sizes the market | | product-research (sibling) | User/product discovery methods | That studies users; this studies the market |
bashpython3 scripts/market_sizer.py --sample python3 scripts/sample_size_planner.py --population 62000 --confidence 0.95 --moe 0.05 python3 scripts/segmentation_scorer.py --sample --output json
The sample market triangulates a ~$1.47B top-down SAM against the bottoms-up figure and flags the divergence; the segmentation sample drops the "solopreneurs who might want analytics" slice for failing the substantiality and accessibility gates.
Walked one at a time by /cs:grill-research-ops or the orchestrator. Recommended answer + canon citation per question. Never bundled.
Recommended: both; reconcile the delta before quoting a number. Canon: Bessemer / a16z market-sizing; Fermi estimation.
Recommended: size to the decision's tolerance, not to a spurious-precision number. Canon: market-model conventions (Gartner/Forrester); decision-driven analysis.
Recommended: power each reported segment, not only the total. Canon: Cochran Sampling Techniques; AAPOR standards.
Recommended: pre-test the wording; cite the bias source. Canon: Schuman & Presser; Dillman Tailored Design Method.
Recommended: drop segments that fail substantiality or accessibility. Canon: Kotler segmentation criteria.
Walk depth-first. Lock 1-2 before opening 3-5. After all are answered, invoke market_sizer.py → sample_size_planner.py → segmentation_scorer.py.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-19 | fail→pass | 16,476 | 6,648 | -60% | 1 | 1 | 0% | 3,059 | 3,443 | +13% | 0 | 0 | — |
case-20 | fail→pass | 15,230 | 16,082 | +6% | 1 | 1 | 0% | 2,916 | 5,317 | +82% | 0 | 0 | — |
case-21 | fail→pass | 11,285 | 10,970 | -3% | 1 | 1 | 0% | 2,012 | 4,091 | +103% | 0 | 0 | — |
case-01 | fail→pass | 16,546 | 25,078 | +52% | 1 | 1 | 0% | 3,362 | 8,135 | +142% | 0 | 0 | — |
case-02 | fail→pass | 11,293 | 17,736 | +57% | 1 | 1 | 0% | 2,193 | 6,026 | +175% | 0 | 0 | — |
case-03 | pass→pass | 10,711 | 10,455 | -2% | 1 | 1 | 0% | 2,382 | 4,160 | +75% | 0 | 0 | — |
case-04 | pass→pass | 10,249 | 6,931 | -32% | 1 | 1 | 0% | 2,082 | 3,665 | +76% | 0 | 0 | — |
case-05 | fail→pass | 11,638 | 2,462 | -79% | 1 | 1 | 0% | 769 | 2,745 | +257% | 0 | 0 | — |
case-06 | pass→pass | 7,855 | 8,303 | +6% | 1 | 1 | 0% | 1,791 | 3,986 | +123% | 0 | 0 | — |
case-07 | pass→pass | 13,206 | 11,180 | -15% | 1 | 1 | 0% | 2,382 | 4,395 | +85% | 0 | 0 | — |
case-08 | fail→pass | 10,099 | 3,651 | -64% | 1 | 1 | 0% | 1,707 | 2,910 | +70% | 0 | 0 | — |
case-09 | fail→pass | 9,342 | 2,773 | -70% | 1 | 1 | 0% | 1,795 | 2,873 | +60% | 0 | 0 | — |
case-10 | fail→pass | 9,864 | 4,847 | -51% | 1 | 1 | 0% | 2,041 | 3,277 | +61% | 0 | 0 | — |
case-11 | pass→pass | 8,449 | 5,693 | -33% | 1 | 1 | 0% | 1,522 | 3,371 | +121% | 0 | 0 | — |
case-12 | pass→pass | 9,646 | 7,783 | -19% | 1 | 1 | 0% | 1,592 | 3,519 | +121% | 0 | 0 | — |
case-13 | pass→pass | 10,918 | 2,747 | -75% | 1 | 1 | 0% | 1,902 | 2,708 | +42% | 0 | 0 | — |
case-14 | pass→pass | 7,474 | 2,840 | -62% | 1 | 1 | 0% | 1,711 | 2,881 | +68% | 0 | 0 | — |
case-15 | fail→fail | 15,357 | 3,787 | -75% | 1 | 1 | 0% | 1,314 | 3,008 | +129% | 0 | 0 | — |
case-16 | pass→pass | 8,833 | 8,034 | -9% | 1 | 1 | 0% | 1,658 | 3,713 | +124% | 0 | 0 | — |
case-17 | fail→pass | 7,694 | 1,852 | -76% | 1 | 1 | 0% | 1,372 | 2,563 | +87% | 0 | 0 | — |
case-18 | fail→pass | 5,786 | 4,325 | -25% | 1 | 1 | 0% | 1,065 | 3,105 | +192% | 0 | 0 | — |
case-22 | pass→pass | 11,946 | 9,163 | -23% | 1 | 1 | 0% | 2,162 | 3,995 | +85% | 0 | 0 | — |
case-23 | pass→pass | 5,275 | 2,114 | -60% | 1 | 1 | 0% | 950 | 2,677 | +182% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.