Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Workflow for generating test cases from requirements (Issue Tracker / Wiki sources), exporting to a Test Management System, etc.
.claude/skills/griddynamics-testgen-flow/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 260% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 217% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 551% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 173% | 0% |
<testgen_flow>
<description_and_purpose>
Systematic requirements analysis from Issue Tracker tickets and Wiki documentation to structured requirements and test scenarios. Extracts data, identifies gaps, clarifies unknowns via HITL, generates requirements document, and produces test cases with export to a Test Management System (Phase 6, user-triggered — the user may choose not to trigger it, but it is a fully-specified phase, not informally skippable). Designed for BA/QA engineers and requirements engineers.
Prerequisite: Rosetta Prep Steps.
Terminology. External systems are named by role throughout this workflow and its phases: Issue Tracker, Wiki, and Test Management System (TMS). Jira, Confluence, and TestRail are canonical examples only — adapt identifiers, URLs, requests, calls, and query syntax to the systems resolved for the current project (from repository-root gain.json, explicit user input, recognizable URLs/handles, and available integrations).
</description_and_purpose>
<workflow_phases>
load-project-context, orchestration, hitltestgen-state.md marks the phase complete and its expected output file exists under plans/testgen-{TICKET-KEY}/; otherwise resume from the earliest incomplete phase. The explicit user instruction skip NEVER applies to the Phase 3 / Phase 6 HITL gates — those are rule 2 of <orchestration_and_escalation> and are never overridden.<orchestration_and_escalation> priority hierarchy.plans/testgen-{TICKET-KEY}/testgen-state.md after each phase.## Metrics count is populated (a thin 0/1 → re-check the artifact), and any HITL approval (Phase 3, 6) is recorded.orchestration.plans/testgen-{TICKET-KEY}/ at start.Analyze requirements for PROJ-123 (also: bare key PROJ-123, full ticket URL). Ticket-only and ticket+Wiki input formats are enumerated in APPLY SKILL FILE phases/testgen-flow-project-config-loading.md step 0.1.phases/testgen-flow-data-collection.md + the data-collection skill's Issue Tracker failure handling.phases/testgen-flow-data-collection.md + the data-collection skill's Wiki failure handling.phases/testgen-flow-question-generation.md <failure_handling> "User explicitly declines to answer".phases/testgen-flow-requirements-document-generation.md <failure_handling> "Missing or empty inputs".data-collection skill's documentation-binding search/ranking behavior.phases/testgen-flow-project-config-loading.md.phases/testgen-flow-test-case-export.md (Threshold (80%) met field + PARTIAL — N/M exported state).phases/testgen-flow-test-case-generation.md <validation_checklist>.subagent_required_model): tier: complex = heavy reasoning / multi-source synthesis / requirements engineering (Opus-class / GPT high-tier); tier: workhorse = structured execution / extraction / generation + export (Sonnet-class / GPT mid-tier).<project_config_loading phase="0" subagent="discoverer" role="Project configuration analyst" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3.1-pro, grok-4.5, gpt-5.6-terra">
phases/testgen-flow-project-config-loading.mdplans/testgen-{TICKET-KEY}/initial-data.md, project config file.sensitive-data (config / initial-data redaction pre-write gate)questioningtestgen-state.md; Phase 0 is not complete until its output spot-check passes.</project_config_loading>
<data_collection phase="1" subagent="discoverer" role="Requirements data collector" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3.1-pro, grok-4.5, gpt-5.6-terra">
phases/testgen-flow-data-collection.mdplans/testgen-{TICKET-KEY}/raw-data.md with Issue Tracker + Wiki data.data-collectiontestgen-state.md; Phase 1 is not complete until its output spot-check passes.</data_collection>
<gap_and_contradiction_analysis phase="2" subagent="architect" role="Requirements gap analyst" subagent_required_model="claude-opus-4-8, gpt-5.5-high, gemini-3.1-pro-high, gpt-5.6-sol">
phases/testgen-flow-gap-and-contradiction-analysis.mdplans/testgen-{TICKET-KEY}/analysis.md with contradictions, gaps, ambiguities.qa-knowledge (gap_analysis mode)testgen-state.md; Phase 2 is not complete until its output spot-check passes.</gap_and_contradiction_analysis>
<question_generation phase="3" subagent="architect" role="Requirements clarification analyst" subagent_required_model="claude-opus-4-8, gpt-5.5-high, gemini-3.1-pro-high, gpt-5.6-sol" type="HITL">
phases/testgen-flow-question-generation.mdplans/testgen-{TICKET-KEY}/questions.md, plans/testgen-{TICKET-KEY}/answers.md.questioningtestgen-state.md; Phase 3 is not complete until its output spot-check passes.</question_generation>
<requirements_document_generation phase="4" subagent="architect" role="Requirements engineer" subagent_required_model="claude-opus-4-8, gpt-5.5-high, gemini-3.1-pro-high, gpt-5.6-sol">
phases/testgen-flow-requirements-document-generation.mdplans/testgen-{TICKET-KEY}/requirements.md.qa-knowledge (synthesis mode)testgen-state.md; Phase 4 is not complete until its output spot-check passes.</requirements_document_generation>
<test_case_generation phase="5" subagent="engineer" role="Test case design engineer" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra">
phases/testgen-flow-test-case-generation.mdplans/testgen-{TICKET-KEY}/test-scenarios.mdtest-scenarios.md before Phase 6 export (phase-file gate, step 5.9) — present a summary and require explicit confirmation; per-phase confirmation per <orchestration_and_escalation> priority (3).qa-knowledge (scenario_design mode + config-resolved TMS FORMAT binding).testgen-state.md; Phase 5 is not complete until its output spot-check passes.coding is NOT used for the default manual-scenario output (writes stay under plans/testgen-{TICKET-KEY}/); apply it only if a write targets tracked repo files outside that folder, per <phase_5_6_standards_gate>.</test_case_generation>
<test_case_export phase="6" subagent="engineer" role="Test case export specialist" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra" type="HITL">
phases/testgen-flow-test-case-export.mdplans/testgen-{TICKET-KEY}/export-report.md (TMS IDs/URLs, per-case status, timestamp). The local receipt is the on-disk evidence Phase 6 ran successfully.qa-knowledge (scenario_design mode + config-resolved TMS EXPORT binding).testgen-state.md; Phase 6 is not complete until its output spot-check passes.coding is NOT used for the default flow (TMS export + receipt under plans/testgen-{TICKET-KEY}/); apply it only if a write targets tracked repo files outside that folder (e.g. embedding TMS IDs into a version-controlled file), per <phase_5_6_standards_gate>.</test_case_export>
</workflow_phases>
<orchestration_and_escalation>
hitl skill): a skip asserted but contradicted by testgen-state.md / disk evidence is refused — announce the specific missing state row / absent artifact, then start the earliest incomplete phase the same turn.<phase_5_6_standards_gate> outside-output-dir confirmation; (2) Phase 3 + Phase 6 HITL gates (answer questions.md / confirm TMS target + export scope) — never skipped by user instruction; (3) per-phase user confirmation; (4) the verification-failure override below.type="HITL" attribute (Phases 3 + 6). The priority-(3) per-phase confirmations (Phases 0, 1, 2, 4, 5) are intentional user-pauses that deliberately carry no type= attribute — here type= marks the never-overridden gates, not every pause.testgen-state.md does not mark it AND the expected output is absent. Action — if testgen-state.md is missing, create it from the Phase 0 <state_file_template> first; log a row into its ## Verification-Failure Overrides (row format owned by that template); then start the earliest incomplete phase the same turn without invoking the hitl ask path. Uncertainty (partial state, ambiguous assertion) → fall back to the hitl ask.testgen-state.md, ask the user; never substitute silently.</orchestration_and_escalation>
<phase_5_6_standards_gate>
coding whenever the phase writes any file outside plans/testgen-{TICKET-KEY}/ — including "mixed outputs" (writes both inside and outside) or when repository edit scope was not explicitly confirmed in chat. When in doubt, apply.cypress/e2e/login.spec.ts; skip when writing only plans/testgen-{TICKET-KEY}/test-scenarios.md.</phase_5_6_standards_gate>
<state_and_outputs>
testgen-state.md (sections ## Phase Completion Status, ## Phase Details, ## Metrics, ## Verification-Failure Overrides) and the per-ticket output-directory layout are initialized and owned by Phase 0 (APPLY SKILL FILE phases/testgen-flow-project-config-loading.md) — see it for the canonical layout. Each subsequent phase updates the state file per <workflow_phases> and writes its output (paths in each phase block) under plans/testgen-{TICKET-KEY}/.
Expected per-ticket artifact set (one validation can confirm all phases ran): initial-data.md + project config (Phase 0) · raw-data.md (1) · analysis.md (2) · questions.md + answers.md (3) · requirements.md (4) · test-scenarios.md (5) · export-report.md (6) · testgen-state.md (all). Full schema/layout owned by Phase 0.
</state_and_outputs>
<references>
Subagents: discoverer · architect · engineer.
Cross-phase skills: qa-knowledge (gap analysis, synthesis, scenario design + TMS bindings — loads its own assets at point of use); data-collection (Issue Tracker + Wiki collection); questioning; sensitive-data (redaction — Phase 0 config/initial-data pre-write gate; Phase 1 collection runs it via data-collection); coding (conditional — only for writes to tracked repo files outside plans/testgen-{TICKET-KEY}/, per <phase_5_6_standards_gate>).
Integrations: Issue Tracker (ticket data extraction), Wiki (documentation retrieval), TMS (Phase 6 export), and additional documentation stores (e.g. Google Drive) when configured — per <description_and_purpose> Terminology.
</references>
<best_practices>
requirements.md + exported cases back to the source ticket (attachment/comment) and archive testgen-state.md + outputs for traceability</best_practices>
<validation_checklist>
</validation_checklist>
<pitfalls>
</pitfalls>
</testgen_flow>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,586 | 4,267 | -50% | 1 | 1 | 0% | 1,399 | 4,429 | +217% | 0 | 0 | — |
case-02 | fail→fail | 7,564 | 5,749 | -24% | 1 | 1 | 0% | 1,111 | 4,282 | +285% | 0 | 0 | — |
case-03 | fail→fail | 7,992 | 6,116 | -23% | 1 | 1 | 0% | 1,107 | 4,554 | +311% | 0 | 0 | — |
case-04 | fail→pass | 9,640 | 5,533 | -43% | 1 | 1 | 0% | 1,376 | 4,949 | +260% | 0 | 0 | — |
case-05 | fail→fail | 7,674 | 7,071 | -8% | 1 | 1 | 0% | 1,202 | 5,203 | +333% | 0 | 0 | — |
case-06 | fail→pass | 10,510 | 3,460 | -67% | 1 | 1 | 0% | 1,450 | 4,596 | +217% | 0 | 0 | — |
case-07 | fail→fail | 7,604 | 2,978 | -61% | 1 | 1 | 0% | 1,133 | 4,463 | +294% | 0 | 0 | — |
case-08 | fail→fail | 7,920 | 6,271 | -21% | 1 | 1 | 0% | 1,298 | 5,115 | +294% | 0 | 0 | — |
case-09 | pass→pass | 5,583 | 2,552 | -54% | 1 | 1 | 0% | 831 | 4,310 | +419% | 0 | 0 | — |
case-10 | pass→pass | 8,574 | 2,671 | -69% | 1 | 1 | 0% | 1,284 | 4,440 | +246% | 0 | 0 | — |
case-11 | fail→pass | 8,841 | 3,927 | -56% | 1 | 1 | 0% | 712 | 4,635 | +551% | 0 | 0 | — |
case-12 | fail→fail | 7,827 | 5,433 | -31% | 1 | 1 | 0% | 1,131 | 4,978 | +340% | 0 | 0 | — |
case-13 | fail→pass | 13,700 | 1,800 | -87% | 1 | 1 | 0% | 2,172 | 4,217 | +94% | 0 | 0 | — |
case-14 | fail→pass | 10,490 | 3,060 | -71% | 1 | 1 | 0% | 1,648 | 4,493 | +173% | 0 | 0 | — |
case-15 | pass→pass | 4,858 | 4,021 | -17% | 1 | 1 | 0% | 868 | 4,658 | +437% | 0 | 0 | — |
case-16 | pass→pass | 9,193 | 5,786 | -37% | 1 | 1 | 0% | 1,569 | 5,010 | +219% | 0 | 0 | — |
case-17 | fail→pass | 8,900 | 2,806 | -68% | 1 | 1 | 0% | 1,350 | 4,330 | +221% | 0 | 0 | — |
case-18 | fail→pass | 9,029 | 2,249 | -75% | 1 | 1 | 0% | 1,407 | 4,326 | +207% | 0 | 0 | — |
case-19 | fail→pass | 6,051 | 3,263 | -46% | 1 | 1 | 0% | 1,022 | 4,466 | +337% | 0 | 0 | — |
case-20 | pass→pass | 6,315 | 6,069 | -4% | 1 | 1 | 0% | 1,024 | 5,086 | +397% | 0 | 0 | — |
case-21 | pass→pass | 9,727 | 8,186 | -16% | 1 | 1 | 0% | 1,595 | 5,201 | +226% | 0 | 0 | — |
case-22 | pass→pass | 5,097 | 9,595 | +88% | 1 | 1 | 0% | 782 | 5,669 | +625% | 0 | 0 | — |
case-23 | fail→pass | 6,970 | 11,479 | +65% | 1 | 1 | 0% | 1,078 | 5,781 | +436% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.