Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Workflow for automated QA: integration and end-to-end UI test automation, page objects, etc.
.claude/skills/griddynamics-ui-aqa-flow/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 134% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 263% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 112% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 194% | 0% |
<ui_aqa_flow>
<description_and_purpose>
End-to-end test automation from requirements gathering to test implementation. Uses test cases, project documentation to create automated tests following existing architecture and coding standards.
Prerequisite: Rosetta Prep Steps.
Terminology. External systems are named by role throughout this workflow and its phases: Test Management System (TMS), Issue Tracker, and Wiki. TestRail, Jira, and Confluence are canonical examples only — adapt identifiers, URLs, requests, calls, and query syntax to the systems resolved for the current project (from repository-root gain.json, explicit user input, recognizable URLs/handles, and available integrations).
</description_and_purpose>
<workflow_phases>
Execution cadence:
load-project-context, orchestration, hitlagents/TEMP/<FEATURE>/ui-aqa-state.md → next; never start a phase until the previous is marked done in ui-aqa-state.md.<orchestration_and_escalation>.No assumptions:
Customization:
Authoritative rules (do not skim past):
coding before any work touching repository tests, page objects, or shared helpers — authoritative for conventions; repository docs win over skill snippets.### Explicit Assertions in the test plan), enforced when Phase 6 implements tests.fixme spec · abort — and WAIT for the user's explicit choice. "Skip clarification" waives clarification questions only; it never authorizes this feasibility/scope call.<data_collection phase="1" applies="ALL" subagent="discoverer" role="UI-AQA data collector" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3.1-pro, grok-4.5, gpt-5.6-terra">
phases/ui-aqa-flow-data-collection.mdgain.json + project context at its configured paths (canonical: docs/CONTEXT.md, docs/ARCHITECTURE.md, agents/IMPLEMENTATION.md). Output: test plan at plans/ui-aqa-<test-name>/test-plan.mddata-collection, sensitive-data, qa-structure, qa-knowledgeagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 1 is not complete until its output spot-check passes.</data_collection>
<requirements_clarification phase="2" applies="ALL" subagent="architect" role="Test requirements analyst" subagent_required_model="claude-opus-4-8, gpt-5.5-high, gemini-3.1-pro-high, gpt-5.6-sol" type="HITL">
phases/ui-aqa-flow-requirements-clarification.mdqa-knowledge (gap_analysis mode), qa-structurequestioningagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 2 is not complete until its output spot-check passes.</requirements_clarification>
<code_analysis phase="3" applies="ALL" subagent="discoverer" role="Test architecture analyst" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3.1-pro, grok-4.5, gpt-5.6-terra">
phases/ui-aqa-flow-code-analysis.mdplans/ui-aqa-<test-name>/code-analysis.md (architecture patterns, existing page objects, test patterns)qa-knowledge (code_analysis mode), reverse-engineering, sensitive-data, qa-structureagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 3 is not complete until its output spot-check passes.</code_analysis>
<selector_identification phase="4" applies="ALL" subagent="engineer" role="Selector identification specialist" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra" type="HITL-CONDITIONAL">
phases/ui-aqa-flow-selector-identification.mdqa-knowledge (implementation_modes — selector mode Part A), qa-structure, sensitive-datatestingagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 4 is not complete until its output spot-check passes.</selector_identification>
<selector_implementation phase="5" applies="ALL" subagent="engineer" role="Selector implementation specialist" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra">
phases/ui-aqa-flow-selector-implementation.mdqa-knowledge (implementation_modes — selector mode Part B), qa-structuretesting, codingagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 5 is not complete until its output spot-check passes.</selector_implementation>
<test_implementation phase="6" applies="ALL" subagent="engineer" role="Test automation engineer" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra" type="HITL">
phases/ui-aqa-flow-test-implementation.mdqa-knowledge (implementation_modes — UI impl), qa-structuretesting, codingagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 6 is not complete until its output spot-check passes.</test_implementation>
<test_report_analysis phase="7" applies="ALL" subagent="engineer" role="Test failure analyst" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra" type="HITL">
phases/ui-aqa-flow-test-report-analysis.mdagents/user-instructions/). Output: failure analysis with root causes + fix recommendationsagents/user-instructions/).qa-knowledge (test_execution_triage mode), sensitive-data, qa-structureagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 7 is not complete until its output spot-check passes.</test_report_analysis>
<test_corrections phase="8" applies="ALL" subagent="engineer" role="Test correction engineer" subagent_required_model="claude-sonnet-5, gpt-5.4-medium, gemini-3-flash, grok-4.5, gpt-5.6-terra" type="HITL">
phases/ui-aqa-flow-test-correction.mdphases/ui-aqa-flow-test-correction.md section <present_for_approval>.qa-knowledge (correction mode), qa-structuredebugging, coding, hitlagents/TEMP/<FEATURE>/ui-aqa-state.md; Phase 8 is not complete until its output spot-check passes.</test_corrections>
</workflow_phases>
<orchestration_and_escalation>
hitl skill): a skip asserted but contradicted by agents/TEMP/<FEATURE>/ui-aqa-state.md / disk evidence is refused — announce the specific missing state row / absent artifact, then start the earliest incomplete phase the same turn. Audit-trail row → the state file's ## Verification-Failure Overrides (template owned by qa-structure).type="HITL" / type="HITL-CONDITIONAL" — those type= attributes are the sole source of truth for this workflow — plus safety/destructive confirmations. Any skip outside the refusal rule above requires explicit user confirmation (HITL).type="HITL" phase the subagent surfaces the question/blocker and returns; the orchestrator runs the gate with the user and only then resumes — never inferred, critical on the destructive Phase 8. Dispatch per USE SKILL orchestration.ui-aqa-state.md, ask the user; never substitute silently.</orchestration_and_escalation>
<workflow_success_criteria>
agents/TEMP/<FEATURE>/ui-aqa-state.md, the artifacts those phases reference exist, and the user accepts the last test outcome or explicitly stops. In-scope = default execution plus user-approved customization/skip decisions under <orchestration_and_escalation>.<test-name> slug (no fabricated/placeholder slug) · P2 — ### Explicit Assertions in the plan (≥1 typed bullet or None-clause) · P3 — code analysis report populated (architecture + page-object inventory + test location) · P4 — ## Selector Management Part A deliverables in the plan · P5 — every identified selector in the updated page objects, lint-clean · P6 — test file exists lint-clean + ## Test Implementation record with all five subsections (incl. ### Uncovered Assertions / None-clause) · P7 — failure analysis when failures occurred, or state rows reconciled to N/A — 0 failures / None · P8 — user-approved edits applied when failures occurred, or state records N/A — no corrections after a zero-failure run.ui-aqa-state.md, flag uncertainty, stop for user guidance.</workflow_success_criteria>
<state_file>
agents/TEMP/<FEATURE>/ui-aqa-state.md — created/updated after each phase from the template owned by qa-structure (its state-file skeleton asset, loaded at Phase 1). It carries ## Phase Completion Status, ## Key Artifacts & Facts (the resume anchor — only what resume-after-compaction needs), and ## Verification-Failure Overrides.
</state_file>
<references>
Subagents: discoverer · architect · engineer.
Cross-phase skills: qa-structure (paths / <test-name> slug / state-file shape) and qa-knowledge (modes, taxonomy, artifact skeletons — loads its own assets at point of use).
Integrations: TMS and Wiki per <description_and_purpose> Terminology, plus browser automation (Playwright is the canonical example).
</references>
</ui_aqa_flow>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 25,153 | 6,033 | -76% | 1 | 1 | 0% | 5,276 | 3,697 | -30% | 0 | 0 | — |
case-02 | fail→fail | 21,862 | 6,229 | -72% | 1 | 1 | 0% | 4,530 | 3,643 | -20% | 0 | 0 | — |
case-03 | pass→pass | 9,664 | 3,562 | -63% | 1 | 1 | 0% | 1,348 | 3,962 | +194% | 0 | 0 | — |
case-04 | fail→pass | 13,833 | 4,510 | -67% | 1 | 1 | 0% | 2,574 | 4,093 | +59% | 0 | 0 | — |
case-05 | pass→pass | 9,074 | 6,993 | -23% | 1 | 1 | 0% | 1,537 | 4,411 | +187% | 0 | 0 | — |
case-06 | fail→fail | 6,274 | 10,193 | +62% | 1 | 1 | 0% | 968 | 4,204 | +334% | 0 | 0 | — |
case-07 | fail→pass | 9,841 | 2,066 | -79% | 1 | 1 | 0% | 1,570 | 3,672 | +134% | 0 | 0 | — |
case-08 | fail→pass | 6,340 | 2,267 | -64% | 1 | 1 | 0% | 1,025 | 3,723 | +263% | 0 | 0 | — |
case-09 | fail→pass | 12,256 | 4,598 | -62% | 1 | 1 | 0% | 1,902 | 4,029 | +112% | 0 | 0 | — |
case-10 | fail→pass | 9,117 | 4,536 | -50% | 1 | 1 | 0% | 1,398 | 4,116 | +194% | 0 | 0 | — |
case-11 | pass→pass | 7,479 | 5,198 | -30% | 1 | 1 | 0% | 1,208 | 4,157 | +244% | 0 | 0 | — |
case-12 | pass→pass | 7,671 | 4,377 | -43% | 1 | 1 | 0% | 1,298 | 4,020 | +210% | 0 | 0 | — |
case-13 | fail→pass | 9,134 | 2,719 | -70% | 1 | 1 | 0% | 1,486 | 3,780 | +154% | 0 | 0 | — |
case-14 | fail→pass | 13,317 | 6,458 | -52% | 1 | 1 | 0% | 2,107 | 4,306 | +104% | 0 | 0 | — |
case-15 | pass→pass | 9,219 | 4,353 | -53% | 1 | 1 | 0% | 1,516 | 3,981 | +163% | 0 | 0 | — |
case-16 | fail→pass | 12,762 | 5,034 | -61% | 1 | 1 | 0% | 1,919 | 4,288 | +123% | 0 | 0 | — |
case-17 | fail→pass | 12,822 | 3,612 | -72% | 1 | 1 | 0% | 1,872 | 3,953 | +111% | 0 | 0 | — |
case-18 | pass→pass | 11,429 | 4,462 | -61% | 1 | 1 | 0% | 1,738 | 4,128 | +138% | 0 | 0 | — |
case-19 | fail→pass | 11,215 | 4,203 | -63% | 1 | 1 | 0% | 1,648 | 4,118 | +150% | 0 | 0 | — |
case-20 | pass→pass | 9,254 | 8,584 | -7% | 1 | 1 | 0% | 1,553 | 4,929 | +217% | 0 | 0 | — |
case-21 | fail→pass | 9,850 | 6,449 | -35% | 1 | 1 | 0% | 1,618 | 4,389 | +171% | 0 | 0 | — |
case-22 | pass→pass | 6,825 | 5,271 | -23% | 1 | 1 | 0% | 1,124 | 4,346 | +287% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.