Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Code and interpret qualitative data from files to generate themes, identify patterns, and explore coverage gaps. Use when analyzing interview transcripts, survey responses, documents, or any text data to identify codes, develop themes, assess analytical coverage, and determine what areas have been effectively explored versus potentially avoided. Supports both reflexive thematic analysis (Braun & Clarke) and codebook approaches for research rigor. Optimized for chat-based collaborative workflows.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 193% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 302% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 247% | 0% |
| case-18 | ✓→✓ | = Same ✓ | 242% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 257% | 0% |
Conduct rigorous thematic analysis on qualitative data from files. This skill guides the process of coding data, generating themes, and interpreting patterns while maintaining analytical transparency. Designed for conversational collaboration through chat.
This skill is optimized for conversational thematic analysis where you work with the user through the analytical process:
┌─────────────────────────────────────────────────────────────────┐
│ PHASE 1: Familiarization ─────► Agent performs │
│ PHASE 2: Initial Coding ─────► Agent performs │
│ │
│ [PAUSE: Present codes to user, get feedback] │
│ │
│ PHASE 3: Searching Themes ─────► Conversational (with user) │
│ PHASE 4: Reviewing Themes ─────► Conversational (with user) │
│ PHASE 5: Defining Themes ─────► Conversational (with user) │
│ │
│ [PAUSE: Get user permission] │
│ │
│ PHASE 6: Producing Report ─────► Agent performs (autonomous) │
└─────────────────────────────────────────────────────────────────┘When the user provides data files (transcripts, documents, etc.), confirm receipt and begin the autonomous phases.
Read through all data, develop familiarization notes, and generate initial codes with emoji prefixes.
Share the coded excerpts using the emoji coding system (see below). Ask for feedback before proceeding.
Work through theme development, review, and definition interactively with the user.
Once themes are finalized, ask: "Shall I proceed to produce the final report?" Then generate autonomously.
Use emoji prefixes to visually distinguish codes in chat-based presentation. This makes patterns immediately visible. The agent should select emojis that intuitively represent each emergent code/theme—there are no pre-defined assignments. Let the data guide the choice.
ALL emojis, codes, and themes shown in the examples below are purely ILLUSTRATIVE. They are NOT templates to follow or default codes to use.
DO NOT:
DO:
> The examples below show the format of how to present coded data. The content must be entirely derived from the actual data provided by the user.
[EMOJI] CODE_NAME: "Coded text excerpt from the data"The agent should choose emojis based on the emergent content of each code:
> Participant P3: "I just feel completely overwhelmed by the constant stream of messages. It's like I'm drowning and can't come up for air." > > Agent Coding: > - 📧 EMAIL-OVERLOAD: "constant stream of messages" > - 🌊 EMOTIONAL-DISTRESS: "I'm drowning and can't come up for air" > - 😵 LOSS-OF-CONTROL: "completely overwhelmed"
Always present a code key at the start of coded data, showing the emoji chosen for each code:
📋 CODE KEY FOR THIS SESSION:
📧 EMAIL-OVERLOAD | 🌊 EMOTIONAL-DISTRESS | 😵 LOSS-OF-CONTROL
🏠 HOME-BOUNDARY | ⚡ ENERGY-DRAIN | 🔄 ADAPTATION> 🔴 Remember: These are EXAMPLE codes only. Your actual codes MUST be derived entirely from the data you receive—not borrowed from these illustrations.
Goal: Develop deep, intimate knowledge of your dataset.
Actions:
Key Questions:
Output: Familiarization notes, initial memos, researcher positionality statement
To User: Present a brief summary: "I've reviewed all [N] transcripts. Initial observations: [2-3 patterns]. Ready to proceed with coding."
Goal: Systematically label interesting features across the entire dataset.
Coding Principles:
Types of Codes:
| Type | Description | Example | |------|-------------|----------| | Descriptive | Labels surface content | 📊 e.g., DATA-TYPE] | | Interpretive | Captures underlying meaning | 🌀 e.g., INTERPRETATION] | | In Vivo | Uses participants' exact words | 📝 "Participant's own words]" | | Process | Uses gerunds (-ing) for actions | 🔄 e.g., DOING-ACTION] |
Examples show format only—use codes derived from YOUR data, not these illustrations.
Chat Presentation Format:
Present coded excerpts to the user with:
Example Presentation:
> 🔴 The codes and emojis below are EXAMPLE ILLUSTRATIONS only. Do not use WORK-BOUNDARIES, ADAPTATION, OVERLOAD, or STRESS as default codes. Generate your own codes and emojis based solely on the actual data.
📋 CODE KEY:
🏠 WORK-BOUNDARIES | 🔄 ADAPTATION | 📡 OVERLOAD | 😵 STRESS
═══════════════════════════════════════════════════════════════
🏠 WORK-BOUNDARIES (appears 12 times across 5 participants)
P2: "I had to learn to turn off notifications after 6pm." 📵 SETTING-LIMITS
P4: "My family knows if my door is closed, I'm working." 🚪 PHYSICAL-BOUNDARIES
═══════════════════════════════════════════════════════════════
🔄 ADAPTATION (appears 8 times across 4 participants)
...
═══════════════════════════════════════════════════════════════
📡 OVERLOAD (appears 15 times across 6 participants)
...
[Continue for all codes...]
───────────────────────────────────────────────────────────────
🤔 I've identified 8 codes across the dataset.
Do these capture what you're seeing? Any codes to merge,
split, or add before we move to theme development?Best Practices:
Common Pitfall: Creating too many tiny codes—group similar codes early
Goal: Group codes into potential themes that tell a story — work through this conversationally with the user.
Understanding Themes vs. Codes:
| Codes | Themes | |-------|--------| | Building blocks (bricks) | Central organizing concepts (walls) | | Micro-level, specific | Macro-level, interpretive | | Many (50-80) | Few (4-8 main themes) | | Describe | Interpret |
Conversational Process:
> 🔴 The theme example below is ILLUSTRATIVE. "Navigating Digital Boundaries," "WORK-BOUNDARIES," "ADAPTATION," and "OVERLOAD" are placeholders. Your actual themes must emerge from YOUR data and codes—not from these examples.
Agent opens:
Let's develop themes together. Looking at the codes, I see potential
connections. Here's one grouping I'm considering:
┌────────────────────────────────────────────────────────┐
│ THEME CANDIDATE: "[Theme Name Derived from Your Data]" │
├────────────────────────────────────────────────────────┤
│ [EMOJI] CODE-1-FROM-YOUR-DATA │
│ [EMOJI] CODE-2-FROM-YOUR-DATA │
│ [EMOJI] CODE-3-FROM-YOUR-DATA │
├────────────────────────────────────────────────────────┤
│ Central Concept: [Interpretive concept based on the │
│ patterns in your specific dataset] │
└────────────────────────────────────────────────────────┘
Does this grouping make sense to you? Do you see alternative
ways to group these codes?Agent facilitates dialogue:
What Makes a Good Theme:
Theme Evolution Example
| Weak (Topic Summary) | Strong (Interpretive Theme) | |---------------------|----------------------------| | "Technology problems" | "Navigating system fractures" | | "Communication issues" | "Emotional disconnection in virtual encounters" | | "Work stress" | "The invisible labor of digital presence" |
These show the pattern of how to evolve themes (descriptive → interpretive). Do not assume your data will contain themes about technology, communication, or work stress.
Output: 4-8 candidate themes with central concepts, preliminary thematic map
Transition to Phase 4: "We now have [N] candidate themes. Ready to review and refine them together?"
Goal: Ensure themes are coherent, distinct, and represent the data accurately — work through this conversationally with the user.
Two-Level Review:
Level 1 - Internal Coherence:
Level 2 - External Validity:
Conversational Review Format:
> 🔴 This is a FORMAT TEMPLATE only. Replace "THEME NAME]," "CODE-1]," etc. with actual themes and codes from YOUR analysis.
Let's review each theme together. Starting with:
[EMOJI] THEME: "[Theme Name from Your Data]"
Codes included:
- [CODE-1] (X instances)
- [CODE-2] (Y instances)
- [CODE-3] (Z instances)
Sample quotes:
"[Actual quote from data]" (P#)
"[Another quote from data]" (P#)
QUESTIONS FOR REVIEW:
□ Does "[Theme Name]" capture what unites these codes?
□ Are there quotes that DON'T fit this theme?
□ Is this theme distinct from the others we've discussed?
What's your sense — does this theme hold together? Anything
you'd change?Patton's Dual Criteria:
| Criterion | Question | Application | |-----------|----------|-------------| | Internal Homogeneity | Do data within the theme fit together? | Check for coherence | | External Heterogeneity | Are themes clearly distinct? | Check for overlap |
Refinement Actions (discuss with user):
| Issue | Action | Example | |-------|--------|---------| | Theme too broad | Split | "Challenges" → "Technical barriers" + "Social barriers" | | Themes overlap | Collapse | "Time pressure" + "Workload" → "Time demands" | | Theme unclear | Redefine | "Stuff about work" → "Negotiating professional identity" | | Insufficient data | Discard | Remove themes with only 1-2 quotes | | Missing pattern | Return to coding | Add new codes for overlooked areas |
Agent's role:
Transition to Phase 5: "Themes are feeling solid. Shall we finalize the definitions and names together?"
Goal: Finalize themes by articulating their essence and creating evocative names — work through this conversationally with the user.
For Each Theme, Develop Together:
Agent presents template, user contributes:
> 🔴 This is a TEMPLATE FORMAT only. The theme "Navigating Digital Boundaries" and its definition are EXAMPLES showing the structure. Replace with themes derived from YOUR data.
Let's define Theme 1: "[Theme Name from Your Data]"
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
DEFINITION:
[Agent proposes based on YOUR data and our conversation]
[Describe the central concept that unites the codes in this theme,
based on patterns found in the actual dataset]
[Ask user] Does this capture the essence? What would you add or change?
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SCOPE:
- Includes: [develop together based on your data]
- Excludes: [develop together based on your data]
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SUB-THEMES:
1. [Agent suggests based on patterns in YOUR data] [description]
2. [Agent suggests based on patterns in YOUR data] [description]
[Ask user] Do these sub-themes match your understanding? Others?
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
CONTRIBUTION TO RESEARCH QUESTION:
[develop together]
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
FINAL NAME OPTIONS:
A. "[Agent's suggestion based on your data]"
B. [User can propose alternatives]
C. "[Alternative evocative name based on your data]"
Which resonates most? Or shall we brainstorm others?Naming Guidelines
| Weak Name | Strong Name | Why | |-----------|-------------|-----| | "Problems" | "Navigating systemic neglect" | Captures action + meaning | | "Feelings" | "Emotional labor of care" | Specific + interpretive | | "Technology" | "Digital disruption of intimacy" | Evocative + conceptual |
These show the pattern of weak → strong naming. Do not assume themes about problems, feelings, or technology will be in the data.
Output: Finalized themes with definitions, scope, sub-themes, and evocative names
Transition to Phase 6: "Themes are now finalized. Shall I proceed to produce the final report? I'll generate it autonomously based on what we've developed together."
Goal: Present findings with vivid examples and analytical narrative — proceed autonomously after receiving user permission.
Trigger: User says "yes," "proceed," "generate report," etc.
Report Structure:
Using Quotes Effectively:
Agent delivers:
Coverage Dimensions:
| Dimension | Questions to Ask | Red Flags | |-----------|------------------|-----------| | Depth | Are themes richly developed with nuance? | Thin themes with few quotes | | Breadth | Do themes capture the full dataset? | Important topics missing | | Balance | Are all participants/themes represented? | Over-reliance on outliers | | Reflexivity | Has researcher position been examined? | No discussion of bias |
Signs of Avoidance:
Techniques to Surface Avoided Areas:
| Research Question | Themes Addressing | Gaps Identified | |-------------------|-------------------|-----------------| | RQ1: Experience of X | Theme A, Theme B | Missing: negative experiences | | RQ2: Response to Y | Theme C | Missing: long-term effects |
| # | Criterion | Check | |---|-----------|-------| | 1 | Data transcribed to appropriate level | ☐ | | 2 | Each data item given equal attention | ☐ | | 3 | Themes not from vivid examples only | ☐ | | 4 | Themes capture something important about data | ☐ | | 5 | Data within themes cohere meaningfully | ☐ | | 6 | Clear distinction between themes | ☐ | | 7 | Themes internally coherent and distinctive | ☐ | | 8 | Analysis matches theoretical framework | ☐ | | 9 | Data extracts relate to identified themes | ☐ | | 10 | Data extracts sufficient to support claims | ☐ | | 11 | Analysis makes sense relative to extracts | ☐ | | 12 | Analysis goes beyond description | ☐ | | 13 | Analysis is nuanced and sophisticated | ☐ | | 14 | Analysis tells convincing story about data | ☐ | | 15 | Analysis addresses research question | ☐ |
| Strategy | Application | |----------|-------------| | Prolonged engagement | Spend sufficient time with data | | Peer debriefing | Regular discussions with user | | Negative case analysis | Actively seek disconfirming evidence | | Thick description | Rich, contextualized reporting | | Audit trail | Document all analytical decisions | | Reflexivity journal | Track assumptions and positionality |
For collaborative or structured approaches, see references/codebook-guide.md.
THEME: [Name]
├── Code 1: [Name]
│ ├── Definition
│ ├── Inclusion criteria
│ ├── Exclusion criteria
│ └── Examples (2-3)
├── Code 2: [Name]
│ └── ...When using codebook approaches:
| Pitfall | Solution | |---------|----------| | Theme = Topic Summary | Ensure themes have central organizing concept | | "Themes emerged" | Use active language: "we identified/developed" | | Positivism Creep | Maintain methodological coherence | | Frequency = Significance | Consider meaning, not just frequency | | Insufficient Familiarization | Read data multiple times before coding | | Theme Overload | Aim for 4-8 main themes | | Ignoring Reflexivity | Document researcher position throughout | | Agent Dominance | Ensure user drives interpretive decisions |
| Use Case | Tool | Notes | |----------|------|-------| | Small projects (<20 interviews) | Spreadsheet + manual coding | Low learning curve | | Free QDA software | Taguette, QualCoder | Good functionality | | Commercial QDA | NVivo, MAXQDA, ATLAS.ti | Advanced features | | Team collaboration | Dedoose | Real-time collaboration | | AI assistance | Use for suggestions only | Human must lead analysis |
Remember: Thematic analysis is iterative, not linear. Expect to cycle back through phases as your understanding deepens. The conversational workflow ensures user insight guides the interpretive work at the crucial theme development stages.
Other measured skills in the registry, with their headline benchmark lift.