▸case-03 I need to establish a consistent feedback classification framework to normalize customer input across community chatter, support tickets, and sales calls. Please create a comprehensive taxonomy layout covering persona details, lifecycle stages, feedback drivers, sentiment dimensions, and metadata attributes, complete with a quarterly governance checklist. | fail→pass | 27,357 | 24,179 | -12% | 1 | 1 | 0% | 4,156 | 4,482 | +8% | 0 | 0 | — |
▸case-02 Our voice of customer data across user interviews and surveys lacks standardized tagging standards before we synthesize it. Could you design a multi-layered classification model for us that captures user personas, customer lifecycle phases, underlying drivers, sentiment metrics, and contextual metadata? Format the output as a clean template outline with recommended fields for a spreadsheet setup. | fail→fail | 18,807 | 16,749 | -11% | 1 | 1 | 0% | 3,153 | 2,986 | -5% | 0 | 0 | — |
▸case-01 We are collecting feedback from various customer channels like support logs and quarterly reviews, but our current tagging is completely disorganized. I need a structured tagging schema covering customer demographics, journey stages, key feedback drivers, sentiment indicators, and business metadata. Please provide this taxonomy as a JSON schema so we can use it for automated tagging in our webhook pipeline. | fail→fail | 22,117 | 20,952 | -5% | 1 | 1 | 0% | 4,152 | 3,603 | -13% | 0 | 0 | — |
▸case-04 We are drafting our annual CSAT and NPS survey forms for enterprise B2B SaaS accounts. We need 5 survey questions designed to measure customer satisfaction, along with rating scale recommendations. Please write the survey questionnaire. | pass→fail | 15,649 | 13,353 | -15% | 1 | 1 | 0% | 2,423 | 2,428 | +0% | 0 | 0 | — |
▸case-05 We are launching a 30-day VoC listening tour to interview top 20 enterprise clients. Please write an interview script and calendar outreach plan for conducting these live interviews. | pass→fail | 20,064 | 23,715 | +18% | 1 | 1 | 0% | 2,997 | 3,675 | +23% | 0 | 0 | — |
▸case-06 We have compiled a dataset of 500 support ticket sentiment scores over Q1 and Q2. Please write a Python script using pandas and matplotlib to calculate quarterly mean sentiment scores and plot a line chart over time. | pass→pass | 16,922 | 12,272 | -27% | 1 | 1 | 0% | 3,093 | 2,318 | -25% | 0 | 0 | — |
▸case-07 We are setting up a CRM dropdown for customer journey stages to tag incoming feedback from sales and CS calls. We usually just use 'Lead', 'Active Customer', and 'Churned'. Please provide the complete set of customer lifecycle stages needed for feedback classification. | fail→fail | 14,667 | 13,032 | -11% | 1 | 1 | 0% | 2,169 | 2,090 | -4% | 0 | 0 | — |
▸case-08 When tagging feedback from user interviews, our team currently only records job title like 'Software Engineer'. What specific persona layer dimensions should be captured to properly evaluate feedback impact? | fail→pass | 15,828 | 11,102 | -30% | 1 | 1 | 0% | 2,467 | 1,747 | -29% | 0 | 0 | — |
▸case-09 Our product team built a taxonomy with 120 granular feedback drivers such as 'button-color-mismatch' and 'modal-close-delay'. Should we keep all 120 drivers or simplify? Explain the recommended limit for feedback drivers to ensure team adoption. | pass→pass | 13,872 | 10,752 | -22% | 1 | 1 | 0% | 2,076 | 1,820 | -12% | 0 | 0 | — |
▸case-10 We want to categorize customer complaints into broad driver categories. A team member proposed using 'Bug', 'Feature Request', 'UX', and 'Praise'. What are the standard high-level driver categories for feedback normalization? | fail→fail | 14,759 | 12,390 | -16% | 1 | 1 | 0% | 2,163 | 2,161 | -0% | 0 | 0 | — |
▸case-11 Our sentiment tagging currently only marks items as Positive, Negative, or Neutral with a High or Low urgency flag. What additional sentiment metrics should be included to evaluate data reliability across customer chatter? | fail→fail | 16,307 | 14,483 | -11% | 1 | 1 | 0% | 2,436 | 2,268 | -7% | 0 | 0 | — |
▸case-12 We need to capture business metadata alongside each tagged piece of feedback in our VoC warehouse. We currently log 'Customer Name' and 'Date'. What standard metadata fields should be attached to each feedback entry? | fail→fail | 17,111 | 12,901 | -25% | 1 | 1 | 0% | 2,648 | 2,414 | -9% | 0 | 0 | — |
▸case-13 We are planning to rename and reorganize several feedback driver categories in our central database. Should we overwrite existing historical database records with the new category names, or how should taxonomy changes be managed? | pass→pass | 15,987 | 14,968 | -6% | 1 | 1 | 0% | 2,288 | 2,414 | +6% | 0 | 0 | — |
▸case-14 How frequently should a customer feedback taxonomy be audited and refreshed to account for product evolution and tag drift, and what artifact should govern this process? | pass→pass | 15,249 | 13,135 | -14% | 1 | 1 | 0% | 2,227 | 2,198 | -1% | 0 | 0 | — |
▸case-15 We are receiving dozens of new feedback signals daily from live interviews and support channels. How should automated tagging of these incoming feedback signals be integrated into operational workflows? | fail→pass | 16,700 | 15,174 | -9% | 1 | 1 | 0% | 2,407 | 2,498 | +4% | 0 | 0 | — |
▸case-20 A teammate suggests defining customer persona solely by account revenue. Why is revenue alone insufficient for persona classification, and what persona attributes should be mapped? | fail→fail | 15,657 | 14,679 | -6% | 1 | 1 | 0% | 2,377 | 2,383 | +0% | 0 | 0 | — |
▸case-16 We are designing a spreadsheet template for customer success managers to tag raw survey comments manually. What structural features and validation mechanisms should be included in the spreadsheet template? | pass→pass | 19,048 | 17,754 | -7% | 1 | 1 | 0% | 2,810 | 2,935 | +4% | 0 | 0 | — |
▸case-17 We suspect our customer feedback dataset from last year has inconsistent tags because three different teams tagged records without clear guidelines. What framework layers should we use to audit and correct dataset drift? | fail→pass | 14,509 | 11,561 | -20% | 1 | 1 | 0% | 2,197 | 1,989 | -9% | 0 | 0 | — |
▸case-18 We just acquired a new company and their support team uses completely different tags than our CS team. What framework should we present to onboard the new team onto shared tagging standards? | fail→pass | 16,477 | 14,497 | -12% | 1 | 1 | 0% | 2,351 | 2,463 | +5% | 0 | 0 | — |
▸case-19 We are building an automated customer feedback ingestion pipeline using webhooks from customer support channels. What template format should be produced for webhook tag ingestion? | fail→pass | 13,101 | 11,010 | -16% | 1 | 1 | 0% | 2,148 | 2,327 | +8% | 0 | 0 | — |
▸case-21 Our feedback tagging lifecycle only tracks 'Pre-purchase' and 'Post-purchase'. Which specific post-onboarding lifecycle stages should be added to track customer progression? | pass→pass | 13,243 | 9,276 | -30% | 1 | 1 | 0% | 2,053 | 1,612 | -21% | 0 | 0 | — |
▸case-22 We categorized feedback drivers into 'bugs' and 'pricing'. However, enterprise customers frequently mention account executive responsiveness and business value ROI. How should these feedback dimensions be categorized in the driver layer? | pass→pass | 13,789 | 10,153 | -26% | 1 | 1 | 0% | 2,011 | 1,987 | -1% | 0 | 0 | — |
▸case-23 In our feedback database, a user comment 'The color scheme is terrible!' and a comment 'System downtime is causing data loss' are both marked as 'Negative'. How does the sentiment layer distinguish between sentiment intensity and operational priority? | pass→pass | 13,194 | 12,121 | -8% | 1 | 1 | 0% | 2,017 | 2,175 | +8% | 0 | 0 | — |