▸case-19 Reports are flooding into our customer queue. Instead of evaluating urgency, agents are handling reports in random chronological order. What specific function must occur immediately after intake? | pass→pass | 5,549 | 4,782 | -14% | 1 | 1 | 0% | 877 | 962 | +10% | 0 | 0 | — |
▸case-20 An agent issued a temporary suspension to a user for harassment and closed the ticket permanently. What final operational step is missing from the workflow lifecycle? | pass→pass | 5,527 | 6,197 | +12% | 1 | 1 | 0% | 793 | 1,220 | +54% | 0 | 0 | — |
▸case-01 We are building an in-house automated toxicity detection engine. Provide a PyTorch script to train a BERT model on the CivilComments dataset with multi-label classification loss. | pass→pass | 17,461 | 21,085 | +21% | 1 | 1 | 0% | 3,956 | 4,569 | +15% | 0 | 0 | — |
▸case-02 Draft a legally binding GDPR and CCPA privacy policy cover page detailing biometric data retention schedules and third-party data broker processing terms for our user registry. | pass→pass | 17,881 | 16,097 | -10% | 1 | 1 | 0% | 3,050 | 3,050 | 0% | 0 | 0 | — |
▸case-03 Our community platform server is experiencing a Layer 7 HTTP flood attack. Provide cloudflare rate-limiting rules and iptables command scripts to scrub incoming abusive traffic immediately. | pass→pass | 14,760 | 14,898 | +1% | 1 | 1 | 0% | 2,782 | 3,069 | +10% | 0 | 0 | — |
▸case-04 We are designing a content moderation framework for a global gaming community. Out of convenience, our team wants to publish a single flat list of forbidden words. How should our policy framework be structured to handle varying severity levels and localized rules? | pass→pass | 17,237 | 17,464 | +1% | 1 | 1 | 0% | 2,689 | 3,018 | +12% | 0 | 0 | — |
▸case-05 Our community manager wants to process incoming content reports by having moderators immediately post publicly on reported threads. How should an incident intake lifecycle be structured step-by-step from arrival to resolution? | pass→pass | 15,669 | 15,731 | +0% | 1 | 1 | 0% | 2,555 | 2,932 | +15% | 0 | 0 | — |
▸case-06 When handling user abuse tickets, our team simply drops tickets into an open channel for anyone to grab, but tickets often sit unresolved for days. What explicit parameters must be attached to the workflow lifecycle to ensure accountability? | pass→pass | 13,870 | 14,393 | +4% | 1 | 1 | 0% | 2,285 | 2,628 | +15% | 0 | 0 | — |
▸case-07 We are setting up our moderation software stack. Engineers suggested relying exclusively on automated regex filters without human oversight software. What core operational tooling components should be implemented for the moderation team? | fail→pass | 16,778 | 13,611 | -19% | 1 | 1 | 0% | 2,509 | 2,352 | -6% | 0 | 0 | — |
▸case-08 A user on our forum posted credible threats of real-world violence alongside severe copyright violations. Our frontline moderators want to handle this entirely within their standard warning queue. Who should be included on the escalation ladder? | fail→pass | 10,983 | 8,355 | -24% | 1 | 1 | 0% | 1,779 | 1,582 | -11% | 0 | 0 | — |
▸case-09 After resolving a major safety incident on our platform, the product team wants to close the ticket and immediately move on without further documentation. What post-incident review practices should be conducted? | fail→fail | 13,295 | 14,625 | +10% | 1 | 1 | 0% | 2,235 | 2,551 | +14% | 0 | 0 | — |
▸case-10 We need to onboard ten new community moderators this week. Rather than giving them arbitrary authority, what specific guidance assets should be included in the moderator operational guide? | fail→pass | 14,606 | 11,793 | -19% | 1 | 1 | 0% | 2,242 | 2,176 | -3% | 0 | 0 | — |
▸case-21 Our team relies on manual review for 100% of incoming spam comments, leading to a 48-hour delay. What system capability should be added to the tooling suite to filter obvious spam instantly? | pass→pass | 9,671 | 10,156 | +5% | 1 | 1 | 0% | 1,630 | 1,845 | +13% | 0 | 0 | — |
▸case-11 Our operations lead tracks moderation effectiveness solely by counting total complaints received. What metrics and mechanisms should be incorporated into the incident log and tracking dashboard? | pass→pass | 15,211 | 15,621 | +3% | 1 | 1 | 0% | 2,639 | 2,863 | +8% | 0 | 0 | — |
▸case-12 During a ongoing high-profile security and harassment breach on our platform, our team plans to post internal slack messages only. What communication outline should be prepared during a crisis? | pass→pass | 13,883 | 15,584 | +12% | 1 | 1 | 0% | 2,251 | 2,747 | +22% | 0 | 0 | — |
▸case-13 Our legal team suggests keeping all community enforcement rules completely secret so users cannot game the system. How should community guidelines and expectations be presented to members? | pass→pass | 15,519 | 12,222 | -21% | 1 | 1 | 0% | 2,355 | 2,159 | -8% | 0 | 0 | — |
▸case-14 Our moderation staff works 12-hour shifts reviewing high-severity abusive content continuously. What shift management strategy should be deployed to mitigate fatigue and psychological impact? | pass→pass | 16,850 | 13,328 | -21% | 1 | 1 | 0% | 2,822 | 2,410 | -15% | 0 | 0 | — |
▸case-22 We are launching a community ambassador program where volunteer leaders handle local sub-forums. How should safety policies be communicated to these non-employee representatives? | fail→fail | 14,822 | 14,201 | -4% | 1 | 1 | 0% | 2,334 | 2,473 | +6% | 0 | 0 | — |
▸case-15 Frontline moderators report feeling overwhelmed after handling severe hate speech queues all day. Beyond shift adjustments, what dedicated support must management provide? | pass→pass | 13,905 | 12,433 | -11% | 1 | 1 | 0% | 2,115 | 2,201 | +4% | 0 | 0 | — |
▸case-16 We are hosting a massive real-time launch event with dynamic live audio channels and chat. How should safety coverage be coordinated during live real-time events? | fail→fail | 16,177 | 19,134 | +18% | 1 | 1 | 0% | 2,423 | 3,256 | +34% | 0 | 0 | — |
▸case-17 A community administrator wants to apply a permanent IP ban to every single offense, regardless of severity. How should penalties be structured within a policy matrix? | fail→pass | 11,698 | 15,292 | +31% | 1 | 1 | 0% | 1,896 | 2,625 | +38% | 0 | 0 | — |
▸case-18 When deploying our user content policies across Europe, East Asia, and North America, operations wants to enforce identical strict local speech bans globally. What dimension must the policy matrix incorporate? | pass→pass | 10,181 | 7,069 | -31% | 1 | 1 | 0% | 1,530 | 1,360 | -11% | 0 | 0 | — |