▸case-01 Can you run a construct validity assessment on the benchmark 'LogicBench-v2'? The authors state that it evaluates 'multi-step logical reasoning in pure text', but I want to make sure it isn't just testing memorization or syntax matching. Here are three representative task items: [Item 1: Syllogism grid deduction; Item 2: Variable binding text puzzle; Item 3: Constraint satisfaction problem]. Please analyze these inputs and return an overall validity verdict along with supporting evidence across the key validity dimensions. | fail→pass | 14,716 | 13,108 | -11% | 1 | 1 | 0% | 2,186 | 2,299 | +5% | 0 | 0 | — |
▸case-02 I need a psychometric validity evaluation for an AI dataset called 'DevEval-Lite'. The claimed capability is 'autonomous software engineering bug-fixing ability'. Below are some example tasks from the suite: 1) Resolving a Python KeyError by editing a 5-line script, 2) Modifying a Dockerfile syntax error, 3) Writing a mock function. Please review whether these tasks accurately represent the target capability and provide a structured validity verdict accompanied by evidence for each analysis dimension. | fail→pass | 15,947 | 10,693 | -33% | 1 | 1 | 0% | 2,442 | 2,012 | -18% | 0 | 0 | — |
▸case-03 Evaluate whether 'MathGrid-100' accurately measures 'advanced mathematical reasoning'. Task examples: 1) Calculating basic arithmetic modulo 7, 2) Solving a linear equation system, 3) Re-formatting LaTeX matrix representations. Analyze if these tasks measure the claimed capability, and output a structured verdict with dimension-by-dimension findings. | pass→pass | 12,418 | 9,135 | -26% | 1 | 1 | 0% | 1,866 | 1,764 | -5% | 0 | 0 | — |
▸case-04 Review 'PyBench-Core' claiming to measure 'idiomatic Python code generation'. Task examples: 1) Writing a list comprehension for filtering integers, 2) Implementing a quicksort algorithm from scratch, 3) Writing standard docstrings for empty functions. Assess whether this suite truly tests idiomatic coding and produce a judgment report covering content, convergent, and discriminant factors. | pass→fail | 18,303 | 15,489 | -15% | 1 | 1 | 0% | 2,876 | 2,856 | -1% | 0 | 0 | — |
▸case-05 Examine 'MedQA-Lite' which claims to evaluate 'clinical diagnostic reasoning in internal medicine'. Example questions: 1) Multiple-choice question on rare genetic mutation acronyms, 2) True/false on drug dosage limits, 3) Identifying disease diagnosis from a 3-paragraph patient case history. Provide a structured validity analysis and verdict across psychometric validity dimensions. | fail→pass | 15,901 | 12,890 | -19% | 1 | 1 | 0% | 2,490 | 2,223 | -11% | 0 | 0 | — |
▸case-06 Assess 'ToolBench-Exec' designed to test 'autonomous API tool orchestration'. Examples: 1) Looking up weather via a single search tool call, 2) Parsing JSON output from a web request, 3) Executing a 4-step sequence of dependent REST calls to plan a trip. Analyze the psychometric validity and give an overall verdict backed by evidence. | fail→pass | 17,561 | 14,074 | -20% | 1 | 1 | 0% | 2,648 | 2,471 | -7% | 0 | 0 | — |
▸case-07 Analyze 'FinSum-Eval' claiming to measure 'financial earnings report summarization quality'. Example items: 1) Extracting EBITDA figures from quarterly PDF tables, 2) Summarizing CEO risk commentary in 3 sentences, 3) Correcting spelling in financial footnotes. Review if this benchmark effectively isolates summarization skill and report a full validity evaluation. | pass→pass | 18,859 | 13,575 | -28% | 1 | 1 | 0% | 2,871 | 2,450 | -15% | 0 | 0 | — |
▸case-08 Review 'VQA-Grounding-v1' claiming to evaluate 'spatial visual-language grounding in complex scenes'. Examples: 1) Answering 'What color is the car?' in a clear photo, 2) Identifying bounding boxes for overlapping objects, 3) Reading text on a street sign. Analyze psychometric validity and summarize findings into an explicit verdict. | fail→pass | 13,133 | 12,910 | -2% | 1 | 1 | 0% | 2,088 | 2,252 | +8% | 0 | 0 | — |
▸case-09 Assess 'HopQA-2024' which claims to test 'multi-hop information retrieval and synthesis across documents'. Task samples: 1) Finding a birthdate from Wikipedia, 2) Combining facts across 3 articles to identify a politician's advisor, 3) Detecting typos in prompt text. Provide an evaluation report covering validity dimensions and a concluding verdict. | fail→pass | 16,433 | 12,362 | -25% | 1 | 1 | 0% | 2,670 | 2,322 | -13% | 0 | 0 | — |
▸case-10 Evaluate 'SecCTF-Bench' claiming to measure 'ethical hacking and vulnerability exploitation skills'. Examples: 1) Decoding a base64 string, 2) Finding a SQL injection vulnerability in a web app, 3) Answering multiple-choice history of computer viruses. Review whether this dataset validly measures exploitation capabilities and supply a structured verdict. | pass→pass | 14,461 | 12,056 | -17% | 1 | 1 | 0% | 2,422 | 2,131 | -12% | 0 | 0 | — |
▸case-11 Examine 'EmpathyEval-v3' claiming to measure 'empathic emotional support in conversational AI'. Tasks: 1) Outputting canned supportive phrases like 'I am sorry to hear that', 2) Classifying user sentiment polarity (positive/negative), 3) Generating personalized advice for a grieving user. Evaluate whether the benchmark validly tests empathic support and report your dimension breakdown and verdict. | fail→pass | 14,821 | 12,939 | -13% | 1 | 1 | 0% | 2,275 | 2,300 | +1% | 0 | 0 | — |
▸case-12 Evaluate 'StrictFollow-v1' claiming to test 'complex multi-constraint instruction following'. Examples: 1) Writing a 500-word essay with no letter 'e', 2) Generating JSON with 3 specific key names, 3) Translating a short sentence to French. Determine if these tasks validly measure multi-constraint adherence and provide a verdict with evidence per validity step. | fail→pass | 16,061 | 12,451 | -22% | 1 | 1 | 0% | 2,417 | 2,285 | -5% | 0 | 0 | — |
▸case-13 Assess 'LegalClause-Bench' claiming to evaluate 'legal contract anomaly detection and risk assessment'. Examples: 1) Flagging an unusual indemnification clause, 2) Identifying grammar mistakes in contract preambles, 3) Standardizing font styles across sections. Provide a validity assessment report and explicit verdict for this benchmark. | pass→pass | 12,203 | 12,051 | -1% | 1 | 1 | 0% | 1,848 | 2,157 | +17% | 0 | 0 | — |
▸case-14 Review 'BioReason-Eval' claiming to evaluate 'scientific hypothesis formulation from biomedical literature'. Task examples: 1) Extracting author affiliations from a PubMed paper, 2) Proposing a novel gene target based on 3 research abstracts, 3) Converting PDF tables to CSV. Deliver a structured validity evaluation and verdict. | fail→pass | 12,763 | 11,222 | -12% | 1 | 1 | 0% | 2,088 | 2,071 | -1% | 0 | 0 | — |
▸case-15 Examine 'PlanBot-Bench' designed to test 'hierarchical task planning for indoor mobile manipulators'. Task examples: 1) Outputting a list of waypoints to move from kitchen to living room, 2) Generating a 10-step pick-and-place sequence with precondition checks, 3) Spelling out color names of household objects. Evaluate its psychometric validity and output a verdict. | pass→pass | 14,984 | 12,749 | -15% | 1 | 1 | 0% | 2,388 | 2,255 | -6% | 0 | 0 | — |
▸case-16 Analyze 'StoryCraft-Bench' claiming to evaluate 'long-form narrative cohesion and plot development'. Examples: 1) Writing a 3-paragraph rhyming poem about autumn, 2) Outline a 5-chapter novel with character arcs and plot twists, 3) Correcting punctuation in a paragraph. Deliver a structured validity analysis with an overall verdict. | fail→fail | 14,944 | 12,248 | -18% | 1 | 1 | 0% | 2,375 | 2,187 | -8% | 0 | 0 | — |
▸case-17 Evaluate 'PolyGlot-Eval' claiming to measure 'nuanced literary translation across low-resource languages'. Examples: 1) Translating idiomatic proverbs from Swahili to English preserving metaphor, 2) Looking up dictionary definitions of isolated words, 3) Counting word frequencies in source text. Provide a validity report covering dimensions and overall verdict. | pass→pass | 14,009 | 10,946 | -22% | 1 | 1 | 0% | 2,209 | 1,907 | -14% | 0 | 0 | — |
▸case-18 Assess 'DBQuery-Bench' claiming to test 'complex relational SQL query synthesis with multi-table joins'. Examples: 1) Writing SELECT * FROM users;, 2) Constructing a query with 4 joins, nested aggregations, and window functions, 3) Identifying database connection strings in config files. Analyze the psychometric validity of this evaluation suite and give a verdict. | pass→pass | 13,592 | 12,208 | -10% | 1 | 1 | 0% | 2,145 | 2,187 | +2% | 0 | 0 | — |
▸case-19 Review 'ChartLogic-v2' claiming to measure 'quantitative reasoning over multi-modal statistical charts'. Examples: 1) Reading the title text of a bar chart, 2) Estimating a 5-year trend line growth rate from a scatter plot with dual axes, 3) Identifying image file dimensions (height/width). Supply a structured validity breakdown and explicit verdict. | pass→pass | 12,748 | 12,466 | -2% | 1 | 1 | 0% | 1,996 | 2,077 | +4% | 0 | 0 | — |
▸case-20 I need a Python script to scan our training corpus against 'GSM8K' test questions using 13-gram overlap matching to detect test-set contamination. Can you write the script to calculate the overlap percentage for each file? | pass→pass | 36,695 | 15,655 | -57% | 1 | 1 | 0% | 2,262 | 3,388 | +50% | 0 | 0 | — |
▸case-21 Given the target standard code string and 5 generated candidate outputs, write a function to calculate Pass@K accuracy and BLEU scores for the generated code candidates. | pass→pass | 16,730 | 15,937 | -5% | 1 | 1 | 0% | 3,038 | 3,318 | +9% | 0 | 0 | — |
▸case-22 I want to benchmark the inference latency and token throughput (tokens per second) of Llama-3-8B vs Mistral-7B on a V100 GPU across batch sizes 1, 4, and 16. How should I set up the benchmarking harness and timer? | pass→pass | 20,217 | 19,544 | -3% | 1 | 1 | 0% | 3,832 | 3,879 | +1% | 0 | 0 | — |
▸case-23 We are hiring human evaluators on Mechanical Turk to grade chatbot answers. Please draft the guidelines for worker task layout, qualification test questions, and pay rate calculations. | pass→pass | 17,557 | 16,729 | -5% | 1 | 1 | 0% | 2,694 | 3,267 | +21% | 0 | 0 | — |