▸case-05 We evaluate LLM agent workflows using ToolBench, WebArena, GAIA, and SWE-bench. We need an audit of our evaluation coverage across the tool-use domain. Build a hierarchical capability taxonomy for agentic tool use and map our benchmark suite against it, highlighting density and missing spots. | fail→fail | 23,598 | 16,884 | -28% | 1 | 1 | 0% | 3,861 | 3,081 | -20% | 0 | 0 | — |
▸case-01 We are evaluating AI coding assistants across various tasks. We currently track performance on HumanEval, MBPP, SWE-bench, and RepoBench. Could you organize the code generation domain into a multi-level capability breakdown, align our benchmark set to those categories, and generate a report showing covered nodes, redundancy spots, and unassessed capabilities? | fail→pass | 20,467 | 21,863 | +7% | 1 | 1 | 0% | 3,500 | 4,033 | +15% | 0 | 0 | — |
▸case-02 Our research group evaluates multimodal LLMs using CLEVR, MMMU, ChartQA, and MathVista. Tempted to just list which benchmark tests chart reading versus visual math, but we need a formal taxonomy audit. Please build a hierarchical domain breakdown for visual reasoning, align our benchmark suite to it, and deliver a coverage report. | pass→fail | 31,367 | 19,785 | -37% | 1 | 1 | 0% | 4,798 | 3,630 | -24% | 0 | 0 | — |
▸case-03 We are auditing our NLP benchmark suite comprising SuperGLUE, MMLU, DROP, and RACE. Rather than a flat table of benchmark summaries, organize language understanding into a literature-grounded hierarchical structure, project our benchmark suite onto it, and report node coverage metrics. | fail→fail | 23,606 | 18,055 | -24% | 1 | 1 | 0% | 4,044 | 3,287 | -19% | 0 | 0 | — |
▸case-04 We want to evaluate LLM math capabilities using GSM8K, MATH, SVAMP, and TabMWP. The engineering team wants to just create a simple tags list, but we need a complete audit. Map these tools onto a hierarchical taxonomy for math reasoning, calculate density per node, and detail over-covered and under-covered areas. | fail→pass | 29,090 | 20,156 | -31% | 1 | 1 | 0% | 4,571 | 3,870 | -15% | 0 | 0 | — |
▸case-06 We track long-context LLM performance using NeedleInAHaystack, L-Eval, Ruler, and LongBench. Instead of a basic benchmark feature grid, construct a hierarchical capability tree for long-context reasoning, map our evaluation suite, and evaluate coverage density and white spaces. | fail→pass | 25,197 | 18,425 | -27% | 1 | 1 | 0% | 4,121 | 3,288 | -20% | 0 | 0 | — |
▸case-07 Our safety team uses AdvGLUE, HarmBench, StrongREJECT, and ToxiGen. A vendor suggested grouping them into three basic safety tags, but we need a systematic mapping. Build a hierarchical taxonomy of safety and alignment capabilities, map our benchmarks, and produce coverage statistics. | pass→pass | 24,058 | 17,827 | -26% | 1 | 1 | 0% | 3,970 | 3,420 | -14% | 0 | 0 | — |
▸case-08 We evaluate cyber security LLM capabilities using Cybench, InterCode-CTF, and SecQA. Construct a literature-backed hierarchical capability taxonomy for cybersecurity, map these benchmarks across the taxonomy nodes, and report coverage density along with over-tested clusters. | fail→fail | 36,343 | 17,838 | -51% | 1 | 1 | 0% | 6,192 | 3,393 | -45% | 0 | 0 | — |
▸case-09 Our medical AI group tracks performance using MedQA, MedMCQA, PubMedQA, and BioASQ. Map these benchmarks against a hierarchical taxonomy of medical AI capabilities and compute coverage statistics to show unassessed medical capabilities. | fail→fail | 17,717 | 19,191 | +8% | 1 | 1 | 0% | 3,038 | 3,554 | +17% | 0 | 0 | — |
▸case-10 We evaluate LLMs in finance using FinQA, ConvFinQA, TAT-QA, and Financial PhraseBank. Organize financial capability into a hierarchical taxonomy, map these benchmarks to the nodes, and detail coverage density and redundant focus areas. | fail→pass | 17,000 | 16,415 | -3% | 1 | 1 | 0% | 2,684 | 3,075 | +15% | 0 | 0 | — |
▸case-11 We assess speech LLMs using LibriSpeech, FLEURS, VoxCeleb, and AudioCaps. Build a multi-tier capability taxonomy for speech and audio processing, map these benchmarks, and output coverage statistics and gaps. | pass→fail | 20,774 | 20,450 | -2% | 1 | 1 | 0% | 3,423 | 3,965 | +16% | 0 | 0 | — |
▸case-12 We evaluate embodied AI models using ManipBench, RoboBench, and VirtualHome. Construct a hierarchical domain taxonomy for robotic manipulation and planning, map our benchmark suite, and output coverage density and white spaces. | fail→fail | 22,427 | 17,936 | -20% | 1 | 1 | 0% | 3,581 | 3,594 | +0% | 0 | 0 | — |
▸case-13 Our lab evaluates models on SciQ, GPQA, ChemBench, and ARC. Create a hierarchical capability breakdown for scientific reasoning, map these benchmark datasets, and summarize density statistics and over-indexed capabilities. | fail→pass | 20,419 | 15,561 | -24% | 1 | 1 | 0% | 3,275 | 2,823 | -14% | 0 | 0 | — |
▸case-14 We track legal LLMs using LegalGLUE, CUAD, LexGLUE, and LegalBench. Grouping them by simple target document types isn't sufficient. Construct a hierarchical taxonomy of legal capabilities, map the benchmarks, and provide coverage stats and white space analysis. | pass→pass | 27,019 | 17,380 | -36% | 1 | 1 | 0% | 4,325 | 3,045 | -30% | 0 | 0 | — |
▸case-15 We evaluate developer tools using HumanEval, SecurityEval, CyberSecEval, and APPS. Build a hierarchical capability taxonomy covering functional code synthesis and code security, project our benchmarks, and report coverage density and redundancy. | pass→pass | 20,057 | 13,807 | -31% | 1 | 1 | 0% | 3,403 | 2,740 | -19% | 0 | 0 | — |
▸case-16 We evaluate multilingual models using FLORES-200, MGSM, TyDi QA, and Belebele. Create a multi-tiered capability taxonomy for multilingual understanding and reasoning, map our benchmarks, and produce a coverage density and gap report. | fail→fail | 18,981 | 17,750 | -6% | 1 | 1 | 0% | 3,327 | 3,502 | +5% | 0 | 0 | — |
▸case-17 Our team evaluates LLMs on creative tasks using WritingPrompts, StoryCloze, FairytaleQA, and CreativeEval. Construct a literature-aligned hierarchical taxonomy of creative text generation capabilities, map these benchmarks, and detail node coverage statistics. | fail→fail | 31,221 | 20,428 | -35% | 1 | 1 | 0% | 6,188 | 4,048 | -35% | 0 | 0 | — |
▸case-18 We assess LLM reasoning with StrategyQA, MuSR, Big-Bench Hard, and PRM800K. Build a hierarchical capability taxonomy for reasoning under uncertainty, align our benchmark set to it, and compute coverage statistics and redundancy spots. | pass→pass | 23,046 | 23,206 | +1% | 1 | 1 | 0% | 3,928 | 4,453 | +13% | 0 | 0 | — |
▸case-19 We evaluate embodiment agents using RxR, REVERIE, Touchdown, and Matterport3D. Build a multi-tier capability taxonomy for vision-language navigation, align our benchmark suite, and output coverage density and white space analysis. | pass→pass | 29,779 | 18,065 | -39% | 1 | 1 | 0% | 5,007 | 3,548 | -29% | 0 | 0 | — |
▸case-20 We need to run local inference evaluations on Llama-3-8B using GSM8K and HumanEval. Could you write a python script using vLLM and lm-evaluation-harness to run these benchmark evaluations and output accuracy logs? | pass→pass | 16,798 | 17,010 | +1% | 1 | 1 | 0% | 3,332 | 3,743 | +12% | 0 | 0 | — |
▸case-21 We are building a new benchmark dataset called AgentBench-V2 to test LLM tool-calling reliability. What schema and data collection guidelines should we follow to curate 500 multi-turn interaction traces? | pass→pass | 24,666 | 19,907 | -19% | 1 | 1 | 0% | 4,332 | 3,600 | -17% | 0 | 0 | — |
▸case-22 We have raw accuracy scores for 5 models across MMLU, GSM8K, MATH, and HumanEval. How should we compute a normalized composite metric score to rank these models on an internal leaderboard? | pass→pass | 18,230 | 17,472 | -4% | 1 | 1 | 0% | 3,259 | 3,275 | +0% | 0 | 0 | — |