▸case-06 We have 12 candidate IT infrastructure projects with estimated costs and expected ROI scores. Our capital budget cap is strictly $10 million. Selecting the top-ranked individual project is insufficient because we need to find the combination of projects that maximizes cumulative ROI without exceeding the $10M budget limit. Please solve this as a 0-1 knapsack capital budgeting optimization. | fail→pass | 12,923 | 30,155 | +133% | 1 | 1 | 0% | 2,602 | 6,848 | +163% | 0 | 0 | — |
▸case-01 Our municipal planning team needs to prioritize eight regional transit expansion proposals using multi-criteria decision analysis. We have gathered data across six evaluation dimensions such as capital expenditure, projected ridership, carbon footprint reduction, and social equity. Please generate a full, ordered league table listing all options with their calculated utility or net preference scores. We also require a comparative evaluation using at least two distinct MCDA techniques, showing their pair-wise rank correlation and highlighting any options that shift position, along with commentary on how sensitive the final order is to adjustments in criterion weights. | fail→fail | 27,309 | 29,888 | +9% | 1 | 1 | 0% | 6,250 | 6,882 | +10% | 0 | 0 | — |
▸case-02 We are evaluating six enterprise cybersecurity platforms for procurement. Our team has scored each vendor across key metrics including cost, deployment speed, compliance coverage, and threat detection accuracy. We need an exhaustive priority ranking of all candidates from top to bottom. Present the result as an ordered table containing each candidate's overall score and status notes, followed by a breakdown comparing how different decision-making models agree (including rank correlation coefficients and divergent candidates), and end with notes regarding weight sensitivity. | fail→fail | 24,043 | 24,585 | +2% | 1 | 1 | 0% | 4,793 | 6,079 | +27% | 0 | 0 | — |
▸case-03 A portfolio review committee wants a comprehensive priority list for seven offshore wind investment targets based on financial yield, environmental compliance, risk profile, and grid readiness. Could you construct a complete ranked hierarchy of all seven alternatives? Please include their relative net preference values, a side-by-side consistency check across different ranking methodologies displaying Kendall rank correlation values, and an analysis of how robust the top choices remain when criterion weights are varied. | fail→fail | 26,097 | 25,609 | -2% | 1 | 1 | 0% | 6,225 | 6,856 | +10% | 0 | 0 | — |
▸case-04 We need to filter 15 candidate pharmaceutical chemical suppliers against strict mandatory regulatory standards: ISO 13485 certification, maximum lead time of 14 days, and zero safety non-conformances in 2 years. Any vendor failing even one criterion must be eliminated. Please perform a binary pass/fail gating analysis and output the list of qualified versus disqualified vendors. | pass→pass | 11,277 | 21,292 | +89% | 1 | 1 | 0% | 2,271 | 5,233 | +130% | 0 | 0 | — |
▸case-05 Our engineering committee needs to derive a normalized weight vector for five key design criteria (cost, durability, weight, manufacturability, and safety) using pairwise comparisons. We are not ranking any physical product options today; we only require the criteria relative importance vector and consistency ratio derived via Analytic Hierarchy Process. | pass→pass | 20,232 | 26,792 | +32% | 1 | 1 | 0% | 4,601 | 6,622 | +44% | 0 | 0 | — |
▸case-07 A regional health network wants to rank five MRI scanner models (Models A, B, C, D, E) across capital expense, image resolution, maintenance SLA, and power consumption. Management is tempted to just pick the highest resolution machine, but we need a complete league table evaluating all five units using multi-criteria ranking, including method consistency check and weight sensitivity. | fail→fail | 25,582 | 27,025 | +6% | 1 | 1 | 0% | 5,155 | 6,845 | +33% | 0 | 0 | — |
▸case-08 An automotive manufacturer is evaluating four tier-1 sub-assembly suppliers (Vanguard, Apex, GlobalCraft, PrecisionTech) across cost, defect rate, delivery reliability, and ESG compliance. The procurement team suggests simply picking the cheapest option, but leadership demands a full multi-criteria ranking using PROMETHEE II and MAVT to evaluate full preference flows and utility values, complete with Kendall tau agreement. | fail→fail | 25,081 | 25,080 | -0% | 1 | 1 | 0% | 6,223 | 6,855 | +10% | 0 | 0 | — |
▸case-09 A financial services firm is prioritizing five cloud hosting providers (AWS, Azure, GCP, OCI, IBM Cloud) for core banking workload migration based on latency, uptime SLA, compliance certifications, and egress costs. The infrastructure team wants to publish a complete priority ranking table along with a comparison of results between MAVT and ELECTRE III methods, specifically identifying any pairs of providers that cannot be strictly ordered. | fail→fail | 26,326 | 26,928 | +2% | 1 | 1 | 0% | 5,326 | 6,853 | +29% | 0 | 0 | — |
▸case-10 A tech firm is selecting among six potential data center sites (Ashburn, Dublin, Frankfurt, Singapore, Tokyo, Santa Clara) evaluating land cost, power grid green-ratio, fiber density, and seismic risk. The site selection team wants a total ordering league table, accompanied by a multi-method agreement table showing Kendall tau values. | fail→fail | 22,700 | 24,537 | +8% | 1 | 1 | 0% | 4,514 | 6,840 | +52% | 0 | 0 | — |
▸case-11 A team lead wants to rank four frontend web frameworks (React, Vue, Angular, Svelte) for a new enterprise dashboard, scoring them on developer velocity, ecosystem breadth, bundle size, and community support. The team is tempted to just use a raw unweighted point tally, but we need a formal multi-criteria synthesis combining weights, normalizing scores, comparing PROMETHEE II with MAVT, and testing weight sensitivity. | fail→fail | 25,827 | 26,962 | +4% | 1 | 1 | 0% | 6,228 | 6,859 | +10% | 0 | 0 | — |
▸case-12 An urban transit authority is evaluating four bus fleet modernization tenders (Tender Alpha, Beta, Gamma, Delta) across lifecycle emissions, purchase price, charging speed, and passenger capacity using ELECTRE III. Because ELECTRE III produces partial orders, explain how non-dominated or incomparable tenders are represented when building a comprehensive ranking table. | fail→fail | 16,013 | 18,363 | +15% | 1 | 1 | 0% | 2,941 | 3,815 | +30% | 0 | 0 | — |
▸case-13 A logistics company wants to rank five freight routes (Routes R1 to R5) based on transit time, toll cost, fuel efficiency, and border clearance reliability. We want to triangulate candidate priorities using both additive utility (MAVT) and net outranking flows (PROMETHEE II), documenting any shifts in route positioning between the two methods. | fail→fail | 25,417 | 25,010 | -2% | 1 | 1 | 0% | 6,212 | 6,843 | +10% | 0 | 0 | — |
▸case-14 A venture capital firm is ranking five AI startups (Alpha, Beta, Gamma, Delta, Epsilon) for Series A funding across market size, team experience, IP strength, and revenue growth. The team has generated a baseline PROMETHEE II ranking but needs to document how sensitive the top candidate's rank is if the weight on revenue growth is adjusted by +/- 20%. | pass→fail | 20,464 | 24,418 | +19% | 1 | 1 | 0% | 4,042 | 6,851 | +69% | 0 | 0 | — |
▸case-15 A real estate investment trust needs to rank six commercial office properties (Properties 101 through 106) based on net yield, tenant occupancy rate, LEED certification, and location walkability score. Please generate a complete league table with net preference scores, compare PROMETHEE II and MAVT using Kendall tau, and evaluate weight sensitivity. | fail→fail | 24,787 | 25,364 | +2% | 1 | 1 | 0% | 6,213 | 6,845 | +10% | 0 | 0 | — |
▸case-16 A delivery corporation needs to prioritize five vehicle depot sites (Depots A, B, C, D, E) for EV charging infrastructure installation. Evaluation criteria include grid connectivity cost, fleet density, local solar potential, and lease duration. The management wants a complete ordinal ranking, multi-method consistency metrics, and robustness analysis. | fail→fail | 27,171 | 24,652 | -9% | 1 | 1 | 0% | 6,206 | 6,838 | +10% | 0 | 0 | — |
▸case-17 A pharmaceutical company is ranking seven R&D drug discovery projects (Projects P1 through P7) across clinical trial success probability, market potential, development cost, and strategic alignment. Rather than picking only the top project, management requires a total priority ordering of all seven projects using multi-criteria triangulation. | fail→fail | 28,449 | 25,273 | -11% | 1 | 1 | 0% | 5,828 | 6,833 | +17% | 0 | 0 | — |
▸case-18 A retail giant is evaluating five distribution center candidate cities (Chicago, Dallas, Atlanta, Columbus, Reno) across labor cost, highway access, rail connectivity, and property tax incentives. Build a full decision synthesis ranking all candidates, comparing MAVT and PROMETHEE II rankings, and testing weight stability. | fail→fail | 26,572 | 25,573 | -4% | 1 | 1 | 0% | 6,203 | 6,835 | +10% | 0 | 0 | — |
▸case-19 A university consortium is ranking four Learning Management System platforms (Canvas, Blackboard, Moodle, Brightspace) based on accessibility standards compliance, integration APIs, annual licensing cost, and mobile user experience. The evaluation committee needs a complete multi-criteria ranking table and consistency correlation. | fail→pass | 19,841 | 27,356 | +38% | 1 | 1 | 0% | 3,946 | 6,827 | +73% | 0 | 0 | — |
▸case-20 A utility provider is evaluating five grid energy storage technologies (Lithium-ion, Vanadium Redox Flow, Compressed Air, Sodium-Sulfur, Flywheel) across round-trip efficiency, capital cost per kWh, calendar lifespan, and environmental footprint. Produce a complete MCDA ranking with method agreement analysis and weight sensitivity notes. | fail→fail | 26,398 | 25,213 | -4% | 1 | 1 | 0% | 6,206 | 6,837 | +10% | 0 | 0 | — |
▸case-21 A municipal water board is ranking six water purification technology proposals (Proposals W1 to W6) across land footprint, chemical consumption rate, capital expenditure, operational reliability, and energy efficiency. Deliver a full league table, a method comparison table, and sensitivity notes. | fail→fail | 24,586 | 27,984 | +14% | 1 | 1 | 0% | 4,599 | 6,826 | +48% | 0 | 0 | — |
▸case-22 A telecom operator needs to rank four 5G network equipment vendors (Vendor Alpha, Beta, Gamma, Delta) across throughput performance, RAN hardware cost, power efficiency, and vendor support SLA. Generate a full ranking output with method comparison and weight sensitivity analysis. | fail→fail | 26,611 | 26,436 | -1% | 1 | 1 | 0% | 6,192 | 6,825 | +10% | 0 | 0 | — |