▸case-01 We are benchmarking a critical cryptography function in Go (`CryptoHash`) using `go test`. Developers want to force each benchmark function to run for exactly 10 seconds of active measurement instead of the default 1 second, while ensuring test suites do not run benchmarks in parallel across CPU cores. Specify the exact `go test` flags required. | pass→pass | 6,192 | 4,034 | -35% | 1 | 1 | 0% | 1,115 | 697 | -37% | 0 | 0 | — |
▸case-02 When benchmarking a CLI binary `data-processor` using Hyperfine, an engineer suggests using standard `--runs 100`. However, short runs cause cold-cache bias and variable system load causes unstable timing. Provide the Hyperfine command line arguments that enforce a strict 3-run warmup period followed by a strict minimum benchmark duration of 15 seconds. | fail→pass | 3,194 | 2,526 | -21% | 1 | 1 | 0% | 605 | 465 | -23% | 0 | 0 | — |
▸case-03 In a Python service test suite using `pytest-benchmark`, team members usually run benchmarks with default parameters, which stop after 1 second or 100 iterations. For long-running linear algebra tasks, 1 second is insufficient for thermal stabilization. Provide the pytest CLI flag that sets the minimum benchmark evaluation time per test to 12 seconds. | pass→pass | 2,798 | 3,288 | +18% | 1 | 1 | 0% | 507 | 529 | +4% | 0 | 0 | — |
▸case-04 When running performance benchmarks with `pytest-benchmark` on memory-intensive data structures, auto-calibration adds variable iteration overhead that distorts fixed-duration runs. Show the pytest command line flag used to explicitly turn off auto-calibration so that duration rules are strictly respected. | fail→fail | 19,561 | 10,512 | -46% | 1 | 1 | 0% | 2,878 | 2,053 | -29% | 0 | 0 | — |
▸case-05 A DevOps engineer wants to run three CLI benchmark commands (`cmd1`, `cmd2`, `cmd3`) simultaneously using shell backgrounding (`&`) in Hyperfine to finish testing faster. Explain why backgrounding invalidates performance results and state the correct Hyperfine command invocation parameter syntax to run them sequentially one after another. | pass→pass | 7,974 | 7,291 | -9% | 1 | 1 | 0% | 1,322 | 1,230 | -7% | 0 | 0 | — |
▸case-06 An API benchmark in k6 is configured with `vus: 50` and `iterations: 1000`. This causes the benchmark duration to vary wildly depending on backend latency spikes. Provide the JS scenario configuration object for k6 that switches the test execution to a fixed duration of 30 seconds using the constant-arrival-rate executor. | pass→pass | 7,581 | 4,956 | -35% | 1 | 1 | 0% | 1,438 | 975 | -32% | 0 | 0 | — |
▸case-07 In a Rust library using `criterion.rs`, a benchmark function `bench_parser` uses the default 5-second measurement time. For a high-variance JSON parser, the team needs to override the benchmark duration to 20 seconds and the warmup period to 5 seconds. Provide the Rust code snippet configuring the `Criterion` struct instance. | pass→pass | 7,666 | 6,605 | -14% | 1 | 1 | 0% | 1,609 | 1,398 | -13% | 0 | 0 | — |
▸case-08 A tester wants to run `ab` against an HTTP endpoint `http://localhost:8080/api` using `-n 50000` (50,000 total requests) with concurrency 1. However, if the server slows down, the test hangs indefinitely. Provide the `ab` command flag that sets a strict maximum run duration of 60 seconds regardless of request count. | pass→pass | 4,674 | 3,782 | -19% | 1 | 1 | 0% | 632 | 651 | +3% | 0 | 0 | — |
▸case-09 For performance regression testing of a microservice, an engineer writes a Locust load test script (`locustfile.py`). Instead of manually stopping the web UI after an arbitrary time, they need a CLI command to run Locust headlessly with 1 user sequentially sending requests for a strict duration of 5 minutes before auto-terminating. Provide the CLI flags. | pass→pass | 4,036 | 3,907 | -3% | 1 | 1 | 0% | 740 | 712 | -4% | 0 | 0 | — |
▸case-10 A developer attempts to run `pytest-benchmark` with `pytest -n 4` using pytest-xdist to execute benchmark cases across 4 parallel worker processes simultaneously. Explain why this breaks benchmark validity and state how to configure pytest execution so benchmarks run sequentially in a single process. | pass→fail | 11,759 | 8,188 | -30% | 1 | 1 | 0% | 1,895 | 1,343 | -29% | 0 | 0 | — |
▸case-11 When executing `go test -bench=.`, an engineer wants each benchmark function to execute 5 independent sequential rounds, with each individual round running for exactly 8 seconds. Provide the `go test` command with the appropriate flags. | pass→pass | 3,676 | 4,138 | +13% | 1 | 1 | 0% | 596 | 658 | +10% | 0 | 0 | — |
▸case-12 When benchmarking disk cleanup operations sequentially using Hyperfine on two target directories (`/tmp/dirA` and `/tmp/dirB`), running the benchmark repeatedly without resetting state distorts subsequent runs because the files are missing. Provide the Hyperfine option used to execute a setup/reset script sequentially before each benchmark run. | pass→pass | 4,342 | 4,200 | -3% | 1 | 1 | 0% | 689 | 750 | +9% | 0 | 0 | — |
▸case-13 An engineer wants to benchmark a Node.js Express route using `autocannon`. By default, autocannon runs for 10 seconds. The team requires a strict duration of 45 seconds while keeping connections set to 1 to maintain sequential request execution. Provide the autocannon CLI command flags for target `http://localhost:3000/health`. | pass→pass | 3,088 | 2,303 | -25% | 1 | 1 | 0% | 562 | 414 | -26% | 0 | 0 | — |
▸case-14 In a JavaScript environment using `Benchmark.js`, a developer creates a test suite. The default suite auto-calculates sample cycles, but for a volatile memory allocation test, the developer needs to set a strict minimum time sample duration of 10 seconds per benchmark task (`maxTime`). Provide the option object property passed to `Suite.add()`. | pass→pass | 4,856 | 3,944 | -19% | 1 | 1 | 0% | 880 | 735 | -16% | 0 | 0 | — |
▸case-15 To detect performance regressions in a Python library, an engineer needs to run benchmarks sequentially for a strict minimum duration of 5 seconds, save the results under the name `v1_baseline`, and ensure no other pytest processes run concurrently. Provide the full pytest command. | pass→pass | 15,631 | 6,265 | -60% | 1 | 1 | 0% | 2,401 | 1,067 | -56% | 0 | 0 | — |
▸case-16 In an enterprise load testing pipeline using Apache JMeter, a team needs to run a load test plan (`test_plan.jmx`) in CLI non-GUI mode with a strict execution duration limit of 180 seconds to avoid runaway build jobs. Provide the JMeter command line options. | pass→pass | 8,765 | 8,993 | +3% | 1 | 1 | 0% | 1,556 | 1,553 | -0% | 0 | 0 | — |
▸case-17 An engineer is benchmarking a compression tool with parameter `level` taking values `1`, `5`, and `9`. To keep the sweep runtime predictable, every parameter combination must be evaluated for at least 10 seconds total benchmark time. Provide the Hyperfine command using parameter sweeps. | fail→fail | 5,735 | 5,485 | -4% | 1 | 1 | 0% | 1,065 | 950 | -11% | 0 | 0 | — |
▸case-18 In Go benchmark code, a developer writes `b.RunParallel(func(pb *testing.PB) { ... })`. Explain why `RunParallel` should be avoided when strict sequential CPU timing is required, and show how the loop should be structured instead. | pass→pass | 10,346 | 10,215 | -1% | 1 | 1 | 0% | 1,867 | 1,784 | -4% | 0 | 0 | — |
▸case-19 When running a long sequential benchmark suite containing 20 heavy CPU benchmarks, earlier benchmarks run at peak clock speeds while later benchmarks suffer from thermal throttling, corrupting relative measurements. What operational option in Hyperfine inserts an explicit delay between benchmark runs to allow thermal cooling? | pass→pass | 10,635 | 13,483 | +27% | 1 | 1 | 0% | 1,908 | 2,391 | +25% | 0 | 0 | — |
▸case-20 A backend developer needs to record a 30-second heap memory allocation profile for a Go HTTP microservice running at `http://localhost:6060/debug/pprof/heap`. Provide the `go tool pprof` command line invocation to download and inspect the heap profile. | pass→pass | 10,881 | 9,989 | -8% | 1 | 1 | 0% | 2,144 | 1,945 | -9% | 0 | 0 | — |
▸case-21 A SRE team needs to execute an extreme concurrent stress test simulating 10,000 simultaneous active Virtual Users (VUs) ramping up over 10 minutes to test the breaking point and auto-scaling triggers of an API gateway using k6. Provide the k6 stage configuration array for ramp-up load testing. | pass→pass | 9,112 | 4,913 | -46% | 1 | 1 | 0% | 1,631 | 869 | -47% | 0 | 0 | — |
▸case-22 A QA engineer is debugging a failing pytest unit test in `tests/test_auth.py` where an unexpected exception is raised during token validation. Provide the pytest flags to drop into the Python debugger (`pdb`) immediately upon test failure. | pass→pass | 3,760 | 3,270 | -13% | 1 | 1 | 0% | 627 | 531 | -15% | 0 | 0 | — |
▸case-23 A Java performance engineer needs to enable detailed garbage collection logging with unified JVM logging flags in Java 17 for an application `app.jar` to diagnose memory leaks. Provide the `java` command line invocation. | pass→pass | 10,781 | 5,773 | -46% | 1 | 1 | 0% | 1,589 | 1,123 | -29% | 0 | 0 | — |