Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Quick web scanning for landscape understanding. Import of web-browsing/web-search skill. Snippets only — no conclusions from snippets alone.
.claude/skills/yogsoth-ai-stress-test-web-search/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-07 | ✓→✗ | ▼ Worse | -78% | 0% |
| case-10 | ✓→✗ | ▼ Worse | -38% | 0% |
| case-12 | ✓→✗ | ▼ Worse | -41% | 0% |
Quick web scanning for landscape understanding.
Import — strictly follow web-browsing/web-search skill protocol.
Snippets only — do not draw conclusions from search result snippets alone. Snippets provide leads for deeper investigation (web-research or paper-search).
Quantity target is set by the calling strategy's budget table. This SOP executes one unit = one brave_web_search call (count=10 results).
web-browsing repo → skills/web-search/SKILL.md
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | stress-test-paper-search | Paper AI summary report via alphaxiv get_paper_content. Import of literature-engine/paper-search skill. Structured AI-generated intermediate report. | | stress-test-web-research | Deep web full-text retrieval via Apify RAG browser. Import of web-browsing/web-research skill. Full page content for substantive analysis. | | web-search | Quick web scanning — discover pages, get snippets, find URLs. For orientation only, not substantive analysis. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,189 | 15,324 | +1% | 1 | 1 | 0% | 272 | 569 | +109% | 0 | 0 | — |
case-02 | fail→fail | 19,031 | 15,545 | -18% | 1 | 1 | 0% | 2,460 | 583 | -76% | 0 | 0 | — |
case-03 | fail→fail | 15,813 | 15,140 | -4% | 1 | 1 | 0% | 280 | 549 | +96% | 0 | 0 | — |
case-04 | fail→fail | 15,349 | 15,924 | +4% | 1 | 1 | 0% | 235 | 596 | +154% | 0 | 0 | — |
case-05 | fail→pass | 24,103 | 23,302 | -3% | 1 | 1 | 0% | 1,532 | 1,931 | +26% | 0 | 0 | — |
case-06 | fail→pass | 18,552 | 7,391 | -60% | 1 | 1 | 0% | 570 | 701 | +23% | 0 | 0 | — |
case-07 | pass→fail | 22,128 | 16,116 | -27% | 1 | 1 | 0% | 2,703 | 607 | -78% | 0 | 0 | — |
case-08 | fail→fail | 15,668 | 16,198 | +3% | 1 | 1 | 0% | 234 | 507 | +117% | 0 | 0 | — |
case-09 | fail→fail | 14,871 | 10,364 | -30% | 1 | 1 | 0% | 193 | 528 | +174% | 0 | 0 | — |
case-10 | pass→fail | 11,357 | 15,464 | +36% | 1 | 1 | 0% | 926 | 576 | -38% | 0 | 0 | — |
case-11 | fail→fail | 10,054 | 16,753 | +67% | 1 | 1 | 0% | 237 | 745 | +214% | 0 | 0 | — |
case-12 | pass→fail | 14,124 | 8,686 | -39% | 1 | 1 | 0% | 1,542 | 911 | -41% | 0 | 0 | — |
case-13 | fail→fail | 17,931 | 16,143 | -10% | 1 | 1 | 0% | 467 | 554 | +19% | 0 | 0 | — |
case-14 | fail→fail | 21,185 | 11,466 | -46% | 1 | 1 | 0% | 794 | 601 | -24% | 0 | 0 | — |
case-15 | fail→fail | 16,227 | 6,487 | -60% | 1 | 1 | 0% | 443 | 666 | +50% | 0 | 0 | — |
case-16 | fail→fail | 16,635 | 10,976 | -34% | 1 | 1 | 0% | 365 | 647 | +77% | 0 | 0 | — |
case-17 | pass→fail | 13,945 | 15,979 | +15% | 1 | 1 | 0% | 2,301 | 585 | -75% | 0 | 0 | — |
case-18 | fail→fail | 15,943 | 16,496 | +3% | 1 | 1 | 0% | 250 | 606 | +142% | 0 | 0 | — |
case-19 | fail→fail | 10,376 | 15,709 | +51% | 1 | 1 | 0% | 239 | 582 | +144% | 0 | 0 | — |
case-20 | fail→fail | 7,952 | 24,619 | +210% | 1 | 1 | 0% | 1,206 | 1,313 | +9% | 0 | 0 | — |
case-21 | fail→fail | 33,362 | 61,777 | +85% | 1 | 1 | 0% | 6,723 | 7,191 | +7% | 0 | 0 | — |
case-22 | fail→fail | 66,922 | 15,317 | -77% | 1 | 1 | 0% | 8,223 | 1,052 | -87% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 0 counted toward the lift figure. The other 22 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. A headline lift is not published for this run.
Other measured skills in the registry, with their headline benchmark lift.