Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Research Elixir/Phoenix/Ecto topics or evaluate Hex libraries (--library). Use when learning about libraries, patterns, or comparing approaches. Searches HexDocs, ElixirForum, GitHub.
.claude/skills/oliver-kriska-phx-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 681% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -56% | 0% |
| case-20 | ✓→✗ | ▼ Worse | -64% | 0% |
Research a topic by searching the web and fetching relevant sources efficiently.
/skill:phx-research Oban unique jobs best practices
/skill:phx-research LiveView file upload with progress
/skill:phx-research --library permitthe text after the skill name = Research topic/question. Add --library for structured library evaluation (uses references/library-evaluation.md template).
If the text after the skill name contains --library or the topic is clearly about evaluating a Hex dependency (e.g., "should we use permit", "evaluate sagents", "compare oban vs exq"):
references/library-evaluation.md for the template.claude/research/{lib}-evaluation.mdCheck .claude/research/{topic-slug}.md. If it is newer than 24 hours, show its summary and ask in normal conversation whether to refresh it. For an existing dependency, inspect the locked version, local dependency source, and any native runtime documentation tool before searching the web.
Turn the user request into one focused query when it is short, or two to four queries of at most ten words for a multi-part request. Never send a long raw prompt to a search provider.
Use the runtime's native web or HTTP capabilities when available. Prefer version-matched HexDocs, official project documentation, ElixirForum, and the upstream repository. Deduplicate URLs and discard irrelevant results. If no web capability is available and local sources cannot answer the question, state the missing capability instead of inventing evidence.
Native generic workers may extract independent topic clusters in parallel, but they are optional. The same-session sequential path must remain complete. Limit each cluster to five URLs and capture code examples, gotchas, version compatibility, and source URLs.
Write .claude/research/{topic-slug}.md (about 5 KB for topic research or 3 KB for a library evaluation) with:
Present the summary and offer, in normal conversation, to plan from it, investigate one finding, research a narrower subtopic, or stop. Never invoke a follow-up workflow automatically.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→fail | 15,203 | 6,677 | -56% | 1 | 1 | 0% | 725 | 1,117 | +54% | 0 | 0 | — |
case-01 | fail→fail | 20,574 | 6,393 | -69% | 1 | 1 | 0% | 3,503 | 1,095 | -69% | 0 | 0 | — |
case-02 | fail→fail | 17,158 | 6,360 | -63% | 1 | 1 | 0% | 2,750 | 1,096 | -60% | 0 | 0 | — |
case-03 | fail→fail | 23,324 | 6,947 | -70% | 1 | 1 | 0% | 3,859 | 1,082 | -72% | 0 | 0 | — |
case-04 | fail→pass | 3,139 | 12,328 | +293% | 1 | 1 | 0% | 342 | 2,670 | +681% | 0 | 0 | — |
case-06 | pass→pass | 16,831 | 29,434 | +75% | 1 | 1 | 0% | 2,913 | 5,403 | +85% | 0 | 0 | — |
case-07 | pass→pass | 13,542 | 10,800 | -20% | 1 | 1 | 0% | 1,995 | 2,301 | +15% | 0 | 0 | — |
case-08 | fail→fail | 15,109 | 9,602 | -36% | 1 | 1 | 0% | 2,535 | 1,235 | -51% | 0 | 0 | — |
case-09 | fail→fail | 21,079 | 9,540 | -55% | 1 | 1 | 0% | 3,881 | 1,409 | -64% | 0 | 0 | — |
case-10 | fail→fail | 15,003 | 8,050 | -46% | 1 | 1 | 0% | 2,631 | 1,534 | -42% | 0 | 0 | — |
case-11 | pass→fail | 14,830 | 4,355 | -71% | 1 | 1 | 0% | 2,227 | 969 | -56% | 0 | 0 | — |
case-12 | pass→pass | 6,954 | 1,645 | -76% | 1 | 1 | 0% | 990 | 921 | -7% | 0 | 0 | — |
case-13 | fail→fail | 9,373 | 8,193 | -13% | 1 | 1 | 0% | 1,494 | 1,122 | -25% | 0 | 0 | — |
case-14 | pass→pass | 13,279 | 9,319 | -30% | 1 | 1 | 0% | 2,234 | 2,212 | -1% | 0 | 0 | — |
case-19 | pass→pass | 16,926 | 15,484 | -9% | 1 | 1 | 0% | 3,552 | 4,046 | +14% | 0 | 0 | — |
case-15 | pass→pass | 8,065 | 4,184 | -48% | 1 | 1 | 0% | 1,180 | 1,412 | +20% | 0 | 0 | — |
case-16 | fail→pass | 8,723 | 1,777 | -80% | 1 | 1 | 0% | 1,300 | 946 | -27% | 0 | 0 | — |
case-17 | fail→pass | 9,681 | 1,634 | -83% | 1 | 1 | 0% | 1,551 | 937 | -40% | 0 | 0 | — |
case-18 | pass→pass | 18,456 | 5,956 | -68% | 1 | 1 | 0% | 2,785 | 1,584 | -43% | 0 | 0 | — |
case-20 | pass→fail | 17,656 | 6,382 | -64% | 1 | 1 | 0% | 2,930 | 1,062 | -64% | 0 | 0 | — |
case-21 | fail→fail | 1,559 | 4,265 | +174% | 1 | 1 | 0% | 217 | 1,028 | +374% | 0 | 0 | — |
case-22 | pass→fail | 9,329 | 4,921 | -47% | 1 | 1 | 0% | 1,794 | 1,063 | -41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 10 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 10 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.