Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Tests a market theory the user defines, automatically, over time. The user states a hypothesis like "when ETH rises more than 3% in 24h, NEAR rises within the next 48h"; the agent checks prices every 6 hours, logs each trigger event, and 48h later checks whether the predicted outcome held — building a confirmed/refuted tally and a hit rate. Every Sunday it sends a Telegram report with the running results and a verdict (promising / inconclusive / rejected), so a hunch becomes evidence instead of
.claude/skills/nearai-crypto-hypothesis-tester/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -65% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 113% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 250% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 21% | 0% |
You test a market hypothesis the user defines by tracking it over time and reporting whether it holds.
hypotheses/active.md with memory_read before writing, then write the full updated file back with memory_write. Never overwrite from scratch and never drop earlier log entries.http tool and the current date/time from the time tool. Never guess a price or a date — the 24h and 48h windows depend on real timestamps.When the user states a hypothesis (e.g. test hypothesis: when ETH rises more than 3% in 24h, NEAR rises within 48h):
hypotheses/active.md with memory_read.time tool), confirmed 0, refuted 0, trigger log empty.memory_write.Testing started. I'll check every 6 hours and report the verdict each Sunday.Stored in hypotheses/active.md:
Hypothesis: [exact text]
- Trigger: [token] [condition]
- Outcome: [token] [window]
- Status: TESTING
- Started: [date]
- Confirmed: [X]
- Refuted: [X]
- Trigger log:
- [date] | triggered: [token] [+X%] | outcome token at trigger: $[price] | checked: pending/CONFIRMED/REFUTEDCreate a routine that runs every 6 hours. The routine goal must contain these full steps as a self-contained prompt, because a routine does not keep any context from this conversation when it runs:
hypotheses/active.md with memory_read.time tool.http tool: https://api.coingecko.com/api/v3/simple/price?ids=[trigger-id],[outcome-id]&vs_currencies=usd&include_24hr_change=true (map symbols to CoinGecko ids yourself — ETH=ethereum, NEAR=near, etc.).memory_write.HEARTBEAT_OK and stop.Weekly report:
🧪 Hypothesis Test — [date]
Theory: [exact hypothesis text]
Status: TESTING (day [X])
✅ Confirmed: [X]
❌ Refuted: [X]
📊 Hit rate: [X] of [Y]
Last trigger: [date] — [trigger move] → [outcome]
Verdict: [PROMISING if confirmed clearly lead / INCONCLUSIVE if close / REJECTED if refuted clearly lead]test hypothesis: [your theory] — start tracking a hypothesishypothesis status — show the current theory, tally, hit rate, and verdict right nowreset hypothesis — clear the current hypothesis and start fresh| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,673 | 6,682 | -23% | 1 | 1 | 0% | 1,603 | 1,501 | -6% | 0 | 0 | — |
case-02 | fail→fail | 2,910 | 4,567 | +57% | 1 | 1 | 0% | 489 | 1,342 | +174% | 0 | 0 | — |
case-03 | fail→fail | 2,991 | 4,379 | +46% | 1 | 1 | 0% | 465 | 1,298 | +179% | 0 | 0 | — |
case-04 | fail→fail | 11,536 | 5,509 | -52% | 1 | 1 | 0% | 1,899 | 1,402 | -26% | 0 | 0 | — |
case-05 | fail→pass | 25,218 | 4,230 | -83% | 1 | 1 | 0% | 5,291 | 1,860 | -65% | 0 | 0 | — |
case-06 | fail→pass | 7,143 | 3,109 | -56% | 1 | 1 | 0% | 1,170 | 1,560 | +33% | 0 | 0 | — |
case-07 | fail→pass | 3,966 | 3,339 | -16% | 1 | 1 | 0% | 791 | 1,683 | +113% | 0 | 0 | — |
case-08 | fail→fail | 9,744 | 1,713 | -82% | 1 | 1 | 0% | 1,601 | 1,329 | -17% | 0 | 0 | — |
case-09 | fail→pass | 2,244 | 1,940 | -14% | 1 | 1 | 0% | 377 | 1,320 | +250% | 0 | 0 | — |
case-10 | pass→pass | 9,849 | 1,945 | -80% | 1 | 1 | 0% | 1,685 | 1,423 | -16% | 0 | 0 | — |
case-11 | pass→pass | 9,840 | 2,830 | -71% | 1 | 1 | 0% | 1,747 | 1,542 | -12% | 0 | 0 | — |
case-12 | fail→fail | 3,138 | 5,637 | +80% | 1 | 1 | 0% | 519 | 1,544 | +197% | 0 | 0 | — |
case-13 | pass→pass | 12,447 | 6,636 | -47% | 1 | 1 | 0% | 2,223 | 2,155 | -3% | 0 | 0 | — |
case-14 | pass→pass | 11,364 | 3,251 | -71% | 1 | 1 | 0% | 2,138 | 1,800 | -16% | 0 | 0 | — |
case-15 | fail→pass | 7,130 | 1,793 | -75% | 1 | 1 | 0% | 1,138 | 1,379 | +21% | 0 | 0 | — |
case-16 | fail→pass | 16,027 | 3,397 | -79% | 1 | 1 | 0% | 2,795 | 1,685 | -40% | 0 | 0 | — |
case-17 | fail→pass | 6,089 | 1,792 | -71% | 1 | 1 | 0% | 961 | 1,381 | +44% | 0 | 0 | — |
case-18 | fail→pass | 6,703 | 1,811 | -73% | 1 | 1 | 0% | 1,098 | 1,349 | +23% | 0 | 0 | — |
case-19 | fail→pass | 4,949 | 2,108 | -57% | 1 | 1 | 0% | 854 | 1,381 | +62% | 0 | 0 | — |
case-20 | fail→pass | 10,211 | 2,292 | -78% | 1 | 1 | 0% | 1,988 | 1,471 | -26% | 0 | 0 | — |
case-21 | fail→pass | 10,634 | 3,119 | -71% | 1 | 1 | 0% | 1,898 | 1,670 | -12% | 0 | 0 | — |
case-22 | fail→fail | 12,002 | 7,936 | -34% | 1 | 1 | 0% | 2,007 | 1,723 | -14% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.