Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill should be used when the user asks about "Elnora agent capabilities", "what can the agent do", "agent tools", "web search", "academic search", "PubMed", "ArXiv", "Exa", "Tavily", "Perplexity", "Valyu", "ToolUniverse", "scientific tools", "agent memory", "code execution", "sandbox", "search papers", "search literature", "drug discovery", "protein analysis", "clinical trials", "file operations", "agent skills", or any question about what the Elnora AI Agent can do when you send it a task
.claude/skills/majiayu000-elnora-agent/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 5% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 155% | 0% |
| case-20 | ✓→✓ | = Same ✓ | 265% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 31% | 0% |
The Elnora Agent is a sandboxed Python environment with ~78 core tools + 2,100 ToolUniverse scientific tools. Interact via elnora_create_task / elnora_send_message — describe what you need in plain language, don't reference internal tool names.
| Tool | Purpose | |------|---------| | elnora_create_task | Create a task with optional initial_message to start generation | | elnora_send_message | Send follow-up message. 30-120s for complex requests. | | elnora_get_task_messages | Read agent responses | | elnora_generate_protocol | Convenience: create task + send message in one call |
| Capability | Examples | |------------|----------| | Web search (34 tools) | Real-time search, neural/semantic search, deep research, URL extraction, site crawling. Providers: Tavily, Exa, Valyu, Perplexity | | Academic databases (12 tools) | PubMed, ArXiv, Semantic Scholar, bioRxiv, Europe PMC, OpenAlex, UniProt, ClinicalTrials.gov, ChEMBL, Wolfram Alpha | | 2,100+ scientific tools (ToolUniverse) | Protein structure (AlphaFold, PDB), genomics (Ensembl, ClinVar), chemistry (PubChem, DrugBank), pathways (KEGG, Reactome), drug safety (OpenFDA), and 21 more categories | | 35 domain skills | Literature review, experimental design, drug discovery workflow, protein engineering, single-cell RNA QC, statistical analysis, scientific writing | | File operations (11 tools) | Create/read/search files, full-text grep, upload attachments, link files to tasks | | Memory (9 tools) | Remember facts across tasks, share findings between agents, recall prior context | | Code execution | Persistent Python REPL with pandas, numpy, biopython. Variables survive across executions. 30s timeout, 1MB output max |
Web research: > "Search for recent CRISPR delivery methods and summarize the top findings"
Literature review: > "Search PubMed for BRCA1 DNA repair papers from 2024, find the most cited ones"
Drug target research: > "Search for compounds targeting EGFR, cross-reference with active clinical trials"
Scientific computation: > "Use ToolUniverse to run AlphaFold on this sequence: MVLSPADKTNVKAAWGKVGA"
Memory: > "Remember that our lab uses Q5 polymerase for all high-fidelity PCR at 62C"
File search: > "Search all project files for mentions of 'annealing temperature' and summarize"
Reference existing files: Use file_ids in elnora_send_message or context_file_ids in elnora_create_task to give the agent context about existing protocols.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,454 | 10,175 | -30% | 1 | 1 | 0% | 2,548 | 1,317 | -48% | 0 | 0 | — |
case-11 | fail→fail | 6,457 | 6,110 | -5% | 1 | 1 | 0% | 979 | 919 | -6% | 0 | 0 | — |
case-02 | fail→fail | 20,789 | 29,625 | +43% | 1 | 1 | 0% | 2,851 | 1,455 | -49% | 0 | 0 | — |
case-03 | fail→fail | 19,636 | 12,010 | -39% | 1 | 1 | 0% | 3,792 | 1,360 | -64% | 0 | 0 | — |
case-04 | fail→fail | 6,012 | 8,689 | +45% | 1 | 1 | 0% | 859 | 1,457 | +70% | 0 | 0 | — |
case-05 | fail→fail | 7,623 | 6,264 | -18% | 1 | 1 | 0% | 1,194 | 1,783 | +49% | 0 | 0 | — |
case-12 | fail→fail | 4,115 | 5,900 | +43% | 1 | 1 | 0% | 623 | 1,038 | +67% | 0 | 0 | — |
case-06 | pass→pass | 4,743 | 7,673 | +62% | 1 | 1 | 0% | 691 | 1,763 | +155% | 0 | 0 | — |
case-07 | pass→fail | 6,892 | 8,857 | +29% | 1 | 1 | 0% | 1,152 | 1,215 | +5% | 0 | 0 | — |
case-08 | fail→fail | 4,838 | 12,625 | +161% | 1 | 1 | 0% | 823 | 1,721 | +109% | 0 | 0 | — |
case-09 | fail→fail | 7,798 | 4,944 | -37% | 1 | 1 | 0% | 1,255 | 1,476 | +18% | 0 | 0 | — |
case-10 | fail→pass | 9,146 | 2,058 | -77% | 1 | 1 | 0% | 1,623 | 981 | -40% | 0 | 0 | — |
case-13 | fail→fail | 9,895 | 5,267 | -47% | 1 | 1 | 0% | 1,506 | 924 | -39% | 0 | 0 | — |
case-14 | fail→fail | 11,749 | 7,229 | -38% | 1 | 1 | 0% | 2,088 | 962 | -54% | 0 | 0 | — |
case-15 | fail→fail | 7,400 | 6,177 | -17% | 1 | 1 | 0% | 1,225 | 945 | -23% | 0 | 0 | — |
case-16 | fail→fail | 13,492 | 12,891 | -4% | 1 | 1 | 0% | 2,391 | 1,724 | -28% | 0 | 0 | — |
case-17 | fail→fail | 4,832 | 5,995 | +24% | 1 | 1 | 0% | 623 | 1,059 | +70% | 0 | 0 | — |
case-18 | fail→fail | 4,913 | 8,418 | +71% | 1 | 1 | 0% | 720 | 1,245 | +73% | 0 | 0 | — |
case-19 | fail→fail | 10,406 | 8,541 | -18% | 1 | 1 | 0% | 1,991 | 958 | -52% | 0 | 0 | — |
case-20 | pass→pass | 2,329 | 2,680 | +15% | 1 | 1 | 0% | 283 | 1,034 | +265% | 0 | 0 | — |
case-21 | pass→pass | 14,209 | 15,598 | +10% | 1 | 1 | 0% | 2,505 | 3,274 | +31% | 0 | 0 | — |
case-22 | pass→pass | 14,994 | 11,216 | -25% | 1 | 1 | 0% | 2,784 | 2,581 | -7% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 7 counted toward the lift figure. The other 15 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 7 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.