Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Retrieve DNA/RNA/protein sequences from NCBI and ENA with disambiguation. Quality hierarchy: RefSeq (NM_/NP_) > RefSeq predicted (XM_/XP_) > GenBank submissions. Use for fetching specific sequences by accession, gene-symbol-to-sequence lookup, transcript-isoform retrieval, and curated-vs-raw-submission preference.
.claude/skills/mims-harvard-tooluniverse-sequence-retrieval/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 162% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 22% | 0% |
| case-08 | ✓→✗ | ▼ Worse | 9% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 94% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -20% | 0% |
Retrieve DNA, RNA, and protein sequences with proper disambiguation and cross-database handling.
IMPORTANT: Always use English terms in tool calls. Only try original-language terms as fallback. Respond in the user's language.
LOOK UP DON'T GUESS: Never assume accession numbers or sequence versions. Always retrieve and verify from NCBI or ENA.
Sequence quality hierarchy: RefSeq (NM_/NP_ = curated) > RefSeq predicted (XM_/XP_) > GenBank (submitted). Prefer the MANE Select transcript for human canonical isoforms. Check version numbers -- annotations improve across versions.
Phase 0: Clarify (if needed) → Phase 1: Disambiguate Gene/Organism → Phase 2: Search & Retrieve → Phase 3: ReportAsk ONLY if: gene exists in multiple organisms, sequence type unclear, or strain matters. Skip for: specific accessions, clear organism+gene combos, complete genome requests with organism.
| Prefix | Type | Use With | |--------|------|----------| | NC_/NM_/NR_/NP_/XM_ | RefSeq | NCBI only | | U/M/K/X/CP/NZ_ | GenBank | NCBI or ENA | | EMBL format | EMBL | ENA preferred |
CRITICAL: Never try ENA tools with RefSeq accessions -- they return 404.
Retrieve silently. Do NOT narrate the search process.
python# Search NCBI Nucleotide result = tu.tools.NCBI_search_nucleotide( operation="search", organism=organism, gene=gene, strain=strain, keywords=keywords, seq_type=seq_type, limit=10 ) # Get accessions from UIDs accessions = tu.tools.NCBI_fetch_accessions(operation="fetch_accession", uids=result["data"]["uids"]) # Retrieve sequence (FASTA or GenBank format) sequence = tu.tools.NCBI_get_sequence(operation="fetch_sequence", accession=accession, format="fasta") # ENA alternative (non-RefSeq accessions only) entry = tu.tools.ena_get_entry(accession=accession) fasta = tu.tools.ena_get_sequence_fasta(accession=accession)
| Primary | Fallback | Notes | |---------|----------|-------| | NCBI_get_sequence | ENA (if GenBank format) | NCBI unavailable | | ena_get_entry | NCBI_get_sequence | ENA doesn't have RefSeq | | NCBI_search_nucleotide | Try broader keywords | No results |
Present as a Sequence Profile Report. Hide search process. Include:
| Tier | Prefix | Description | |------|--------|-------------| | RefSeq Reference (best) | NC_, NM_, NP_ | NCBI-curated, gold standard | | RefSeq Predicted | XM_, XP_, XR_ | Computationally predicted | | GenBank Validated | Various | Submitted, some curation | | GenBank Direct | Various | Direct submission | | Third Party | TPA_ | Third-party annotation |
Sequence quality: Prefer RefSeq over GenBank. Check version numbers. Sequences with "PREDICTED" in definition are not experimentally validated.
Accession guidance: RefSeq = NCBI-only. GenBank = mirrored in ENA/EMBL. Default to RefSeq mRNA (NM_) for human/model organisms; most complete genome assembly for microbial queries.
Cross-database reconciliation: Same sequence may have different accessions (e.g., GenBank U00096 = RefSeq NC_000913 for E. coli K-12). Always report both when available. Discrepancies between GenBank/RefSeq typically indicate RefSeq curation corrected submission errors.
| Error | Response | |-------|----------| | "No search criteria provided" | Add organism, gene, or keywords | | "ENA 404 error" | Likely RefSeq -- use NCBI only | | "No results found" | Broaden search, check spelling, try synonyms | | "Sequence too large" | Note size, provide download link instead |
NCBI Tools: NCBI_search_nucleotide (search), NCBI_fetch_accessions (UID→accession), NCBI_get_sequence (retrieve) ENA Tools (GenBank/EMBL only): ena_get_entry (metadata), ena_get_sequence_fasta (FASTA), ena_get_entry_summary (summary)
NCBI_search_nucleotide: operation="search", organism (scientific name), gene (symbol), strain, keywords, seq_type (complete_genome/mrna/refseq), limit
NCBI_get_sequence: operation="fetch_sequence", accession, format (fasta/genbank)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,092 | 5,211 | -74% | 1 | 1 | 0% | 3,778 | 1,790 | -53% | 0 | 0 | — |
case-02 | fail→fail | 12,920 | 8,318 | -36% | 1 | 1 | 0% | 2,353 | 1,919 | -18% | 0 | 0 | — |
case-17 | pass→pass | 12,988 | 7,763 | -40% | 1 | 1 | 0% | 2,176 | 2,775 | +28% | 0 | 0 | — |
case-03 | fail→fail | 11,426 | 4,587 | -60% | 1 | 1 | 0% | 2,059 | 1,763 | -14% | 0 | 0 | — |
case-04 | pass→pass | 6,794 | 6,766 | -0% | 1 | 1 | 0% | 1,375 | 2,709 | +97% | 0 | 0 | — |
case-05 | pass→pass | 22,831 | 30,289 | +33% | 1 | 1 | 0% | 4,531 | 7,661 | +69% | 0 | 0 | — |
case-06 | pass→pass | 16,432 | 26,927 | +64% | 1 | 1 | 0% | 3,686 | 7,649 | +108% | 0 | 0 | — |
case-18 | pass→pass | 12,093 | 10,297 | -15% | 1 | 1 | 0% | 2,140 | 3,388 | +58% | 0 | 0 | — |
case-07 | pass→fail | 8,340 | 8,444 | +1% | 1 | 1 | 0% | 1,707 | 2,088 | +22% | 0 | 0 | — |
case-08 | pass→fail | 8,707 | 6,099 | -30% | 1 | 1 | 0% | 1,728 | 1,879 | +9% | 0 | 0 | — |
case-09 | pass→fail | 5,753 | 7,343 | +28% | 1 | 1 | 0% | 960 | 1,861 | +94% | 0 | 0 | — |
case-10 | pass→pass | 4,331 | 5,167 | +19% | 1 | 1 | 0% | 708 | 2,361 | +233% | 0 | 0 | — |
case-11 | pass→fail | 12,882 | 7,165 | -44% | 1 | 1 | 0% | 2,199 | 1,770 | -20% | 0 | 0 | — |
case-12 | pass→fail | 13,782 | 7,684 | -44% | 1 | 1 | 0% | 2,395 | 1,874 | -22% | 0 | 0 | — |
case-13 | pass→fail | 9,448 | 5,014 | -47% | 1 | 1 | 0% | 1,403 | 1,633 | +16% | 0 | 0 | — |
case-14 | fail→fail | 18,134 | 9,735 | -46% | 1 | 1 | 0% | 2,481 | 2,048 | -17% | 0 | 0 | — |
case-15 | fail→pass | 4,789 | 4,109 | -14% | 1 | 1 | 0% | 835 | 2,190 | +162% | 0 | 0 | — |
case-16 | pass→pass | 10,163 | 7,911 | -22% | 1 | 1 | 0% | 2,032 | 3,076 | +51% | 0 | 0 | — |
case-19 | pass→fail | 9,899 | 6,354 | -36% | 1 | 1 | 0% | 1,742 | 1,944 | +12% | 0 | 0 | — |
case-20 | pass→fail | 14,969 | 8,271 | -45% | 1 | 1 | 0% | 2,374 | 1,996 | -16% | 0 | 0 | — |
case-21 | pass→pass | 13,992 | 5,247 | -63% | 1 | 1 | 0% | 2,352 | 2,478 | +5% | 0 | 0 | — |
case-22 | pass→pass | 6,122 | 3,905 | -36% | 1 | 1 | 0% | 892 | 2,123 | +138% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 10 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -32 percentage points is the difference between those two pass rates over the 10 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.