Install any skill in seconds. Free to start, no credit card required.
Get Started Free →AI-guided de novo protein design — RFdiffusion backbone generation, ProteinMPNN sequence design, structure validation (pLDDT, pTM, MPNN scores). Use for designing therapeutic protein binders, novel scaffolds, enzyme variants, and miniprotein/protein-interface design before experimental validation.
.claude/skills/mims-harvard-tooluniverse-protein-therapeutic-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 162% | 0% |
AI-guided de novo protein design using RFdiffusion backbone generation, ProteinMPNN sequence optimization, and structure validation for therapeutic protein development.
KEY PRINCIPLES:
Therapeutic protein design starts with the target interaction. What binding surface do you need to cover? A small pocket = nanobody or peptide. A large flat surface = designed protein. Stability, immunogenicity, and manufacturability constrain the design space.
When uncertain about any scientific fact, SEARCH databases first rather than reasoning from memory. A database-verified answer is always more reliable than a guess.
When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.
Apply when user asks to:
Phase 1: Target Characterization
Get structure (PDB, EMDB cryo-EM, AlphaFold), identify binding epitope
Phase 2: Backbone Generation (RFdiffusion)
Define constraints, generate >= 5 backbones, filter by geometry
Phase 3: Sequence Design (ProteinMPNN)
Design >= 8 sequences per backbone, sample with temperature control
Phase 4: Structure Validation (ESMFold/AlphaFold2)
Predict structure, compare to backbone, assess pLDDT/pTM
Phase 5: Developability Assessment
Aggregation, pI, expression prediction
Phase 6: Report Synthesis
Ranked candidates, FASTA, experimental recommendations[TARGET]_protein_design_report.md first with section headers[TARGET]_designed_sequences.fasta and [TARGET]_top_candidates.csvEvery design MUST include: Sequence, Length, Target, Method, and Quality Metrics (pLDDT, pTM, MPNN score, binding prediction).
| Tool | Purpose | Key Parameter | |------|---------|---------------| | NvidiaNIM_rfdiffusion (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) | Backbone generation | diffusion_steps (NOT num_steps) | | NvidiaNIM_proteinmpnn (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) | Sequence design | pdb_string (NOT pdb) | | ESMFold_predict_structure | Fast validation | sequence (NOT seq) | | NvidiaNIM_alphafold2 (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) | High-accuracy structure inference from sequence | sequence, algorithm | | NvidiaNIM_esm2_650m (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) | Sequence embeddings | sequences, format |
| Tool | Wrong | Correct | |------|-------|---------| | NvidiaNIM_rfdiffusion (requires NVIDIA_API_KEY) | num_steps=50 | diffusion_steps=50 | | NvidiaNIM_proteinmpnn (requires NVIDIA_API_KEY) | pdb=content | pdb_string=content | | ESMFold_predict_structure | seq="MVLS..." | sequence="MVLS..." | | NvidiaNIM_alphafold2 (requires NVIDIA_API_KEY) | seq="MVLS..." | sequence="MVLS..." |
NVIDIA_API_KEY environment variable required| Tool | Purpose | Key Parameters | |------|---------|----------------| | PDBe_get_uniprot_mappings | Find PDB structures | uniprot_id | | RCSBData_get_entry | Download PDB file | pdb_id | | alphafold_get_prediction | Get AlphaFold DB structure | accession | | EMDB_search_structures | Search cryo-EM maps | query | | EMDB_get_structure | Get entry details | entry_id | | UniProt_get_entry_by_accession | Get target sequence | accession | | InterPro_get_protein_domains | Get domains | accession |
| Tier | Criteria | |------|----------| | T1 (best) | pLDDT >85, pTM >0.8, low aggregation, neutral pI | | T2 | pLDDT >75, pTM >0.7, acceptable developability | | T3 | pLDDT >70, pTM >0.65, developability concerns | | T4 | Failed validation or major developability issues |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 35,811 | 6,963 | -81% | 1 | 1 | 0% | 6,227 | 2,151 | -65% | 0 | 0 | — |
case-02 | fail→fail | 30,908 | 36,950 | +20% | 1 | 1 | 0% | 6,237 | 2,439 | -61% | 0 | 0 | — |
case-03 | fail→fail | 34,893 | 7,949 | -77% | 1 | 1 | 0% | 6,232 | 2,200 | -65% | 0 | 0 | — |
case-04 | fail→pass | 12,441 | 7,243 | -42% | 1 | 1 | 0% | 2,310 | 2,937 | +27% | 0 | 0 | — |
case-05 | fail→pass | 10,079 | 3,230 | -68% | 1 | 1 | 0% | 1,811 | 2,189 | +21% | 0 | 0 | — |
case-06 | pass→pass | 8,990 | 3,031 | -66% | 1 | 1 | 0% | 1,533 | 2,139 | +40% | 0 | 0 | — |
case-07 | fail→pass | 8,720 | 2,780 | -68% | 1 | 1 | 0% | 1,568 | 2,149 | +37% | 0 | 0 | — |
case-08 | pass→pass | 12,196 | 18,660 | +53% | 1 | 1 | 0% | 1,991 | 3,730 | +87% | 0 | 0 | — |
case-09 | pass→pass | 14,850 | 11,065 | -25% | 1 | 1 | 0% | 2,782 | 3,598 | +29% | 0 | 0 | — |
case-10 | pass→pass | 14,323 | 3,514 | -75% | 1 | 1 | 0% | 2,288 | 2,202 | -4% | 0 | 0 | — |
case-11 | fail→pass | 18,952 | 2,946 | -84% | 1 | 1 | 0% | 1,631 | 2,151 | +32% | 0 | 0 | — |
case-12 | pass→pass | 14,074 | 4,762 | -66% | 1 | 1 | 0% | 2,183 | 2,467 | +13% | 0 | 0 | — |
case-13 | fail→pass | 18,781 | 3,581 | -81% | 1 | 1 | 0% | 830 | 2,177 | +162% | 0 | 0 | — |
case-14 | fail→pass | 26,221 | 3,704 | -86% | 1 | 1 | 0% | 2,005 | 2,085 | +4% | 0 | 0 | — |
case-15 | fail→pass | 13,386 | 5,173 | -61% | 1 | 1 | 0% | 2,191 | 2,530 | +15% | 0 | 0 | — |
case-16 | pass→pass | 16,979 | 15,594 | -8% | 1 | 1 | 0% | 2,994 | 3,997 | +34% | 0 | 0 | — |
case-17 | pass→fail | 9,151 | 6,162 | -33% | 1 | 1 | 0% | 1,634 | 1,916 | +17% | 0 | 0 | — |
case-18 | fail→pass | 7,894 | 3,755 | -52% | 1 | 1 | 0% | 1,379 | 2,226 | +61% | 0 | 0 | — |
case-19 | pass→pass | 13,155 | 24,787 | +88% | 1 | 1 | 0% | 2,025 | 5,716 | +182% | 0 | 0 | — |
case-20 | fail→fail | 11,710 | 7,111 | -39% | 1 | 1 | 0% | 2,164 | 2,066 | -5% | 0 | 0 | — |
case-21 | fail→fail | 17,084 | 9,444 | -45% | 1 | 1 | 0% | 3,003 | 2,027 | -33% | 0 | 0 | — |
case-22 | fail→fail | 13,077 | 8,367 | -36% | 1 | 1 | 0% | 1,872 | 1,968 | +5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 13 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 13 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.