---
name: xuzhougeng/evo2
source: https://app.decimal.ai/s/xuzhougeng-evo2@1/SKILL.md
source_sha256: a633ea00c5bd
---

# Evo 2 — DNA Language Model

## Prerequisites

| Requirement | Minimum | Recommended      |
| ----------- | ------- | ---------------- |
| Python      | 3.11    | 3.12 (<3.13)     |
| CUDA        | 12.1+   | 12.4+            |
| GPU VRAM    | 24 GB (7B bf16) | 80 GB (40B) |
| RAM         | 32 GB   | 128 GB           |

## How to run

### Installation

```bash
pip install evo2
# Weights pulled from Hugging Face on first model load.
```

### Loading and scoring

```python
from evo2 import Evo2

model = Evo2("evo2_7b")        # or "evo2_40b" — see model table
seqs = ["ATCG" * 50, "GGGCTTAA" * 25]
ll = model.score_sequences(seqs)   # → list[float], mean per-token log-likelihood
print(ll)
```

### Generation

```python
out = model.generate(
    prompt_seqs=["ATGAAAGCT"],
    n_tokens=256,
    temperature=0.7,
)
print(out.sequences[0])
```

## Models

| Name        | Params | Context | VRAM (bf16) | Notes                              |
| ----------- | ------ | ------- | ----------- | ---------------------------------- |
| `evo2_7b`   | 7 B    | 1 M nt  | ~22 GB      | Default; fits on a single 24 GB+ GPU |
| `evo2_40b`  | 40 B   | 1 M nt  | ~78 GB      | H100 80 GB or multi-GPU            |
| `evo2_1b_base` | 1 B | 8 K nt  | ~6 GB       | FP8 path requires sm_89+ (H100)    |

## Output format

`score_sequences` returns a `list[float]` (or `np.ndarray`) of mean log-likelihoods,
one per input sequence. More negative ⇒ less likely under the model. For variant
effect, compute `Δll = ll_alt - ll_ref` over a fixed window.

`generate` returns a `GenerationOutput` with `.sequences` (list[str]), `.logits`
(list[Tensor]), and `.logprobs_mean` (list[float]) — always populated, no flag required.

## Decision tree

```
Need a DNA model?
│
├─ Per-base/per-sequence likelihood, generation → Evo 2 ✓
├─ Predict experimental tracks (expression, accessibility) → borzoi
└─ Protein, not DNA → fair-esm2 / esmfold2
```


## Remote compute

7B/40B inference is GPU-bound (≥24 GB / 80 GB VRAM). Use a selected and
probed `ssh:<alias>` context and load `remote-compute-ssh`. Confirm that the
environment imports Evo 2 and that the desired weights are cached. Submit a
self-contained scoring script through one `run_in_context` call:

```json
{
  "context_id": "ssh:gpu-box",
  "title": "Evo 2 variant scoring",
  "command": "source ~/miniforge3/etc/profile.d/conda.sh && conda activate evo2 && HF_HOME=/srv/model-cache HF_HUB_OFFLINE=1 python score_evo2.py --output /home/me/wisp-results/evo2/scores.json",
  "timeout_secs": 1800,
  "input_paths": ["runs/score_evo2.py"],
  "output_specs": [
    {
      "glob": "ssh://gpu-box/home/me/wisp-results/evo2/scores.json",
      "kind": "json",
      "residency": "remote"
    }
  ]
}
```

Replace context, environment, cache, and output paths with discovered values.
Call `monitor_run` once to wait, `get_run` once for a snapshot, or `cancel_run`
to stop. Set `HF_HUB_OFFLINE=1` only after confirming the cache is complete, so
the loader does not try to write `refs/` into a read-only mount. Weight footprint:
~15 GB (7B), ~80 GB (40B).


## Typical performance

| Task                        | 7B on H100 | Notes                       |
| --------------------------- | ---------- | --------------------------- |
| Model load (cached)         | ~5-7 min   | First call hydrates weights |
| `score_sequences`, 200×200bp| ~10-20 s   | After load                  |
| `generate`, 1×512 nt        | ~15 s      |                             |

## Troubleshooting

| Symptom                              | Cause                          | Fix                                        |
| ------------------------------------ | ------------------------------ | ------------------------------------------ |
| `Transformer Engine not installed`   | No FP8 — falls back to bf16    | Informational only on non-H100; ignore     |
| OOM on load                          | 40B on <80 GB GPU              | Use `evo2_7b` or shard with `device_map`   |
| HF tries to write `refs/main`        | `HF_HOME` points at RO mount   | Set `HF_HUB_OFFLINE=1`                     |
| `dtype mismatch` in `score_sequences`| Passing tensors not strings    | Pass `list[str]`; the API tokenises for you |

---

**Next**: pair with `borzoi` to predict track-level effects of the same
variants.