Install any skill in seconds. Free to start, no credit card required.
Get Started Free →**Individual metrics have weak predictive power for binding**. Research shows:
.claude/skills/protein-qc/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | — | — |
| case-17 | ✗→✓ | ▲ Improved | — | — |
| case-12 | ✗→✓ | ▲ Improved | — | — |
| case-16 | ✗→✓ | ▲ Improved | — | — |
| case-03 | ✗→✓ | ▲ Improved | — | — |
Individual metrics have weak predictive power for binding. Research shows:
These thresholds filter out poor designs but do NOT predict binding affinity.
QC is organized by purpose and level:
| Purpose | What it assesses | Key metrics | |---------|------------------|-------------| | Binding | Interface quality, binding geometry | ipTM, PAE, SC, dG, dSASA | | Expression | Manufacturability, solubility | Instability, GRAVY, pI, cysteines | | Structural | Fold confidence, consistency | pLDDT, pTM, scRMSD |
Each category has two levels:
| Category | Metric | Standard | Stringent | Source | |----------|--------|----------|-----------|--------| | Structural | pLDDT | > 0.85 | > 0.90 | AF2/Chai/Boltz | | | pTM | > 0.70 | > 0.80 | AF2/Chai/Boltz | | | scRMSD | < 2.0 Å | < 1.5 Å | Design vs pred | | Binding | ipTM | > 0.50 | > 0.60 | AF2/Chai/Boltz | | | PAE_interaction | < 12 Å | < 10 Å | AF2/Chai/Boltz | | | Shape Comp (SC) | > 0.50 | > 0.60 | PyRosetta | | | interface_dG | < -10 | < -15 | PyRosetta | | Expression | Instability | < 40 | < 30 | BioPython | | | GRAVY | < 0.4 | < 0.2 | BioPython | | | ESM2 PLL | > 0.0 | > 0.2 | ESM2 |
| Pattern | Risk | Action | |---------|------|--------| | Odd cysteine count | Unpaired disulfides | Redesign | | NG/NS/NT motifs | Deamidation | Flag/avoid | | K/R >= 3 consecutive | Proteolysis | Flag | | >= 6 hydrophobic run | Aggregation | Redesign |
See: references/binding-qc.md, references/expression-qc.md, references/structural-qc.md
pythonimport pandas as pd designs = pd.read_csv('designs.csv') # Stage 1: Structural confidence designs = designs[designs['pLDDT'] > 0.85] # Stage 2: Self-consistency designs = designs[designs['scRMSD'] < 2.0] # Stage 3: Binding quality designs = designs[(designs['ipTM'] > 0.5) & (designs['PAE_interaction'] < 10)] # Stage 4: Sequence plausibility designs = designs[designs['esm2_pll_normalized'] > 0.0] # Stage 5: Expression checks (design-level) designs = designs[designs['cysteine_count'] % 2 == 0] # Even cysteines designs = designs[designs['instability_index'] < 40]
Individual metrics alone are too weak. Use composite scoring:
pythondef composite_score(row): return ( 0.30 * row['pLDDT'] + 0.20 * row['ipTM'] + 0.20 * (1 - row['PAE_interaction'] / 20) + 0.15 * row['shape_complementarity'] + 0.15 * row['esm2_pll_normalized'] ) designs['score'] = designs.apply(composite_score, axis=1) top_designs = designs.nlargest(100, 'score')
For advanced composite scoring, see references/composite-scoring.md.
| Level | Use Case | Stringency | |-------|----------|------------| | Default | Standard design | Most stringent | | Relaxed | Need more designs | Higher failure rate | | Peptide | Designs < 30 AA | ~5-10x lower success |
bashboltzgen run ... \ --budget 60 \ --alpha 0.01 \ --filter_biased true \ --refolding_rmsd_threshold 2.0 \ --additional_filters 'ALA_fraction<0.3'
alpha=0.0: Quality-only rankingalpha=0.01: Default (slight diversity)alpha=1.0: Diversity-onlyFor pattern-based checks, use severity scoring:
| Severity Level | Score | Action | |----------------|-------|--------| | LOW | 0-15 | Proceed | | MODERATE | 16-35 | Review flagged issues | | HIGH | 36-60 | Redesign recommended | | CRITICAL | 61+ | Redesign required |
| Metric | AUC | Use | |--------|-----|-----| | ipTM | ~0.64 | Pre-screening | | PAE | ~0.65 | Pre-screening | | ESM2 PLL | ~0.72 | Best single metric | | Composite | ~0.75+ | Always use |
Key insight: Metrics work as filters (eliminating failures) not predictors (ranking successes).
Quick assessment of your design campaign:
| Pass Rate | Status | Interpretation | |-----------|--------|----------------| | > 15% | Excellent | Above average, proceed | | 10-15% | Good | Normal, proceed | | 5-10% | Marginal | Below average, review issues | | < 5% | Poor | Significant problems, diagnose |
Low pLDDT across campaign
├── Check scRMSD distribution
│ ├── High scRMSD (>2.5Å): Backbone issue
│ │ └── Fix: Regenerate backbones with lower noise_scale (0.5-0.8)
│ └── Low scRMSD but low pLDDT: Disordered regions
│ └── Fix: Check design length, simplify topology
├── Try more sequences per backbone
│ └── modal run modal_proteinmpnn.py --num-seq-per-target 32 --sampling-temp 0.1
├── Use SolubleMPNN instead of ProteinMPNN
│ └── Better for expression-optimized sequences
└── Consider different design tool
└── BindCraft (integrated design) may work betterLow ipTM across campaign
├── Review hotspot selection
│ ├── Are hotspots surface-exposed? (SASA > 20Ų)
│ ├── Are hotspots conserved? (check MSA)
│ └── Try 3-6 different hotspot combinations
├── Increase binder length (more contact area)
│ └── Try 80-100 AA instead of 60-80 AA
├── Check interface geometry
│ ├── Is target flat? → Try helical binders
│ └── Is target concave? → Try smaller binders
└── Try all-atom design tool
└── BoltzGen (all-atom, better packing)Sequences don't specify intended structure
├── ProteinMPNN issue
│ ├── Lower temperature: --sampling-temp 0.1
│ ├── Increase sequences: --num-seq-per-target 32
│ └── Check fixed_positions aren't over-constraining
├── Backbone geometry issue
│ ├── Backbones may be unusual/strained
│ ├── Regenerate with lower noise_scale (0.5-0.8)
│ └── Reduce diffuser.T to 30-40
└── Try different sequence design
└── ColabDesign (AF2 gradient-based) may work betterIn silico metrics don't predict affinity
├── Generate MORE designs (10x current)
│ └── Computational metrics have high false positive rate
├── Increase diversity
│ ├── Higher ProteinMPNN temperature (0.2-0.3)
│ ├── Different backbone topologies
│ └── Different hotspot combinations
├── Try different design approach
│ ├── BindCraft (different algorithm)
│ ├── ColabDesign (AF2 hallucination)
│ └── BoltzGen (all-atom diffusion)
└── Check if target is druggable
└── Some targets are inherently difficultSuspiciously high pass rate
├── Check if thresholds are too lenient
│ └── Use stringent thresholds: pLDDT > 0.90, ipTM > 0.60
├── Verify prediction quality
│ ├── Are predictions actually running? Check output files
│ └── Are complexes being predicted, not just monomers?
├── Check for data issues
│ ├── Same sequence being predicted multiple times?
│ └── Wrong FASTA format (missing chain separator)?
└── Apply diversity filter
└── Cluster at 70% identity, take top per clusterpythonimport pandas as pd df = pd.read_csv('designs.csv') # Pass rates at each stage print(f"Total designs: {len(df)}") print(f"pLDDT > 0.85: {(df['pLDDT'] > 0.85).mean():.1%}") print(f"ipTM > 0.50: {(df['ipTM'] > 0.50).mean():.1%}") print(f"scRMSD < 2.0: {(df['scRMSD'] < 2.0).mean():.1%}") print(f"All filters: {((df['pLDDT'] > 0.85) & (df['ipTM'] > 0.5) & (df['scRMSD'] < 2.0)).mean():.1%}") # Identify top issue if (df['pLDDT'] > 0.85).mean() < 0.1: print("ISSUE: Low pLDDT - check backbone or sequence quality") elif (df['ipTM'] > 0.50).mean() < 0.1: print("ISSUE: Low ipTM - check hotspots or interface geometry") elif (df['scRMSD'] < 2.0).mean() < 0.5: print("ISSUE: High scRMSD - sequences don't specify backbone")
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 20 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.