Open-source models for biology and chemistry

These models work with proteins, molecules and genomic data: they predict structures and properties of compounds and support drug discovery. They are tools for labs, biotech companies and research teams. Check the license, since some models restrict commercial use, as well as hardware needs and input data formats.

5 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Biology and chemistry2022–2026

OpenFold / OpenFold3

AlQuraishi Lab (Columbia University) and the OpenFold consortium · USA

A fully open reproduction of AlphaFold 2 and then AlphaFold 3 under Apache 2.0, with training data. OpenFold3 predicts complexes of proteins, nucleic acids and ligands.

  • Predicting structures of proteins and ligand complexes
  • Fine-tuning on the company's own data (training code is open)
  • An in-house structural analysis service without sending data outside
Sizes
a single set of weights per version
Hardware
from: 1 GPU
Commercial use allowedDetails
Biology and chemistry2022–2026

ESM (ESM-2, ESM3, ESM C, ESMFold2)

EvolutionaryScale / Chan Zuckerberg Biohub (ESM-2 — Meta AI) · USA

Protein language models: they understand amino acid sequences, predict structure (ESMFold2) and help with protein design. Since 2026 all open versions are under MIT.

  • Protein embeddings for predicting properties (stability, solubility)
  • Predicting 3D structures of proteins and complexes
  • Screening enzyme and antibody design candidates before lab work
Sizes
8M – 15B (ESM-2), 300M – 6B (ESM C), 1.4B (open ESM3)
Hardware
from: Laptop
Commercial use allowedDetails
Biology and chemistry2024–2026

Arc Institute Evo / Evo 2

Arc Institute (with Together AI, Stanford, NVIDIA) · USA

DNA language models with context up to a million nucleotides: they assess the impact of mutations, annotate genomes and generate sequences. Evo 2 is trained on genomes from all domains of life.

  • Assessing the likely harmfulness of genetic variants for research
  • Annotating genomes of microorganisms and plants
  • Finding promising sequences in breeding and synthetic biology
Sizes
1B – 40B
Hardware
from: 1 GPU
Commercial use allowedDetails
Biology and chemistry2024–2025

Boltz (Boltz-1, Boltz-2, BoltzGen)

MIT (Jameel Clinic) and Recursion · USA

An open MIT-licensed alternative to AlphaFold 3: predicts structures of protein, DNA and small-molecule complexes; Boltz-2 estimates binding strength, BoltzGen designs new binding proteins.

  • Predicting how a candidate molecule binds to a target protein
  • Ranking compounds by predicted binding strength before synthesis
  • Designing binder proteins for a given target
Sizes
checkpoints of about 2 GB
Hardware
from: 1 GPU
Commercial use allowedDetails
Biology and chemistry2024

AlphaFold 3

Google DeepMind and Isomorphic Labs · UK

The reference model for the structure of biomolecules and their complexes. Weights are provided for non-commercial research only; companies need commercial access via Google Cloud or open alternatives (Boltz, OpenFold3).

  • Academic research on protein and complex structures
  • Benchmarking open alternatives against the reference on your own targets
Sizes
a single set of weights
Hardware
from: 1 GPU
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment