Chan Zuckerberg Biohub (CZ Biohub) has published the ESM Atlas, a database of approximately 1.1 billion predicted protein structures, together with the ESMFold2 model and its weights under an open-source, no-commercial-restrictions licence. According to the organisation, the release and an accompanying paper were published the same day in Nature.
Size and scope
- The ESM Atlas contains about 1.1 billion predicted structures. The announcement notes this is roughly five times the size of the DeepMind AlphaFold database, which it cites as having 200 million entries.
- CZ Biohub states that ESMFold2 was trained on "billions" of metagenomic sequences, a training set the announcement says was not used by AlphaFold.
Reported performance and uses
CZ Biohub reports that designs generated with ESMFold2 and tested in the laboratory achieved high hit rates against cancer and immune system targets. The organisation emphasises that the model and the dataset are available for unrestricted use, with no commercial-use prohibitions.
The open vs closed model debate
The release is presented in contrast to recent trends where some successor models and their weights have remained closed. The announcement explicitly refers to AlphaFold3 weights as closed-source. CZ Biohub frames the ESM Atlas and ESMFold2 release as an example of how open tools and larger datasets can shift where competitive advantage lies: from proprietary base models toward whoever builds services, applications, or downstream research on top of accessible foundations.
Why it matters
A large, openly available protein-structure database combined with an open prediction model can accelerate biological research and lower barriers for drug discovery and biotechnological innovation. By releasing a substantially larger dataset and removing usage restrictions, CZ Biohub's announcement raises questions about how value and control in structure prediction will be distributed going forward.
Summary
The Chan Zuckerberg Biohub has made available 1.1 billion predicted protein structures and an open ESMFold2 model without commercial restrictions. The organisation positions the move as significant both for the scale of the dataset and for the implications of keeping foundational models and data fully open.



