PREMVAL

Benchmarking generative models of protein conformational ensembles against all-atom molecular dynamics.

Leaderboard

Model Data RMWD ↓ RMSF r ↑ MD-PCA W₂ ↓ Weak J ↑ Trans J ↑ Time/chain ↓ N
AlphaFlow-MD held-out (weak) 2.373.81 0.8640.814 1.462.12 0.6200.602 0.4130.398 n/a 82
AlphaFlow-MD (distilled) held-out (weak) 3.204.76 0.8320.783 1.682.52 0.5130.493 0.2850.268 n/a 82
ESMFlow-MD held-out (weak) 3.405.98 0.7670.704 1.542.49 0.5520.527 0.3500.328 n/a 82
ESMFlow-MD (distilled) held-out (weak) 4.186.43 0.7700.699 1.932.70 0.4740.454 0.2990.275 365s 82
BioEmu held-out 5.186.17 0.7940.751 1.592.13 0.3330.349 0.0660.074 718s 81
ESMDiff uncertain 5.797.74 0.6930.664 1.832.44 0.4850.462 0.2580.246 237s 82
Reference floors (baselines, not generative models)
ANM baseline (Cα) held-out 2.823.50 0.7760.731 1.532.07 0.3120.295 0.1750.174 1s 82
Static structure held-out 3.063.71 n/a 2.052.61 0.0000.000 0.0000.000 2s 82

Metric cells read median mean over the split; rows are ranked by median RMWD. Click any column to rank by it instead.

Reference floors are trivially reproducible baselines, scored by the identical code but not ranked against the models: ANM is a training-free elastic network sampled along its slowest normal modes, and Static is that same structure with no motion at all. Both are handed the ATLAS input structure, so their mean structure is right by construction while the models above start from sequence. Static's RMSF r is n/a because the correlation is undefined for an ensemble that does not move, and its contact Jaccards are 0 by construction: a contact that never fluctuates is never weak and never transient.

RMWD root-mean Wasserstein distance (Å), lower is better · RMSF r per-residue flexibility correlation, higher is better · MD-PCA W₂ distance in the reference's principal-motion plane (Å), lower is better · Weak/Trans J contact-set overlap for breaking and forming contacts, higher is better · Time/chain median inference time per chain (s), lower is faster · N chains scored

Data how strong is the ATLAS held-out guarantee? · held-out never trained on ATLAS/MD · held-out (weak) trained on ATLAS train, temporal split only · uncertain unknown

Browse chains →

Scored on the ATLAS test split, against all-atom MD trajectories from the ATLAS dataset.

References

  1. Jing, B., Berger, B., & Jaakkola, T. (2024). AlphaFold Meets Flow Matching for Generating Protein Ensembles. International Conference on Machine Learning (ICML). arXiv:2402.04845. [link]
  2. Atilgan, A. R., Durell, S. R., Jernigan, R. L., Demirel, M. C., Keskin, O., & Bahar, I. (2001). Anisotropy of fluctuation dynamics of proteins with an elastic network model. Biophysical Journal, 80(1), 505-515. Sampled with ProDy (Bakan et al., Bioinformatics 2011). [link]
  3. Lewis, S., et al. (2024). Scalable emulation of protein equilibrium ensembles with generative deep learning. bioRxiv 2024.12.05.626885. [link]
  4. Lu, J., Chen, X., Lu, S. Z., Shi, C., Guo, H., Bengio, Y., & Tang, J. (2025). Structure Language Models for Protein Conformation Generation. International Conference on Learning Representations (ICLR). arXiv:2410.18403. [link]
  5. Vander Meersche, Y., Cretin, G., Gheeraert, A., Gelly, J.-C., & Galochkina, T. (2024). ATLAS: protein flexibility description from atomistic molecular dynamics simulations. Nucleic Acids Research, 52(D1), D384-D392. [link]