Benchmarking generative models of protein conformational ensembles against all-atom molecular dynamics.
| Model | Data | RMWD ↓ | RMSF r ↑ | MD-PCA W₂ ↓ | Weak J ↑ | Trans J ↑ | Time/chain ↓ | N |
|---|---|---|---|---|---|---|---|---|
| AlphaFlow-MD | held-out (weak) | 2.373.81 | 0.8640.814 | 1.462.12 | 0.6200.602 | 0.4130.398 | n/a | 82 |
| AlphaFlow-MD (distilled) | held-out (weak) | 3.204.76 | 0.8320.783 | 1.682.52 | 0.5130.493 | 0.2850.268 | n/a | 82 |
| ESMFlow-MD | held-out (weak) | 3.405.98 | 0.7670.704 | 1.542.49 | 0.5520.527 | 0.3500.328 | n/a | 82 |
| ESMFlow-MD (distilled) | held-out (weak) | 4.186.43 | 0.7700.699 | 1.932.70 | 0.4740.454 | 0.2990.275 | 365s | 82 |
| BioEmu | held-out | 5.186.17 | 0.7940.751 | 1.592.13 | 0.3330.349 | 0.0660.074 | 718s | 81 |
| ESMDiff | uncertain | 5.797.74 | 0.6930.664 | 1.832.44 | 0.4850.462 | 0.2580.246 | 237s | 82 |
| Reference floors (baselines, not generative models) | ||||||||
| ANM baseline (Cα) | held-out | 2.823.50 | 0.7760.731 | 1.532.07 | 0.3120.295 | 0.1750.174 | 1s | 82 |
| Static structure | held-out | 3.063.71 | n/a | 2.052.61 | 0.0000.000 | 0.0000.000 | 2s | 82 |
Metric cells read median mean over the split; rows are ranked by median RMWD. Click any column to rank by it instead.
Reference floors are trivially reproducible baselines, scored by the identical code but not ranked against the models: ANM is a training-free elastic network sampled along its slowest normal modes, and Static is that same structure with no motion at all. Both are handed the ATLAS input structure, so their mean structure is right by construction while the models above start from sequence. Static's RMSF r is n/a because the correlation is undefined for an ensemble that does not move, and its contact Jaccards are 0 by construction: a contact that never fluctuates is never weak and never transient.
RMWD root-mean Wasserstein distance (Å), lower is better · RMSF r per-residue flexibility correlation, higher is better · MD-PCA W₂ distance in the reference's principal-motion plane (Å), lower is better · Weak/Trans J contact-set overlap for breaking and forming contacts, higher is better · Time/chain median inference time per chain (s), lower is faster · N chains scored
Data how strong is the ATLAS held-out guarantee? · held-out never trained on ATLAS/MD · held-out (weak) trained on ATLAS train, temporal split only · uncertain unknown
Browse chains →Scored on the ATLAS test split, against all-atom MD trajectories from the ATLAS dataset.