FPBench: A New Benchmark for Evaluating Foundation Potentials
Published:
We introduce FPBench, an application-oriented benchmark for evaluating machine-learned foundation potentials (FPs) — models that promise to replace expensive first-principles calculations across atomistic simulations.
While FPs often report near-DFT accuracy on average energy and force errors, we found that these averages frequently fail to predict how a model actually performs on the tasks that matter in practice: force prediction for atomistic simulations, energy ranking of substitutional and vacancy orderings, and ion/vacancy migration.
FPBench introduces error-decomposition metrics that resolve force errors among highly accurate, large-error, and far-from-equilibrium atoms; relative-energy errors among competing orderings, phases, and compositions; and endpoint and along-path errors in ion migration — pinpointing exactly where and why a model breaks down, and providing targeted guidance for improving foundation potentials.
We have released an open benchmark, evaluation code, and a public leaderboard for the community to assess and compare foundation potentials.
Read the paper on arXiv, view the code, or explore the leaderboard.
