FPBench: A New Benchmark for Evaluating Foundation Potentials

less than 1 minute read

Published:

We introduce FPBench, an application-oriented benchmark for evaluating machine-learned foundation potentials (FPs) — models that promise to replace expensive first-principles calculations across atomistic simulations.

While FPs often report near-DFT accuracy on average energy and force errors, we found that these averages frequently fail to predict how a model actually performs on the tasks that matter in practice: force prediction for atomistic simulations, energy ranking of substitutional and vacancy orderings, and ion/vacancy migration.

FPBench introduces error-decomposition metrics that resolve force errors among highly accurate, large-error, and far-from-equilibrium atoms; relative-energy errors among competing orderings, phases, and compositions; and endpoint and along-path errors in ion migration — pinpointing exactly where and why a model breaks down, and providing targeted guidance for improving foundation potentials.

We have released an open benchmark, evaluation code, and a public leaderboard for the community to assess and compare foundation potentials.

Read the paper on arXiv, view the code, or explore the leaderboard.