tl;dr: MatFlowBench is a private benchmark for AI agents working on materials research workflows. It spans computational and experimental workflows, asking agents for bounded scientific results under a limited budget and precise physical constraints. Submissions are graded by task-specific verifiers.
Benchmarks for AI in materials science often use one-shot prompts and grade only the final answer, such as a predicted property or the correct option on a question. This leaves the multi-step workflow of choosing methods, running calculations and checking evidence outside the evaluation.
We introduce Materials Workflow Benchmark (MatFlowBench) to measure whether LLM agents can carry out multi-step computational and experimental workflows to determine material properties and select candidates. Agents run and analyze calculations in computational environments, or plan measurements and interpret characterization data in experimental settings. Each task fixes the material or candidate pool along with the constraints and supplies either calculation assumptions or a measurement archive. The agent decides how to proceed, and the verifier checks the evidence behind the answer, not just the answer itself.
Numerical methods in materials science approximate the physics, so settings that work for one material or property may be inadequate for another. More demanding settings can improve numerical accuracy, but they cost compute and time. Choosing them means deciding how much precision the question needs. A calculation can also finish without errors yet settle into the wrong magnetic state. Getting a property worth reporting usually takes several calculations, with checks along the way. Those checks matter in experimental work too, where each result shapes what happens next.
Computational workflows
Computational workflows use Density Functional Theory (DFT) and related methods to investigate materials properties. Agents work in prepared software environments that include Quantum ESPRESSO, Wannier90, and Phonopy for lattice dynamics. The environments also provide Python libraries for structure handling and numerical analysis, including NumPy, SciPy, Matplotlib, ASE, and pymatgen.
The agent organizes the calculation sequence itself. It prepares inputs and runs the simulations, then inspects each output to decide whether to move to the next step or repeat the calculation. Grading relies on the computational artifacts the agent produces rather than its own account of the results, which guards against reward hacking. Acceptance tolerances are set by the physics of each property, reflecting the accuracy the method can deliver for it.
Passing also takes scientific judgment, since the agent must decide whether its calculations are enough to answer the question under the stated physical assumptions. A Wannier model built for later analysis has to reproduce the DFT bands it came from. Competing states that differ only slightly in energy need tighter convergence than routine settings provide, and some tasks require the agent to look beyond the states a routine calculation would sample.
Experimental workflows
Experimental workflows cover measurement planning and materials selection. Agents access archived instrument measurements through a measurement service or analyze supplied characterization data directly. When an agent requests a measurement, the service returns stored results, allowing evaluation of experimental planning and analysis without operating laboratory equipment.
Materials selection requires deciding which samples to investigate and which measurements to request within a finite budget. Objectives include finding optical films that satisfy spectral requirements and selecting catalysts that retain activity after conditioning. Inexpensive screening can narrow the candidate pool before more costly measurements establish whether a promising sample qualifies, making the allocation of the measurement budget part of the scientific problem.
How we score tasks
Each task is scored pass or fail. For DFT tasks, the verifier does not rely on the number the agent reports. It reads the agent’s calculation output files and checks the result against the acceptance criteria. For budgeted materials searches, the verifier looks at which measurements the agent requested. An agent that recommends a candidate must have measured it, and an agent that reports no material qualifies must have measured enough of the pool to rule each one out. Guessing the right answer without those measurements still fails.
A passing score means the agent met those checks for a specific task. It is evidence of progress on a bounded research step, not proof that the agent discovered a material or could carry out a research program on its own.
Results and leaderboard
Across 500 computational rollouts covering 25 tasks, with two attempts per task for each of ten models, Opus achieved the highest pass rate at 64%, followed by Astra at 58%, Sonnet at 56%, and Sol at 46%. The remaining models scored between 0% and 18%. Astra passed the widest range of tasks, solving 18 of 25 at least once, while Opus, Sonnet, and Sol each solved 17. Three tasks received no passing submission from any model. These results show that leading agents can complete a substantial share of bounded materials workflows, but producing results that satisfy the required evidence checks remains inconsistent.
Computational workflows: Failures occurred both in executing calculations and in validating their results. Observed problems included Wannier models that did not reproduce their underlying DFT bands, incomplete searches for nodes or extrema, and incorrect property values despite successful solver execution. Some agents also declared completion after rejected commands or without producing required artifacts. These cases show that running a calculation is only part of the challenge: agents must check that it supports the final scientific claim.
Experimental workflows: Failures often involved selecting candidates that missed one or more requirements, or recommending candidates without obtaining the required confirmation measurements. Some agents moved to expensive measurements before cheaper screening provided a clear signal about which candidates warranted further investigation. Others exhausted their measurement allowance before confirming a suitable candidate or reached execution limits before completing their analysis. These cases highlight the difficulty of allocating a limited budget across screening, selection, and confirmation.
Future Steps
The first question for MatFlowBench is whether an agent can finish a defined piece of materials work with the evidence that piece requires. In an experimental search, that evidence is the set of measurements that justify the recommendation, while a calculation has to pass the checks attached to the property it claims. Knowing which of these steps an agent can complete gives researchers a concrete starting point for deciding where it can contribute to a larger discovery workflow. We will publish a paper with the full results, including a breakdown by task family, a detailed failure analysis and a discussion of what the verifiers can and cannot establish.
https://matflowbench.com/





