AXI Strain 1.0: Evaluating Frontier AI Reasoning in Materials Science

August 18, 2026

A multimodal benchmark for measuring how AI systems see, retrieve, and reason in materials science.

AI models are becoming increasingly fluent in science. But scientific intelligence is more than knowing the right facts—it requires understanding physical systems, interpreting evidence, and reasoning across different forms of scientific information. AXI Strain 1.0 is our evaluation for measuring these capabilities in materials science. Built around expert-authored problems, it focuses on materials cognition: whether frontier AI systems can reason through challenging scientific problems rather than simply recognize or retrieve familiar information. Models are evaluated in a single-shot setting, with each problem issued as an independent query and no opportunity for iterative prompting, feedback, or self-correction across turns.

From Materials Vision to Materials Cognition: Materials science is inherently multimodal. Scientists reason across structures, microscopy images, diffraction patterns, spectra, phase diagrams, plots, and numerical data. AXI Strain 1.0 incorporates this complexity directly into its design. A distinctive feature of the evaluation is its ability to compare performance across different representations of the same scientific information. This allows us to examine materials vision—what a model can extract from scientific figures—alongside materials cognition—what it can infer once the relevant information is available.

Separating Retrieval from Reasoning: Knowing more does not necessarily mean reasoning better. AXI Strain 1.0 can be evaluated under different information-access settings, including with and without external retrieval. Comparing these conditions is designed to distinguish limitations in access to scientific knowledge from limitations in reasoning with that knowledge. In the current leaderboard, GPT 5.5 scored 28.76% with Search OFF and 27.55% with Search ON; Claude Opus 4.8 scored 23.39% and 21.37%, respectively; Claude Opus 4.7 scored 22.58% in both settings; and Claude Fable 5 scored 20.43% with Search OFF and 21.37% with Search ON. Across these models, enabling web search therefore produced relatively small and inconsistent changes in performance, rather than a uniform improvement. Together, these dimensions make AXI Strain 1.0 more than a single benchmark score. It provides a way to examine how frontier models see, retrieve, and reason in materials science.

Example: A model was shown an experimental materials-science figure along with a relevant experimental explanation and asked to determine a set of properties that are not explicitly stated but can be derived from the data through a series of high-level reasoning steps. In this example, an electrochemistry data set is provided together with the experimental conditions and fitted parameters. The model must interpret the results, evaluate contributing factors, calculate multiple values, apply several scientific reasoning steps, and refit the data to determine the requested findings. The same problem can also be presented with the figure values supplied directly as text. Comparing these two representations helps distinguish limitations in scientific-figure interpretation from limitations in downstream reasoning. In this way, AXI Strain 1.0 is designed to separate materials vision from broader materials cognition without reducing performance to a single aggregate score.