Public computational databases for inorganic materials, chiefly the Materials Project, OQMD, and AFLOW, have been genuinely important for academic research in materials informatics. They provide DFT-calculated formation energies, band structures, elastic tensors, and derived properties for hundreds of thousands of inorganic compounds, all freely accessible via API. The scale of these datasets makes them useful training sources for machine-learned models, and they have enabled a significant body of research on property prediction and stability estimation across broad inorganic chemistry space.
For industrial battery cathode R&D, however, there is a persistent gap between what these databases contain and what the R&D team actually needs. Understanding this gap, and what it takes to close it for a specific cathode development program, is the difference between a materials informatics project that produces actionable candidate lists and one that produces academic output that does not translate to the lab.
What the Public Databases Contain and What They Do Not
The public databases are primarily populated with calculations performed at standard DFT conditions: 0 K, GGA or GGA+U exchange-correlation functional, idealized crystallographic unit cells with nominal stoichiometry, and no electrolyte or surface chemistry. The structures are typically equilibrium or near-equilibrium configurations, not the range of defect concentrations, dopant distributions, or surface terminations that arise in actual cathode materials under synthesis and cycling conditions.
Formation energies and convex hull stability predictions from these databases are reliable for comparing thermodynamic stability of bulk phases, and this is genuinely useful for identifying candidate compositions worth investigating. The gap emerges when you need to predict properties that depend on conditions not represented in the database: Li diffusivity at relevant operating temperatures, capacity fade under realistic cycling protocols, thermal stability under overcharge conditions, or the effect of dopant concentration on cation ordering.
For NMC cathodes specifically (LiNixMnyCo1-x-yO2 family), the composition space is three-dimensional and the electrochemically relevant properties are strongly sensitive to cation ordering patterns that cannot be directly read from the nominal composition. A database entry for a particular NMC stoichiometry represents one idealized ordered structure. The actual powder synthesized from that stoichiometry has a distribution of local cation environments, and that distribution depends on the synthesis conditions (temperature profile, cooling rate, precursor morphology) in ways that are not captured in the database entry.
The Calibration Problem
The central calibration challenge for battery cathode informatics is systematic bias between database predictions and experimental measurements under industrial synthesis conditions. This bias has several sources.
First, GGA exchange-correlation functionals underestimate the on-site Coulomb interactions for Mn, Ni, and Co d-electrons. This affects the predicted lithiation voltage and the relative phase stability of different oxidation-state configurations. GGA+U corrects part of this, but the U parameter value varies across the literature and is not consistently applied across database entries. A model trained on a mix of GGA and GGA+U calculations for transition metal oxide cathodes will have systematic inconsistencies in the training labels that manifest as composition-dependent prediction bias.
Second, standard DFT calculations in public databases use the PBE pseudopotential library with fixed parameters. Industrial synthesis targets often include dopants (Al, Ti, Mg, W, Nb at dopant concentrations below five percent) that are underrepresented or absent from the databases at the relevant concentration levels. The model performance on doped compositions typically degrades relative to its performance on the parent undoped compound.
Third, the synthesis conditions relevant to industrial cathode manufacturing (air calcination at 700 to 900 degrees Celsius, specific precursor morphology, controlled lithium excess) produce phases that may differ from the database's idealized structure in their defect chemistry and surface area characteristics. Mapping the DFT-calculated stability back to what actually forms under industrial conditions requires calibration data that links the DFT quantities to the experimentally measured quantities on samples made by the intended synthesis route.
Bridging the Gap: What Calibration Requires
Useful calibration for a specific cathode development program requires paired experimental and computational data on materials made by the intended synthesis route, not generic data from published literature. The published literature on cathode materials typically reports measurements on samples made in small academic labs without tight control on synthesis conditions. Those measurements may correlate with the DFT predictions in a general way but provide poor calibration for a specific industrial synthesis route.
The practical implication is that building a materials informatics model that reliably guides a cathode R&D program requires investment in a calibration dataset: a set of compositions made by the intended synthesis route, characterized by the relevant measurement techniques (electrochemical cycling, ICP-OES stoichiometry, XRD lattice parameter measurement, BET surface area), and paired with DFT calculations on the corresponding crystal structures. The size of that dataset depends on the composition complexity and the property targets, but a minimum of 30 to 60 compositions is typically needed before calibration offsets are reliable enough to trust for screening guidance.
This is not a criticism of the public databases, which serve a genuinely different purpose. It is a statement about what materials informatics requires to be actionable for an industrial cathode program, versus what it requires to be publishable as academic research. The two are different problems and different amounts of work.
What an Informatics-Guided Cathode Program Looks Like in Practice
The sequence we have found most productive for a cathode R&D program starts with using the public databases for initial candidate generation: identifying the composition families with thermodynamically favorable formation energies and reasonable predicted voltages. This gives a large initial pool at low compute cost. The second step is applying the MLIP screening engine to the full pool to rank by a broader set of stability criteria, reducing the pool to a manageable shortlist. The third step is DFT validation on the shortlist, producing the paired experimental data for calibration. The fourth step is using the calibrated model for the next iteration of candidate generation, now informed by what the synthesis actually produces.
The critical point is that the informatics model improves over iterations as calibration data accumulates. The first iteration uses a poorly calibrated model and produces a shortlist with a mediocre hit rate. The fifth iteration uses a model calibrated on data from the first four iterations and produces a substantially better shortlist. The value of an informatics approach to cathode development is front-loaded in terms of data investment and back-loaded in terms of discovery productivity.
For teams at the beginning of this process who want to discuss how to structure the calibration phase and what data collection investments are most efficient, the science page describes our model architecture and calibration approach in more detail. The applications page covers specific aspects of battery cathode screening. We are also happy to talk through the specific composition families and synthesis routes a team is working with via a direct conversation at contact.