# FAIR Chemistry: Complete Documentation Canonical site: https://facebookresearch.github.io/fairchem/ --- Source: `docs/index.md` +++ {"class": "col-page-inset"} ```{image} assets/fair-chemistry-logo-light-mode.png :alt: FAIR Chemistry :width: 600px :align: center :class: dark:hidden ``` ```{image} assets/fair-chemistry-logo-dark-mode.png :alt: FAIR Chemistry :width: 600px :align: center :class: hidden dark:block ``` ## Open data and universal models for atomic systems FAIR Chemistry develops open datasets and machine-learning models for molecules, materials, and catalysts. **UMA** brings these domains together in one pretrained model while preserving each dataset's level of theory. ```{code-block} bash :filename: Install pip install fairchem-core ``` {button}`Understand FAIR Chemistry → <./core/introduction.md>` {button}`Run your first calculation → <./core/quickstart.md>` +++ {"class": "col-page-inset"} ## From open datasets to UMA UMA learns from more than 500 million density functional theory calculations. At inference time, you choose a task that matches your chemistry domain and the corresponding level of theory. :::::{grid} 1 2 3 5 ::::{card} Heterogeneous catalysis :link: catalysts/datasets/summary.md ```{image} assets/icons/catalysis.svg :alt: Catalysis :width: 60px :align: center ``` Surface reactions, adsorption, and catalyst design. +++ **Datasets:** OC20, OC22, OC25
**UMA tasks:** `oc20`, `oc22`, `oc25` :::: ::::{card} Inorganic materials :link: inorganic_materials/datasets/summary.md ```{image} assets/icons/inorganic.svg :alt: Inorganic materials :width: 60px :align: center ``` Bulk materials, phonons, and elastic properties. +++ **Dataset:** OMat24
**UMA task:** `omat` :::: ::::{card} Molecules & polymers :link: molecules/datasets/summary.md ```{image} assets/icons/molecules.svg :alt: Molecules and polymers :width: 60px :align: center ``` Conformers, reactions, and electronic properties. +++ **Dataset:** OMol25
**UMA task:** `omol` :::: ::::{card} Molecular crystals :link: molecules/datasets/omc25.md ```{image} assets/icons/molecular-crystals.svg :alt: Molecular crystals :width: 60px :align: center ``` Packed organic molecules in crystal structures. +++ **Dataset:** OMC25
**UMA task:** `omc` :::: ::::{card} MOFs for direct air capture :link: dac/datasets/summary.md ```{image} assets/icons/mofs-dac.svg :alt: Metal-organic frameworks for direct air capture :width: 60px :align: center ``` CO₂ and H₂O adsorption in metal-organic frameworks. +++ **Datasets:** ODAC23, ODAC25
**UMA task:** `odac` :::: ::::: :::{card} UMA: one model family, multiple levels of theory UMA uses a learned task embedding to apply one pretrained model across these domains without mixing their reference methods. Start with [`uma-s-1p2p1`](./core/uma.md), the fastest current UMA model with state-of-the-art accuracy on most supported benchmarks. ::: +++ {"class": "col-page-inset"} ## See what UMA can do ::::{grid} 1 2 2 2 :::{card} Interactive UMA playground :link: https://aidemos.atmeta.com/uma?view=playground Explore atomistic simulations without installing anything. Animation of the interactive UMA playground +++ [Open the playground →](https://aidemos.atmeta.com/uma?view=playground) ::: :::{card} Faster molecular simulation :link: ./core/uma_changelog.md ```{image} https://raw.githubusercontent.com/facebookresearch/fairchem/main/benchmarks/omol/omol_force_mae_vs_speed.png :alt: Force accuracy versus molecular-dynamics runtime for UMA and other models. :width: 100% :align: center ``` +++ [Read the benchmark details →](https://github.com/facebookresearch/fairchem/tree/main/benchmarks/omol) ::: :::: +++ {"class": "col-page-inset"} ## Choose your next step :::::{grid} 1 2 3 3 ::::{card} Install FAIR Chemistry :link: ./core/install.md Set up `fairchem-core` and request access to UMA checkpoints. +++ [Install →](./core/install.md) :::: ::::{card} Hello World :link: ./core/quickstart.md Run a molecular calculation and relax an inorganic crystal. +++ [Get started →](./core/quickstart.md) :::: ::::{card} Explore UMA :link: ./core/introduction.md#what-you-can-do-with-uma Find the right task and workflow for your scientific problem. +++ [Explore capabilities →](./core/introduction.md#what-you-can-do-with-uma) :::: ::::{card} Model guide :link: ./core/uma.md Understand UMA tasks, inputs, architecture, and limitations. +++ [Read the guide →](./core/uma.md) :::: ::::{card} Common workflows :link: ./core/common_tasks/summary.md Scale from ASE calculations to training and batched inference. +++ [Browse workflows →](./core/common_tasks/summary.md) :::: ::::{card} Learning resources :link: ./videos.md Watch introductory videos and technical presentations. +++ [Start learning →](./videos.md) :::: ::::: Copyright © Meta Platforms, Inc | [Terms of Use](https://opensource.fb.com/legal/terms) | [Privacy Policy](https://opensource.fb.com/legal/privacy) --- Source: `docs/core/install.md` # Installation & License ## Installation :::{warning} FAIRChem V2 is a major breaking change from V1 and is not compatible with previous pretrained models. If you need the old V1 code, install version 1 with `pip install fairchem-core==1.10`. ::: To install `fairchem-core` you will need to setup the `fairchem-core` environment. We support either pip or uv. Conda is no longer supported and has also been dropped by pytorch itself. Note you can still create environments with conda and use pip to install the packages. :::{tip} We recommend installing fairchem inside a virtual environment instead of directly onto your system. ::: ### Step 1: Create a virtual environment ```bash virtualenv -p python3.12 fairchem source fairchem/bin/activate ``` ### Step 2: Install the package ```bash pip install fairchem-core ``` :::{admonition} For developers contributing to fairchem :class: dropdown Clone the repo and install in editable mode: ```bash git clone git@github.com:facebookresearch/fairchem.git cd fairchem pip install -e packages/fairchem-core[dev] ``` ::: :::{note} In V2, we removed all dependencies on 3rd party libraries such as torch-geometric, pyg, torch-scatter, torch-sparse etc that made installation difficult. So no additional steps are required! ::: ## Subpackages In addition to `fairchem-core`, there are related packages for specialized tasks or applications. Each can be installed with `pip` or `uv` just like `fairchem-core`: ### Data Packages Utilities for generating input configurations and working with specific datasets: ::::{grid} 1 2 3 3 :::{card} fairchem-data-oc Code for generating adsorbate-catalyst input configurations ::: :::{card} fairchem-data-omat Code for generating OMat24 input configurations and VASP input sets ::: :::{card} fairchem-data-omc Code for generating OMC (Molecular Crystals) VASP inputs ::: :::{card} fairchem-data-omol Code for generating OMOL input configurations ::: :::{card} fairchem-data-odac Code for ODAC MOF configurations and VASP input sets for direct air capture ::: :::: ### Application Packages Higher-level applications built on top of FAIRChem models: ::::{grid} 1 2 3 3 :::{card} fairchem-applications-adsorbml Module for calculating minimum adsorption energies ::: :::{card} fairchem-applications-cattsunami Accelerating transition state energy calculations with pre-trained GNNs ::: :::{card} fairchem-applications-fastcsp Accelerated molecular crystal structure prediction with UMA ::: :::{card} fairchem-applications-ocx Bridging experiments to computational models ::: :::: ### Integration & Demo Packages Tools for integrating with other software or demo APIs: ::::{grid} 1 2 3 3 :::{card} fairchem-lammps Use FAIRChem models with LAMMPS for large-scale MD simulations ::: :::{card} fairchem-demo-ocpapi Python client library for the Open Catalyst Demo API ::: :::: ## Access to gated models on HuggingFace To access gated models like UMA, you need to get a HuggingFace account and request access to the UMA models. :::{admonition} HuggingFace Setup Steps :class: tip 1. Get and login to your HuggingFace account 2. Request access to 3. Create a HuggingFace token at with the permission "Read access to contents of all public gated repos you can access" 4. Add the token as an environment variable using `huggingface-cli login` or by setting the `HF_TOKEN` environment variable ::: ## License ### Repository software The software in this repo is licensed under an MIT license unless otherwise specified. :::{admonition} MIT License :class: dropdown ```text MIT License Copyright (c) Meta, Inc. and its affiliates. Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. ``` ::: ### Terms of use & privacy policy Please read the following [Terms of Use](https://opensource.fb.com/legal/terms) and [Privacy Policy](https://opensource.fb.com/legal/privacy) covering usage of `fairchem` software and models. ### Model checkpoints and datasets Please check each dataset and model for their own licenses. --- Source: `docs/core/quickstart.md` # Hello World with UMA This tutorial takes you from an ASE structure to a UMA prediction. You will load one pretrained model, use it for two chemistry domains, and learn how to substitute your own structure. :::{note} Before you start Complete the [installation and Hugging Face access steps](./install.md). The first run downloads the gated `uma-s-1p2p1` checkpoint. ::: ## The four-step workflow Every basic calculation follows the same pattern: 1. Represent the system as an ASE `Atoms` object. 2. Load a pretrained UMA model. 3. Choose the task matching the system and attach a `FAIRChemCalculator`. 4. Ask ASE for energies and forces or use an ASE simulation method. ## Load UMA once ```{code-cell} ipython3 from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", device="cuda", ) ``` The predictor contains the shared UMA model. The calculator created for each system supplies the domain-specific task. ## Example 1: calculate a molecular spin gap For the `omol` task, set the molecule's total charge and spin multiplicity in `atoms.info`. Here we compare singlet and triplet states of CH₂. ```{code-cell} ipython3 from ase.build import molecule singlet = molecule("CH2_s1A1d") singlet.info.update({"charge": 0, "spin": 1}) singlet.calc = FAIRChemCalculator(predictor, task_name="omol") triplet = molecule("CH2_s3B1d") triplet.info.update({"charge": 0, "spin": 3}) triplet.calc = FAIRChemCalculator(predictor, task_name="omol") spin_gap = triplet.get_potential_energy() - singlet.get_potential_energy() print(f"Triplet-singlet energy difference: {spin_gap:.3f} eV") ``` ## Example 2: relax an inorganic crystal The `omat` task predicts stress as well as energy and forces, so ASE can relax both the atoms and the unit cell. ```{code-cell} ipython3 from ase.build import bulk from ase.filters import FrechetCellFilter from ase.optimize import FIRE iron = bulk("Fe") iron.calc = FAIRChemCalculator(predictor, task_name="omat") optimizer = FIRE(FrechetCellFilter(iron), logfile=None) optimizer.run(fmax=0.05, steps=100) print(f"Relaxed energy: {iron.get_potential_energy():.3f} eV") print("Relaxed cell (Å):") print(iron.cell) ``` ## Try your own structure ASE reads many common chemistry file formats, including XYZ, CIF, POSCAR, and trajectory files. Replace the filename and task below with values appropriate for your system. ```{code-cell} ipython3 :tags: [skip-execution] from ase.io import read atoms = read("my-structure.xyz") # Required for molecules evaluated with the omol task. atoms.info.update({"charge": 0, "spin": 1}) atoms.calc = FAIRChemCalculator(predictor, task_name="omol") energy = atoms.get_potential_energy() forces = atoms.get_forces() ``` Choose the task by scientific domain, not merely by which task accepts the structure: | Domain | Task | Start here | | --- | --- | --- | | Molecules and polymers | `omol` | [OMol25](../molecules/datasets/omol25.md) | | Inorganic materials | `omat` | [OMat24](../inorganic_materials/datasets/omat24.md) | | Heterogeneous catalysis | `oc20`, `oc22`, or `oc25` | [Catalysis datasets](../catalysts/datasets/summary.md) | | Molecular crystals | `omc` | [OMC25](../molecules/datasets/omc25.md) | | MOFs and direct air capture | `odac` | [ODAC datasets](../dac/datasets/summary.md) | :::{warning} Different tasks reproduce different levels of theory. Do not compare energies across tasks as though they came from one reference calculation. ::: ## Next steps - [Explore UMA's capabilities](./introduction.md#what-you-can-do-with-uma) by domain. - Review task limitations in the [UMA model guide](./uma.md). - Learn about inference settings in the [ASE calculator guide](./common_tasks/ase_calculator.md). - Try the [playground](https://aidemos.atmeta.com/uma?view=playground). --- Source: `docs/core/introduction.md` # Introduction to FAIR Chemistry FAIR Chemistry is Meta FAIR's open ecosystem for machine learning in atomistic simulation. It brings together large quantum-chemistry datasets, pretrained models, and tools that connect those models to familiar simulation workflows. The central idea is simple: expensive density functional theory (DFT) calculations can be used to train machine-learned interatomic potentials. Once trained, those models estimate energies and forces much faster, making it possible to explore more structures and longer trajectories before confirming the most important results with higher-fidelity methods. ## How the pieces fit together ```{image} ../assets/uma-diagram-light-mode.png :alt: UMA connects several chemistry domains through one universal model. :width: 700px :align: center :class: dark:hidden ``` ```{image} ../assets/uma-diagram-dark-mode.png :alt: UMA connects several chemistry domains through one universal model. :width: 700px :align: center :class: hidden dark:block ``` FAIR Chemistry provides three connected pieces: 1. **Open datasets** contain atomistic structures and DFT labels for distinct chemistry domains. 2. **UMA** is a family of Universal Models for Atoms pretrained across those domains. 3. **`fairchem`** connects UMA to tools such as ASE, LAMMPS, and quacc for calculations and simulations. ## One model, several tasks Each dataset was calculated with a particular scientific method and set of approximations. UMA preserves those distinctions through a **task** input. You select the task that matches your system and the level of theory you want UMA to emulate. | Domain | Representative training data | UMA task | Example uses | | --- | --- | --- | --- | | Organic molecules and polymers | OMol25 | `omol` | Conformers, reactions, molecular dynamics | | Inorganic materials | OMat24 | `omat` | Relaxations, phonons, elastic properties | | Heterogeneous catalysts | OC20, OC22, OC25 | `oc20`, `oc22`, `oc25` | Adsorption, surfaces, reaction pathways | | Molecular crystals | OMC25 | `omc` | Crystal packing and polymorph ranking | | MOFs and direct air capture | ODAC23 | `odac` | CO₂ and H₂O adsorption | Task selection matters because predictions from different tasks generally represent different DFT levels of theory. They should not be mixed in one energy comparison without careful validation. See the [UMA model guide](./uma.md) for task-specific caveats. ## What you can do with UMA UMA provides energies, forces, and—for supported periodic tasks—stresses through the standard ASE calculator interface. Those predictions can drive many atomistic workflows without changing models as you move between domains. :::{tip} Start with [`uma-s-1p2p1`](./uma.md) and choose the task that matches the dataset and level of theory relevant to your system. ::: ::::{grid} 1 2 2 2 :::{card} Molecules and polymers · `omol` Calculate conformer energies, spin gaps, vibrations, and molecular dynamics. [Explore molecular data →](../molecules/datasets/summary.md) ::: :::{card} Inorganic materials · `omat` Relax atomic positions and cells, calculate elastic properties, and construct phonon spectra. [Explore materials tutorials →](../inorganic_materials/examples_tutorials/summary.md) ::: :::{card} Heterogeneous catalysts · `oc20`, `oc22`, `oc25` Study adsorption, surface stability, reaction thermochemistry, and transition states with the task appropriate to the interface. [Explore catalysis tutorials →](../catalysts/examples_tutorials/summary.md) ::: :::{card} Molecular crystals · `omc` Score periodic molecular crystals and support crystal-structure prediction workflows. [Explore OMC25 →](../molecules/datasets/omc25.md) ::: :::{card} MOFs and direct air capture · `odac` Estimate adsorption energies and study framework deformation for CO₂ and H₂O. [Explore the adsorption tutorial →](../dac/examples_tutorials/adsorption_energy.md) ::: :::{card} Scaled simulation Use batched inference, multiple GPUs, LAMMPS, or workflow engines for larger and more numerous simulations. [Browse common workflows →](./common_tasks/summary.md) ::: :::: ## A typical workflow 1. Install `fairchem-core` and obtain access to the gated UMA repository. 2. Create or load an atomic structure as an ASE `Atoms` object. 3. Load `uma-s-1p2p1` and select the appropriate task. 4. Attach a `FAIRChemCalculator` to the structure. 5. Run an energy, force, relaxation, dynamics, or downstream-property calculation. 6. Inspect the structure and validate important conclusions against reference data or higher-fidelity calculations. :::{important} UMA accelerates atomistic modeling; it does not remove the need to check the training domain, level of theory, physical constraints, and uncertainty for your application. ::: ## Try UMA without writing code The [Meta AI Demo Lab UMA playground](https://aidemos.atmeta.com/uma?view=playground) is the recommended browser-based experience. Use it to manipulate structures and build intuition before setting up a local workflow. The separate [guided UMA demo](https://facebook-fairchem-uma-demo.hf.space/) contains additional worked examples and is maintained outside this repository. ## Where to go next ::::{grid} 1 2 2 2 :::{card} Install :link: ./install.md Prepare Python and Hugging Face access. ::: :::{card} Hello World :link: ./quickstart.md Run two small end-to-end calculations. ::: :::{card} Common workflows :link: ./common_tasks/summary.md Scale from ASE calculations to training and batched inference. ::: :::{card} Playground :link: https://aidemos.atmeta.com/uma?view=playground Explore UMA interactively in a browser. ::: :::: --- Source: `docs/core/uma.md` # UMA Models [UMA](https://ai.meta.com/research/publications/uma-a-family-of-universal-models-for-atoms/) is an equivariant GNN that leverages a novel technique called Mixture of Linear Experts (MoLE) to give it the capacity to learn the largest multi-modal dataset to date (500M DFT examples), while preserving energy conservation and inference speed. Even a 6M active parameter (145M total) UMA model is able to achieve SOTA accuracy on a wide range of domains such as materials, molecules and catalysis. [arXiv Paper](https://arxiv.org/abs/2506.23971) For most applications, start with `uma-s-1p2p1`. It is the latest patch release of the small UMA 1.2 model and offers the best default balance of speed, accuracy, and memory use. ![UMA model architecture](uma.svg "UMA model architecture") ## The UMA Mixture-of-Linear-Experts routing function :::{note} The Mixture-of-Linear-Expert (MoLE) architecture is key to UMA's performance. It enables very high parameter count with fast inference speeds by using a single output head with dynamic routing. ::: The UMA model uses a Mixture-of-Linear-Expert (MoLE) architecture to achieve very high parameter count with fast inference speeds with a single output head. In order to route the model to the correct set parameters, the model must be given a set of inputs. The following information is required for the input to the model: ::::{grid} 2 :::{card} Task The task specifies the level of theory/DFT calculations to emulate: `omol`, `oc20`, `omat`, `odac`, `omc` ::: :::{card} Charge Total known charge of the system (only used for `omol` task, defaults to 0) ::: :::{card} Spin Total spin multiplicity of the system (only used for `omol` task, defaults to 1) ::: :::{card} Elemental Composition The unordered total elemental composition. Each element has an atom embedding and the composition embedding is the mean over all atom embeddings. ::: :::: ### The UMA task UMA is trained on 5 different DFT datasets with different levels of theory. An UMA **task** refers to a specific level of theory associated with that DFT dataset. UMA learns an embedding for the given **task**. Thus at inference time, the user must specify which one of the 5 embeddings they want to use to produce an output with the DFT level of theory they want. :::{tip} Choose your task based on the application domain. Each task corresponds to a specific DFT level of theory optimized for that domain. ::: | Task | Dataset | DFT Level of Theory | Relevant Applications | Usage Notes | |------|---------|---------------------|----------------------|-------------| | **omol** | [OMol25](https://arxiv.org/abs/2505.08762) | wB97M-V/def2-TZVPD as implemented in ORCA6, including non-local dispersion. All solvation should be explicit. | Biology, organic chemistry, protein folding, small-molecule pharmaceuticals, organic liquid properties, homogeneous catalysis, polymers | Requires total charge and spin multiplicity. If you don't know what these are, you should be very careful if modeling charged or open-shell systems. This can be used to study radical chemistry or understand the impact of magnetic states on the structure of a molecule. All training data is aperiodic, so any periodic systems should be treated with some caution. Probably won't work well for inorganic materials. | | **omc** | [OMC25](https://arxiv.org/abs/2508.02651) | PBE+D3 as implemented in VASP. | Pharmaceutical packaging, bio-inspired materials, organic electronics, organic LEDs | UMA has not seen varying charge or spin multiplicity for the OMC task, and expects total_charge=0 and spin multiplicity=0 as model inputs. | | **omat** | [OMat24](https://arxiv.org/abs/2410.12771) | PBE/PBE+U as implemented in VASP using Materials Project suggested settings, except with VASP 54 pseudopotentials. No dispersion. | Inorganic materials discovery, solar photovoltaics, advanced alloys, superconductors, electronic materials, optical materials | UMA has not seen varying charge or spin multiplicity for the OMat task, and expects total_charge=0 and spin multiplicity=0 as model inputs. Spin polarization effects are included, but you can't select the magnetic state. Further, OMat24 did not fully sample possible spin states in the training data. | | **oc20** | [OC20*](https://arxiv.org/abs/2010.09990) | RPBE as implemented in VASP, with VASP5.4 pseudopotentials. No dispersion. | Renewable energy, catalysis, fuel cells, energy conversion, sustainable fertilizer production, chemical refining, plastics synthesis/upcycling | UMA has not seen varying charge or spin multiplicity for the OC20 task, and expects total_charge=0 and spin multiplicity=0 as model inputs. No oxides or explicit solvents are included in OC20. The model works surprisingly well for transition state searches given the nature of the training data, but you should be careful. RPBE works well for small molecules, but dispersion will be important for larger molecules on surfaces. | | **odac** | [ODAC23](https://arxiv.org/abs/2311.00341) | PBE+D3 as implemented in VASP, with VASP5.4 pseudopotentials. | Direct air capture, carbon capture and storage, CO2 conversion, catalysis | UMA has not seen varying charge or spin multiplicity for the ODAC task, and expects total_charge=0 and spin multiplicity=0 as model inputs. The ODAC23 dataset only contains CO2/H2O water absorption, so anything more than might be inaccurate (e.g. hydrocarbons in MOFs). Further, there is a limited number of bare-MOF structures in the training data, so you should be careful if you are using a new MOF structure. | | **oc25** | [OC25](https://arxiv.org/abs/2509.17862) | RPBE+D3 as implemented in VASP, with VASP6.4 pseudopotentials, and dipole corrections. | Renewable energy, (electro)catalysis, fuel cells, energy conversion, sustainable fertilizer production, chemical refining, plastics synthesis/upcycling | UMA has not seen varying charge or spin multiplicity for the OC25 task, and expects total_charge=0 and spin multiplicity=0 as model inputs. The model works surprisingly well for charged systems despite not being explicitly provided that information, but one should be careful. Work functions are not provided by UMA, subsequent DFT calculations are required to extract such information, if desired. Only available in UMA-1.2 | | **oc22** | [OC22](https://arxiv.org/abs/2206.08917) | PBE+U as implemented in VASP with VASP5.4 pseudopotentials and spin-polarization | Oxide catalysis or supports, energy conversion and storage | UMA has not seen varying charge or spin multiplicity for the OC22 task, and expects total_charge=0 and spin multiplicity=0 as model inputs. No explicit solvents are included in OC22. Only available in UMA-1.2 | :::{note} *OC20 was updated from the original OC20 and recomputed to produce total energies instead of adsorption energies. ::: --- Source: `docs/core/uma_changelog.md` # UMA Release Changelog This page documents the release history of UMA models, including new features, improvements, and bug fixes. --- ## Library compatibility — UMA 1.0 deprecation UMA 1.0 checkpoints are no longer supported (their MoE `include_self` behavior diverges from later releases); loading one raises a `RuntimeError`. To use a 1.0 checkpoint, `pip install 'fairchem-core<=2.21.0'`. Loading now requires a `model_id` to be defined. UMA 1.1 (which ships without one) is auto-tagged `model_id = "UMA-1.1"` in memory; UMA 1.2+ already carry `model_id`; a checkpoint with neither `model_id` nor `backbone.model_version` is rejected. When training, set `model_id` on the model config (e.g. `model_id: UMA-1.2.1`). --- ## UMA 1.2 :::{admonition} Latest Release :class: tip **~50% faster, ~40% more accurate on Open Molecules test set, and expanded data coverage!** ::: ### Model Highlights ::::{grid} 2 :::{grid-item-card} 🔋 Better Charge Handling Leading to improvements across molecular systems: - **~50%** relative improvement on biomolecules vs UMA-S-1.1 - **~30%** relative improvement on electrolytes vs UMA-S-1.1 ::: :::{grid-item-card} 📊 Expanded Data Coverage Now **~520M total DFT calculations** including: - OC22 (oxide catalysts) - OC25 (electrolyte/inorganic interfaces) - Expanded OMol25 data including polymers (OPoly26) ::: :::{grid-item-card} ⚡ Faster Inference **~50% speedup** for UMA-S with turbo mode compared to previous releases ::: :::{grid-item-card} 🔧 Bug Fixes & Stability - Numerical and stability improvements for Hessians and phonons - More accurate diatomics and ionization potentials ::: :::: ### Available Tasks (7 total) | Task | Dataset | Domain | Status | |------|---------|--------|--------| | **oc20** | [OC20](https://arxiv.org/abs/2010.09990) | Heterogeneous Catalysis | No update | | **omat** | [OMat24](https://arxiv.org/abs/2410.12771) | Inorganic Materials | 📊 **Updated** — dimer data | | **omol** | [OMol-1.0](https://arxiv.org/abs/2505.08762) | Organic Molecules | 📊 **Updated** — polymers (OPoly26), ionization potentials, small molecule clusters, dimers | | **odac** | [ODAC23](https://arxiv.org/abs/2311.00341) | MOFs for Direct Air Capture | No update | | **omc** | [OMC25](https://arxiv.org/abs/2508.02651) | Molecular Crystals | No update | | **oc25** | [OC25](https://arxiv.org/abs/2509.17862) | Electrolyte/Inorganic Interfaces | 🆕 **Added to UMA** | | **oc22** | [OC22](https://arxiv.org/abs/2206.08917) | Oxide Catalysts | 🆕 **Added to UMA** | ### Large-Scale Inference & Molecular Dynamics :::{note} UMA is built to easily scale up to multi-node, multi-GPU parallel inference with [Ray](https://www.ray.io/) under the hood. ::: ::::{grid} 2 :::{grid-item-card} 🚀 Large-Scale MD Simulations Run MD with ASE and LAMMPS on large-scale systems with ns/day speeds: - Battle tested up to **hundreds of GPUs** - Systems up to **1M atoms** - UMA handles all parallelism—no need to manually configure LAMMPS, Kokkos, MPI, CUDA, etc. ::: :::{grid-item-card} 📦 Batched Simulations Client–server Ray framework enables batched simulations: - Structural relaxations - Molecular dynamics - **Up to 4x speedup** over serial calculations for batches of small systems on a single H100 GPU ::: :::: --- ## UMA 1.1 :::{admonition} Bug Fix Release :class: warning Fixed bug with size extensivity. ::: ### Available Tasks (5 total) | Task | Dataset | Domain | Status | |------|---------|--------|--------| | **oc20** | [OC20](https://arxiv.org/abs/2010.09990) | Heterogeneous Catalysis | No update | | **omat** | [OMat24](https://arxiv.org/abs/2410.12771) | Inorganic Materials | No update | | **omol** | [OMol25](https://arxiv.org/abs/2505.08762) | Organic Molecules | No update | | **odac** | [ODAC23](https://arxiv.org/abs/2311.00341) | MOFs for Direct Air Capture | No update | | **omc** | [OMC25](https://arxiv.org/abs/2508.02651) | Molecular Crystals | No update | --- ## UMA 1.0 :::{admonition} Initial Release — May 2025 :class: note The first public release of UMA, introducing a unified model for atoms across multiple domains. ::: ### Available Tasks (5 total) | Task | Dataset | Domain | Status | |------|---------|--------|--------| | **oc20** | [OC20](https://arxiv.org/abs/2010.09990) | Heterogeneous Catalysis | 🆕 **New** | | **omat** | [OMat24](https://arxiv.org/abs/2410.12771) | Inorganic Materials | 🆕 **New** | | **omol** | [OMol25](https://arxiv.org/abs/2505.08762) | Organic Molecules | 🆕 **New** | | **odac** | [ODAC23](https://arxiv.org/abs/2311.00341) | MOFs for Direct Air Capture | 🆕 **New** | | **omc** | [OMC25](https://arxiv.org/abs/2508.02651) | Molecular Crystals | 🆕 **New** | ### Key Features - **Mixture-of-Linear-Experts (MoLE)** architecture for high parameter count with fast inference - **Equivariant GNN** preserving energy conservation - **Multi-domain training** on 500M+ DFT calculations - **ASE Calculator integration** for seamless workflow adoption --- Source: `docs/core/intro.md` # GNNs for Chemistry The most recent, state of the art machine learned potentials in atomistic simulations are based on graph models that are trained on large (1M+) datasets. These models can be downloaded and used in a wide array of applications ranging from catalysis to materials properties. These pre-trained models can be used on their own, to accelerate DFT calculation, and they can also be used as a starting point to fine-tune new models for specific tasks. ## Background on DFT and machine learning potentials Density functional theory (DFT) has been a mainstay in molecular simulation, but its high computational cost limits the number and size of simulations that are practical. Over the past two decades machine learning has increasingly been used to build surrogate models to supplement DFT. We call these models machine learned potentials (MLP). In the early days, neural networks were trained using the cartesian coordinates of atomistic systems as features with some success. These features lack important physical properties, notably they lack invariance to rotations, translations and permutations, and they are extensive features, which limit them to the specific system being investigated. About 15 years ago, a new set of features called symmetry functions were developed that were intensive, and which had these invariances. These functions enabled substantial progress in MLP, but they had a few important limitations. First, the size of the feature vector scaled quadratically with the number of elements, practically limiting the MLP to 4-5 elements. Second, composition was usually implicit in the functions, which limited the transferability of the MLP to new systems. Finally, these functions were "hand-crafted", with limited or no adaptability to the systems being explored, thus one needed to use judgement and experience to select them. While progress has been made in mitigating these limitations, a new approach has overtaken these methods. Today, the state of the art in machine learned potentials uses graph convolutions to generate the feature vectors. In this approach, atomistic systems are represented as graphs where each node is an atom, and the edges connect the nodes (atoms) and roughly represent interactions or bonds between atoms. Then, there are machine learnable convolution functions that operate on the graph to generate feature vectors. These operators can work on pairs, triplets and quadruplets of nodes to compute "messages" that are passed to the central node (atom) and accumulated into the feature vector. This feature generate method can be constructed with all the desired invariances, the functions are machine learnable, and adapt to the systems being studied, and it scales well to high numbers of elements (the current models handle 50+ elements). These kind of MLPs began appearing regularly in the literature around 2016. :::{note} Today an MLP consists of three things: 1. A **model** that takes an atomistic system, generates features and relates those features to some output. 2. A **dataset** that provides the atomistic systems and the desired output labels. This label could be energy, forces, or other atomistic properties. 3. A **checkpoint** that stores the trained model for use in predictions. ::: ## FAIR Chemistry models FAIRChem provides a number of GNNs in this repository. Each model represents a different approach to featurization, and a different machine learning architecture. The models can be used for different tasks, and you will find different checkpoints associated with different datasets and tasks. Read the papers for details, but we try to highlight here the core ideas and advancements from one model to the next. :::{warning} Since Fairchem version 2.0.0, we are currently only supporting the UMA model code. For all other models please checkout fairchem version 1 of the repo while we bring them back to the new repo. ::: ::::{grid} 1 :::{card} **Universal Model for Atoms (UMA)** - Current SOTA :link: uma **Core Idea:** UMA is an equivariant GNN that leverages a novel technique called Mixture of Linear Experts (MoLE) to give it the capacity to learn the largest multi-modal dataset to date (500M examples and 50B atoms), while preserving energy conservation and inference speed. Even a 6M active parameter (145M total) UMA model is able to achieve SOTA accuracy on a wide range of domains such as materials, molecules and catalysis. **Paper:** ::: :::{card} **equivariant Smooth Energy Network (eSEN)** **Core Idea:** Scaling GNNs to train on hundreds of millions of structures required a number of engineering decisions that led to SOTA models for some tasks, but led to challenges in other tasks. eSEN started with the eSCN network, carefully analyzed which decisions were necessary to build smooth and energy conserving models, and used those learnings to train a new model that is SOTA (as of early 2025) across many domains. **Paper:** ::: :::{card} **Equivariant Transformer V2 (EquiformerV2)** **Core Idea:** We adapted and scaled the Equiformer model to larger datasets using a number of small tweaks/tricks to accelerate training and inference, and incorporating the eSCN convolution operation. This model was also the first shown to be SOTA on OC20 without requiring the underlying structures to be tagged as surface/subsurface atoms, a major improvement in usability. **Paper:** ::: :::{card} **Equivariant Spherical Channel Network (eSCN)** **Core Idea:** The SCN network was high performance, but the approach broke equivariance in the resulting models. eSCN enabled equivariance in these models, and introduced an SO(2) convolution operation that allowed the approach to scale to even higher order spherical harmonics. The model was shown to be equivariant in the limit of infinitely fine grid for the convolution operation. **Paper:** ::: :::{card} **Spherical Channel Network (SCN)** **Core Idea:** We developed a message convolution operation, inspired by the vision AI/ML community, that led to more scalable networks and allowed for higher-order spherical harmonics. This model was SOTA on OC20 on release, but introduced some limitations in equivariance addressed later by eSCN. **Paper:** ::: :::{card} **GemNet-OC** **Core Idea:** GemNet-OC is a faster and more scalable version of GemNet, a model that incorporated some clever features like triplet/quadruplet information into GNNs, and provided SOTA performance when released on OC20. **Paper:** ::: :::: --- Source: `docs/core/common_tasks/summary.md` # Common Tasks This section provides practical guides for common tasks you will encounter when working with FAIRChem models. Whether you are running inference, training models, or integrating with simulation tools, these guides will help you get started quickly. ::::{grid} 1 2 2 3 :::{card} ASE Calculator :link: ase_calculator Use FAIRChem models with ASE for single-point calculations, relaxations, and molecular dynamics. ::: :::{card} Batch Inference :link: batch_inference Efficiently run predictions on many structures using batched inference. ::: :::{card} InferenceBatcher :link: inference_batcher Run many independent ASE simulations with concurrent batched inference for improved GPU utilization. ::: :::{card} Dataset Creation :link: ase_dataset_creation Create custom datasets from ASE databases, CIF files, trajectories, and other formats. ::: :::{card} Fine-tuning :link: fine_tuning Fine-tune pretrained UMA models on your own datasets for improved accuracy. ::: :::{card} Training :link: training Train models from scratch using the FAIRChem training framework. ::: :::{card} Evaluation :link: evaluation Evaluate pretrained models on standard benchmarks and metrics. ::: :::{card} Benchmarks :link: benchmark Run downstream property benchmarks like relaxations, elastic tensors, and phonons. ::: :::{card} LAMMPS Integration :link: lammps Use FAIRChem models with LAMMPS for large-scale molecular dynamics simulations. ::: :::{card} Workflows :link: workflows Integrate FAIRChem with QuAcc for complex molecular simulation workflows. ::: :::: --- Source: `docs/core/common_tasks/ase_calculator.md` # Inference using ASE and Predictor Interface Inference is done using [MLIPPredictUnit](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/units/mlip_unit/mlip_unit.py#L867). The [FairchemCalculator](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/calculate/ase_calculator.py#L3) (an ASE calculator) is simply a convenience wrapper around the MLIPPredictUnit. :::{tip} For simple cases such as demos or education, the ASE calculator is very easy to use. For more complex cases such as running MD or batched inference, we recommend using the predictor directly for better performance. ::: ```{code-cell} python3 from __future__ import annotations from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1", device="cuda") calc = FAIRChemCalculator(predictor, task_name="oc20") ``` ## Adding a DFT-D3(BJ) dispersion correction ```{code-cell} python3 :tags: [skip-execution] from fairchem.core import DFTD3Calculator, FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", device="cuda", inference_settings="turbo" ) base_calc = FAIRChemCalculator(predictor, task_name="omat") # Use "pbe" for PBE-trained models such as UMA's OMat head, or "r2scan" # for models fine-tuned on r2SCAN data. calc = DFTD3Calculator(base_calc, functional="pbe", device="cuda") atoms.calc = calc ``` The wrapper adds the D3 energy, forces, and analytic stress to the base calculator. The `pbe` and `r2scan` presets both use Becke-Johnson damping, a 15 Å cutoff, and C5 smoothing over the outer 20% of the cutoff. Pass `param_file` and `auto_download=False` to use a local D3 parameter table without network access. | `functional` | `a1` | `a2` (Bohr) | `s6` | `s8` | | --- | ---: | ---: | ---: | ---: | | `pbe` | 0.4289 | 4.4407 | 1.0 | 0.7875 | | `r2scan` | 0.49484001 | 5.73083694 | 1.0 | 0.78981345 | Both presets use `k1=16.0` and `k3=-4.0`. The PBE-D3(BJ) parameters are from: > S. Grimme, S. Ehrlich, and L. Goerigk, “Effect of the damping function in > dispersion corrected density functional theory,” *J. Comput. Chem.* **32**, > 1456–1465 (2011). [doi:10.1002/jcc.21759](https://doi.org/10.1002/jcc.21759) The r2SCAN-D3(BJ) parameters were reported and benchmarked alongside D4 in: > S. Ehlert, U. Huniar, J. Ning, J. W. Furness, J. Sun, A. D. Kaplan, > J. P. Perdew, and J. G. Brandenburg, “r²SCAN-D4: Dispersion corrected > meta-generalized gradient approximation for general chemical applications,” > *J. Chem. Phys.* **154**, 061101 (2021). > [doi:10.1063/5.0041008](https://doi.org/10.1063/5.0041008) ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ## Default mode UMA defaults to the `merge_mole + compile` fast mode with TF32 disabled. This fast path requires fixed composition, task, charge, and spin across repeated evaluations. If a later evaluation changes any of these, the calculator prints a warning and permanently falls back to the unmerged, uncompiled model. Batching is supported; a mixed batch across any of the same parameters triggers the same fallback. ## Batch mode Use batch mode for heterogeneous batches whose systems differ in composition, task, charge, or spin. It currently keeps MOLE unmerged and leaves compilation disabled. The named mode provides a stable entry point for future batch-specific optimizations, such as compilation without MOLE merging. ```{code-cell} python3 predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", device="cuda", inference_settings="batch" ) ``` ## Turbo mode Turbo mode uses the same `merge_mole + compile` fast path as default mode and additionally enables TF32. TF32 can improve performance on compatible hardware at a small precision trade-off. Similar to default mode, any changes in composition, task, charge, and spin across different evaluations trigger a fallback to the unoptimized execution path. ```{code-cell} python3 predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", device="cuda", inference_settings="turbo" ) ``` ## Custom modes for advanced users The advanced user might quickly see that **default**, **batch**, and **turbo** modes are special cases of our [inference settings api](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/units/mlip_unit/api/inference.py#L47). You can customize it for your application if you understand what you are doing. The following table provides more information. | Setting Flag | Description | | ----- | ----- | | tf32 | enables torch [tf32](https://docs.pytorch.org/docs/stable/notes/cuda.html) format for matrix multiplication. This will speed up inference at a slight trade-off for precision. In our tests, it makes minimal difference to most applications. It is able to preserve equivariance, energy conservation for long rollouts. However, if you are computing higher order derivatives such as Hessians or other calculations that requires strict numerical precision, we recommend turning this off | | activation_checkpointing | this uses a custom chunked activation checkpointing algorithm and allows significant savings in memory for a small inference speed penalty. If you are predicting on systems >1000 atoms, we recommend keeping this on. However, if you want the absolute fastest inference possible for small systems, you can turn this off | | merge_mole | This is useful in long rollout applications where the system composition stays constant. By pre-merge the MoLE weights, we can save both memory and compute. | | compile | This uses torch.compile to significantly speed up computation. Due to the way pytorch traces the internal graph, it requires a long compile time during the first iteration and can even recompile anytime it detected a significant change in input dimensions. It is not recommended if you are computing frequently on very different atomic systems. | | external_graph_gen | Only use this if you want to use an external graph generator. This should be rarely used except for development | | internal_graph_gen_version | currently we support v2[default], an internal implementation that is better suited for parallelism and v3 the neighborlist from Nvidia Alchemi library which is faster for single gpu operations. | | edge_chunk_size | Experimental. Used for padding edge sizes. This helps reduce re-compilations from torch compile, default to None | | use_quaternion_wigner | enable quaternion-based Wigner D matrix computation. If false we fall back to euler-angle based rotations. default True. | | base_precision_dtype | governs the main precision type of the computation, default to FP32, FP64 is also supported | | execution_mode | This allows manually toggling custom backends to maximize speed ups. default to "None", when set to "None", the predictor will automatically determine the best backend. For example, "umas-fast-gpu" will introduce 30-40% speedup for uma-s line of models. | For example, for an MD simulation use-case for a system of ~500 atoms, we can choose to use a custom mode like the following: ```{code-cell} python3 from fairchem.core.units.mlip_unit.api.inference import InferenceSettings settings = InferenceSettings( tf32=True, activation_checkpointing=False, merge_mole=True, compile=True, external_graph_gen=False, internal_graph_gen_version=2, ) predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", device="cuda", inference_settings=settings ) ``` ## Enabling gradient stress or Hessian prediction Some tasks, for example omol, odac, or oc20/25, were not trained using stress labels. Similarly, no tasks were supervised to predict Hessians. However, predictions of untrained derivatives of energy, such as stress and Hessians, can be enabled by using the following inference settings flags, | Setting Flag | Description | | ----- | ----- | | predict_untrained_forces | A set of task/dataset names (e.g., `{"omol", "oc20"}`) for which forces will be computed via autograd even though the checkpoint was not trained with a forces head for those tasks. | | predict_untrained_stress | A set of task/dataset names for which stress tensors will be computed via autograd even though the checkpoint was not trained with a stress head for those tasks. The default empty set disables this. | | predict_untrained_hessian | A set of task/dataset names for which the Hessian matrix will be computed via autograd. | For example, to enable stress and Hessian predictions with `omol` level of theory, the following settings can be used, ```{code-cell} python3 settings = InferenceSettings( predict_untrained_stress={'omol'}, predict_untrained_hessian={'omol'} ) predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", device="cuda", inference_settings=settings ) ``` ## Multi-GPU Inference UMA supports Graph Parallel inference natively. The graph is chunked into each rank and both the forward and backwards communication is handled by the built-in graph parallel algorithm with torch distributed. Because Multi-GPU inference requires special setup of communication protocols within a node and across nodes, we leverage [ray](https://www.ray.io/) to launch Ray Actors for each GPU-rank under the hood. This allows us to seamlessly scale to any infrastructure that can run Ray. To make things simple for the user that wants to run multi-gpu inference locally, we provide a drop-in replacement for MLIPPredictUnit, called [ParallelMLIPPredictUnit](https://github.com/facebookresearch/fairchem/blob/85bd83535fedbc1d99eee4c12e175603ccc44ef7/src/fairchem/core/units/mlip_unit/predict.py#L415) :::{note} Multi-GPU inference requires Ray. Install it with `pip install fairchem-core[ray]`. ::: For example, we can create a predictor with 8 GPU workers in a very similar way to MLIPPredictUnit and perform an MD calculation with the ASE calculator. This mode of operation is also compatible with our LAMMPS integration. ```python from ase import units from ase.md.langevin import Langevin from fairchem.core import pretrained_mlip, FAIRChemCalculator import time from fairchem.core.datasets.common_structures import get_fcc_crystal_by_num_atoms predictor = pretrained_mlip.get_predict_unit( "uma-s-1p2p1", inference_settings="turbo", device="cuda", workers=1 ) calc = FAIRChemCalculator(predictor, task_name="omat") atoms = get_fcc_crystal_by_num_atoms(8000) atoms.calc = calc dyn = Langevin( atoms, timestep=0.1 * units.fs, temperature_K=400, friction=0.001 / units.fs, ) # warmup 10 steps dyn.run(steps=10) start_time = time.time() dyn.attach( lambda: print( f"Step: {dyn.get_number_of_steps()}, E: {atoms.get_potential_energy():.3f} eV, " f"QPS: {dyn.get_number_of_steps()/(time.time()-start_time):.2f}" ), interval=1, ) dyn.run(steps=1000) ``` :::{tip} This will automatically create a Ray server on your local machine and use a local client to connect to it. If you have set up a Ray cluster, you can leverage it to run parallel inference on as many nodes as you like. ::: --- Source: `docs/core/common_tasks/lammps.md` # LAMMPS Integration We provide an integration with the [LAMMPS](https://www.lammps.org) Molecular Simulator through the [`fix external`](https://docs.lammps.org/fix_external.html) command. This simple integration hands control of the neighborlist (graph) generation, parallelism, energy, force, and stress calculations all to UMA. :::{danger} Security Warning **Never run YAML configuration files from untrusted sources.** FAIRChem uses [Hydra](https://hydra.cc/) to instantiate Python objects from YAML configs via the `_target_` key. A maliciously crafted config file can execute arbitrary code on your machine. Only use configs that you have written yourself or that come from trusted sources. This is analogous to the security risks of Python's `pickle` and `torch.load()`. ::: :::{tip} The main advantage is that we can optimize UMA for distributed parallel inference directly without modifying LAMMPS. The user would also not need to deal with building LAMMPS from source (see conda install option below) nor [Kokkos](https://docs.lammps.org/Speed_kokkos.html), which is notoriously difficult to build correctly. ::: There is some Python overhead, but for very fast empirical force fields where Python would be a limiting factor, this is negligible at the speeds of current MLIPs (10s - 100s of ms per step). This is the same reason nearly all modern LLM inference uses Python engines. Additionally, to easily scale to multi-node parallelism regimes, we designed the architecture using a client-server interface so LAMMPS would only see the client and the server code running inference can be optimized completely independently later. Since the `fix external` integration simply wraps the UMA predictor interface, the way inference is run is identical to using the [MLIPPredictUnit, ASE Calculator or ParallelMLIPPredictUnit for Multi-GPU inference](https://facebookresearch.github.io/fairchem/core/common_tasks/ase_calculator.html). ## Usage Notes :::{warning} Please note the following differences from regular LAMMPS workflows: ::: - We currently only support `metal` [units](https://docs.lammps.org/units.html), i.e., energy in `eV` and forces in `eV/A` - Users can write LAMMPS scripts in the usual way (see lammps_in_example.file) - Users should **NOT** define other types of forces such as "pair_style", "bond_style" in their scripts. These forces will get added together with UMA forces and most likely produce false results - UMA uses atomic numbers so we try to guess the atomic number from the provided atomic masses in your LAMMPS scripts. Just make sure you provide the right masses for your atom types - this makes it easy so that you don't need to redefine atomic element mappings with LAMMPS :::{note} This assumption fails if you use isotopes or non-standard atomic masses, but we don't expect our models to work in those cases anyway. ::: ## Install and Run Users can install LAMMPS however they like, but the simplest is to install via conda ([https://docs.lammps.org/Install_conda.html](https://docs.lammps.org/Install_conda.html)) if you don't need any bells and whistles. For conda install, activate the conda env with LAMMPS and install fairchem into it. For manual LAMMPS installs, you need to provide python paths so LAMMPS can find fairchem. :::{note} We separate the LAMMPS integration code into a standalone package (`fairchem-lammps`). Please note fairchem-lammps uses the GnuV2 License as is required by any code that uses LAMMPS, instead of the MIT License used by the FAIRChem repository. ::: ```bash # first install conda and lammps following the instructions above, ie: conda install lammps # then activate the environment and install fairchem conda activate lammps-env pip install fairchem-core[extras] pip install fairchem-lammps ``` Assuming you have a classic LAMMPS .in script, make the following changes: 1. Remove all other forces from your LAMMPS script (e.g., pair_style, etc.) 2. Make sure the units are in "metal" 3. Make sure there is only 1 run command at the bottom of the script, if you have multiple run segments, ie: NVT followed by NPT, you can separate them into separate scripts To run, use the Python entrypoint `lmp_fc` (shortcut name for the [python lammps_fc.py script](https://github.com/facebookresearch/fairchem/pull/1454)): ```bash lmp_fc lmp_in="lammps_in_example.file" task_name="omol" ``` ## Multi-GPU Parallelism Our LAMMPS integration is fully compatible out of the box with our Multi-GPU inference API. :::{note} Multi-GPU inference requires Ray. Install it with `pip install fairchem-core[ray]`. ::: :::{tip} The only change required is to pass the `ParallelMLIPPredictUnit` [here](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/lammps/lammps_fc_config.yaml#L20) instead of the regular predict unit when initializing the LAMMPS fairchem script. No need to install anything new such as Kokkos or add communication code. ::: For example: ```bash lmp_fc lmp_in="lammps_in_example.file" task_name="omol" predict_unit='${parallel_predict_unit}' ``` --- Source: `docs/core/common_tasks/batch_inference.md` # Batch Inference with UMA Models ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` If your application requires predictions over many systems, you can run batch inference using UMA models to use compute more efficiently and improve GPU utilization. :::{tip} To learn more about the different inference settings supported, see the [Prediction interface documentation](https://facebookresearch.github.io/fairchem/core/common_tasks/ase_calculator.html). ::: ## Generate Batches at Runtime The recommended way to create batches at runtime is to convert ASE `Atoms` objects into `AtomicData` as follows, ```python from ase.build import bulk, molecule from fairchem.core import pretrained_mlip from fairchem.core.datasets.atomic_data import AtomicData, atomicdata_list_to_batch atoms_list = [bulk("Pt"), bulk("Cu"), bulk("NaCl", crystalstructure="rocksalt", a=2.0)] # you need to assign the task_name desired atomic_data_list = [ AtomicData.from_ase(atoms, task_name="omat") for atoms in atoms_list ] batch = atomicdata_list_to_batch(atomic_data_list) predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1", device="cuda") preds = predictor.predict(batch) ``` The predictions are returned in a dictionary with single `torch.Tensor` value for each property predicted. system level properties can be accessed using the same index for the system in the `atomic_data_list`, atom level properties like forces can be obtained for a single system in the batch using the `batch.batch` attribute, ```python # energy of the first system in the batch preds["energy"][0] # forces of the first system in the batch preds["forces"][batch.batch == 0] ``` ## Batch Inference Using a Dataset and DataLoader If you are running predictions over more structures than you can fit in memory, you can run inference using a torch DataLoader: ```python from torch.utils.data import DataLoader from fairchem.core.datasets import AseDBDataset from fairchem.core.datasets.atomic_data import atomicdata_list_to_batch dataset = AseDBDataset( config=dict(src="path/to/your/dataset.aselmdb", a2g_args=dict(task_name="omol")) ) loader = DataLoader(dataset, batch_size=200, collate_fn=atomicdata_list_to_batch) predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1", device="cuda") for batch in loader: preds = predictor.predict(batch) ``` ## Inference over Heterogeneous Batches :::{note} For cases where you want to batch systems to be computed with different task predictions (e.g., molecules and materials), you can take advantage of UMA models and do it in a single batch. ::: ```python from ase.build import bulk, molecule from fairchem.core import pretrained_mlip from fairchem.core.datasets.atomic_data import AtomicData, atomicdata_list_to_batch # a molecule h2o = molecule("H2O") h2o.info.update({"charge": 0, "spin": 1}) # a bulk pt = bulk("Pt") # a catalytic surface slab = fcc100("Cu", (3, 3, 3), vacuum=8, periodic=True) adsorbate = molecule("CO") add_adsorbate(slab, adsorbate, 2.0, "bridge") atomic_data_list = [ # note that we put the molecule in a large box AtomicData.from_ase( h2o, task_name="omol", r_data_keys=["spin", "charge"], molecule_cell_size=12 ), AtomicData.from_ase(pt, task_name="omat"), AtomicData.from_ase(slab, task_name="oc20"), ] batch = atomicdata_list_to_batch(atomic_data_list) predictions = predictor.predict(batch) ``` --- Source: `docs/core/common_tasks/inference_batcher.md` # Batched Atomic Simulations with InferenceBatcher ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` :::{warning} The `InferenceBatcher` class and underlying concurrent batching implementations are experimental and under current development. The API may change. If you have suggestions for improvements, please open an issue or submit a pull request. ::: :::{note} `InferenceBatcher` requires Ray. Install it with `pip install fairchem-core[ray]`. ::: When running many independent ASE calculations (relaxations, molecular dynamics, etc.) on small to medium-sized systems, you can significantly improve GPU utilization by batching model inference calls together. The `InferenceBatcher` class provides a high-level API to do this with minimal code changes. The key idea is simple: instead of running each simulation sequentially, `InferenceBatcher` collects inference requests from multiple concurrent simulations and batches them together for more efficient GPU computation. ## Basic Setup To use `InferenceBatcher`, you need to: 1. Create a predict unit as usual 2. Wrap it with `InferenceBatcher` 3. Use `batcher.batch_predict_unit` instead of the original predict unit in your simulation functions ```python from fairchem.core import pretrained_mlip from fairchem.core.calculate import FAIRChemCalculator, InferenceBatcher # Create a predict unit predict_unit = pretrained_mlip.get_predict_unit("uma-s-1p2p1") # Wrap it with InferenceBatcher batcher = InferenceBatcher( predict_unit, concurrency_backend_options=dict(max_workers=32) ) ``` :::{tip} The `max_workers` parameter controls how many concurrent simulations can run concurrently. Adjust this based on your system's memory and the size of your structures. ::: ## Writing Simulation Functions The only requirement for using `InferenceBatcher` is to write your simulation logic as a function that takes an `Atoms` object and a predict unit as arguments: ```python from ase.build import bulk from ase.filters import FrechetCellFilter from ase.optimize import LBFGS def run_relaxation(atoms, predict_unit): """Run a structure relaxation and return the final energy.""" calc = FAIRChemCalculator(predict_unit, task_name="omat") atoms.calc = calc opt = LBFGS(FrechetCellFilter(atoms), logfile=None) opt.run(fmax=0.02, steps=100) return atoms.get_potential_energy() ``` ## Running Batched Relaxations Once you have your simulation function, you can run it in batched mode using the executor's `map` or `submit` methods: ### Using `executor.map` ```python from functools import partial # Create a list of structures to relax prim_atoms = [ bulk("Cu"), bulk("MgO", "rocksalt", a=4.2), bulk("Si", "diamond", a=5.43), bulk("NaCl", "rocksalt", a=3.8), ] atoms_list = [make_supercell(atoms, 3 * np.identity(3)) for atoms in prim_atoms] for atoms in atoms_list: atoms.rattle(0.1) # Create a partial function with the batch predict unit run_relaxation_batched = partial( run_relaxation, predict_unit=batcher.batch_predict_unit ) # Run all relaxations in parallel with batched inference relaxed_energies = list(batcher.executor.map(run_relaxation_batched, atoms_list)) ``` ### Using `executor.submit` for more control If you need more control over the execution or want to process results as they complete: ```python # Create a new list of structures to relax atoms_list = [make_supercell(atoms, 3 * np.identity(3)) for atoms in prim_atoms] for atoms in atoms_list: atoms.rattle(0.1) # Submit all jobs futures = [ batcher.executor.submit(run_relaxation, atoms, batcher.batch_predict_unit) for atoms in atoms_list ] # Collect results relaxed_energies = [future.result() for future in futures] ``` ## Running Batched Molecular Dynamics The same pattern works for molecular dynamics simulations: ```python from ase import units from ase.md.langevin import Langevin from ase.md.velocitydistribution import MaxwellBoltzmannDistribution def run_nvt_md(atoms, predict_unit, temperature, traj_fname): """Run NVT molecular dynamics simulation.""" calc = FAIRChemCalculator(predict_unit, task_name="omat") atoms.calc = calc MaxwellBoltzmannDistribution(atoms, temperature, force_temp=True) dyn = Langevin( atoms, timestep=2 * units.fs, temperature_K=temperature, friction=0.1, trajectory=traj_fname, loginterval=5, ) dyn.run(100) # Run batched MD simulations run_md_batched = partial( run_nvt_md, predict_unit=batcher.batch_predict_unit, temperature=300 ) futures = [ batcher.executor.submit(run_md_batched, atoms, traj_fname=f"traj_{i}.traj") for i, atoms in enumerate(atoms_list) ] # Wait for all simulations to complete [future.result() for future in futures] ``` ## When to Use InferenceBatcher `InferenceBatcher` is most beneficial when: - Running many independent simulations on small to medium-sized systems - GPU utilization is low with serial execution - Each individual simulation has many inference steps (relaxations, MD) :::{note} When running batch inference over static structures, consider using the [batch inference approach](batch_inference.md) with `AtomicData` directly instead. For single large systems, consider using the `MLIPParallelPredictUnit` for graph parallel inference. ::: --- Source: `docs/core/common_tasks/ase_dataset_creation.md` # FAIRChem and Custom Datasets ## Datasets in FAIRChem `fairchem` provides training and evaluation code for tasks and models that take arbitrary chemical structures as input to predict energies, forces, positions, and stresses. It can be used as a base scaffold for research projects. For an overview of tasks, data, and metrics, please read the documentation and respective papers: - [OC20](catalysts/datasets/oc20) - [OC22](catalysts/datasets/oc22) - [ODAC23](dac/datasets/odac) - [OC20Dense](catalysts/datasets/oc20dense) - [OC20NEB](catalysts/datasets/oc20neb) - [OMat24](inorganic_materials/datasets/omat24) - [OMol25](https://ai.meta.com/blog/meta-fair-science-new-open-source-releases/) - [OMC25](molecules/datasets/omc25) :::{note} There are multiple ways to train and evaluate FAIRChem models on data other than OC20 and OC22. Writing an LMDB is the most performant option. However, ASE-based dataset formats are also included as a convenience for people with existing data who simply want to try fairchem tools without needing to learn about LMDBs. ::: ## Custom ASE Databases If your data is already in an [ASE Database](https://databases.fysik.dtu.dk/ase/ase/db/db.html), no additional preprocessing is necessary before running training/prediction! :::{tip} Although the ASE DB backends may not be sufficiently high throughput for all use cases, they are generally considered "fast enough" to train on a reasonably-sized dataset with 1-2 GPUs or predict with a single GPU. If you want to effectively utilize more resources than this, consider writing your data to an LMDB. ::: :::{admonition} Performance Tip :class: dropdown If your dataset is small enough to fit in CPU memory, use the `keep_in_memory: True` option to avoid I/O bottlenecks and significantly speed up training. ::: To use this dataset, we will just have to change our config files to use the ASE DB Dataset rather than the LMDB Dataset: ```yaml dataset: format: ase_db train: src: # The path/address to your ASE DB connect_args: # Keyword arguments for ase.db.connect() select_args: # Keyword arguments for ase.db.select() # These can be used to query/filter the ASE DB a2g_args: r_energy: True r_forces: True # Set these if you want to train on energy/forces # Energy/force information must be in the ASE DB! keep_in_memory: False # Keeping the dataset in memory reduces random reads and is extremely fast, but this is only feasible for relatively small datasets! include_relaxed_energy: False # Read the last structure's energy and save as "y_relaxed" for IS2RE-Direct training val: src: a2g_args: r_energy: True r_forces: True test: src: a2g_args: r_energy: False r_forces: False # It is not necessary to have energy or forces if you are just making predictions. ``` ## Using ASE-Readable Files It is possible to train/predict directly on ASE-readable files. :::{warning} This is only recommended for smaller datasets, as directories of many small files do not scale efficiently on all computing infrastructures. ::: There are two options for loading data with the ASE reader: ### Single-Structure Files This dataset assumes a single structure will be obtained from each file: ```yaml dataset: format: ase_read train: src: # The folder that contains ASE-readable files pattern: # Pattern matching each file you want to read (e.g. "*/POSCAR"). Search recursively with two wildcards: "**/*.cif". include_relaxed_energy: False # Read the last structure's energy and save as "y_relaxed" for IS2RE-Direct training ase_read_args: # Keyword arguments for ase.io.read() a2g_args: # Include energy and forces for training purposes # If True, the energy/forces must be readable from the file (ex. OUTCAR) r_energy: True r_forces: True keep_in_memory: False ``` ### Multi-structure Files This dataset supports reading files that each contain multiple structures (for example, an ASE .traj file). :::{tip} Using an index file, which tells the dataset how many structures each file contains, is recommended. Otherwise, the dataset is forced to load every file at startup and count the number of structures! ::: ```yaml dataset: format: ase_read_multi train: index_file: Filepath to an index file which contains each filename and the number of structures in each file. e.g.: /path/to/relaxation1.traj 200 /path/to/relaxation2.traj 150 ... # If using an index file, the src and pattern are not necessary src: # The folder that contains ASE-readable files pattern: # Pattern matching each file you want to read (e.g. "*.traj"). Search recursively with two wildcards: "**/*.xyz". ase_read_args: # Keyword arguments for ase.io.read() a2g_args: # Include energy and forces for training purposes r_energy: True r_forces: True keep_in_memory: False ``` --- Source: `docs/core/common_tasks/training.md` # Training Models from Scratch This repo is used to train large state-of-the-art graph neural networks from scratch on datasets like OC20, OMol25, or OMat24, among others. :::{danger} Security Warning **Never run YAML configuration files from untrusted sources.** FAIRChem uses [Hydra](https://hydra.cc/) to instantiate Python objects from YAML configs via the `_target_` key. A maliciously crafted config file can execute arbitrary code on your machine. Only use configs that you have written yourself or that come from trusted sources. This is analogous to the security risks of Python's `pickle` and `torch.load()`. ::: :::{tip} We now provide a simple CLI to handle this using your own custom datasets, but we suggest fine-tuning one of the existing checkpoints first before trying a from-scratch training. ::: ## FAIRChem Training Framework Overview The FAIRChem training framework currently uses a simple SPMD (Single Program Multiple Data) paradigm. It is made of several components: 1. **User CLI and Launcher** - The `fairchem` CLI can run jobs locally using [torch distributed elastic](https://docs.pytorch.org/docs/stable/distributed.elastic.html) or on [SLURM](https://slurm.schedmd.com/documentation.html). More environments may be supported in the future. 2. **Configuration** - We strictly use [Hydra YAMLs](https://hydra.cc/docs/intro/) for configuration. 3. **Runner Interface** - The core program code that is replicated to run on all ranks. An optional Reducer is also available for evaluation jobs. Runners are distinct user functions that run on a single rank (i.e., GPU). They describe separate high-level tasks such as Train, Eval, Predict, Relaxations, MD, etc. Anyone can write a new runner if its functionality is sufficiently different than the ones that already exist. 4. **Trainer** - We use [TorchTNT](https://docs.pytorch.org/tnt/stable/) as a light-weight training loop. This allows us to cleanly separate the data loading from the training loop. :::{note} TNT is PyTorch's replacement for PyTorch Lightning - which has become severely bloated and difficult to use over the years; so we opted for the simpler option. Units are concepts in TorchTNT that provide a basic interface for training, evaluation, and prediction. These replace trainers in fairchemv1. You should write a new unit when the model paradigm is significantly different, e.g., training a Multitask-MLIP is one unit, training a diffusion model should be another unit. ::: ## FAIRChem v2 CLI FAIRChem uses a single [CLI](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/_cli.py) for running jobs. It accepts a single argument, the location of the Hydra YAML. :::{warning} This is intentional to make sure all configuration is fully captured and avoid bloating of the command line interface. Because of the flexibility of Hydra YAMLs, you can still provide additional parameters and overrides using the [Hydra override syntax](https://hydra.cc/docs/advanced/override_grammar/basic/). ::: The CLI can launch jobs locally using [torch distributed elastic](https://docs.pytorch.org/docs/stable/distributed.elastic.html) OR on [SLURM](https://slurm.schedmd.com/documentation.html). ### FAIRChem v2 Config Structure A FAIRChem config is composed of only 2 valid top-level keys: **job** (Job Config) and **runner** (Runner Config). Additionally, you can add key/values that are used by the OmegaConf interpolation syntax to replace fields. Other than these, no other top-level keys are permitted. - **JobConfig** represents configuration parameters that describe the overall job (mostly infra parameters) such as number of nodes, log locations, loggers, etc. This is a structured config and must strictly adhere to the JobConfig class. - **Runner Config** describes the user code. This part of config is recursively instantiated at the start of a job using Hydra instantiation framework. ### Example Configurations **Local run:** ```yaml job: device_type: CUDA scheduler: mode: LOCAL ranks_per_node: 4 run_name: local_training_run ``` **SLURM run:** ```yaml job: device_type: CUDA scheduler: mode: SLURM ranks_per_node: 8 num_nodes: 4 slurm: account: ${cluster.account} qos: ${cluster.qos} mem_gb: ${cluster.mem_gb} cpus_per_task: ${cluster.cpus_per_task} run_dir: /path/to/output run_name: slurm_run_example ``` ### Config Object Instantiation To keep our configs explicit (configs should be thought of as an extension of code), we prefer to use the Hydra instantiation framework throughout; the config is always fully described by a corresponding Python class and should never be a standalone dictionary. :::{admonition} Good vs Bad Config Patterns :class: dropdown **Bad pattern** - We have no idea where to find the code that uses runner or where variables x and y are actually used: ```yaml runner: x: 5 y: 6 ``` **Good pattern** - Now we know which class runner corresponds to and that x, y are just initializer variables of runner. If we need to check the definition or understand the code, we can simply go to runner.py: ```yaml runner: _target_: fairchem.core.components.runner.Runner x: 5 y: 6 ``` ::: ### Runtime Instantiation with Partial Functions While we want to use static instantiation as much as possible, there will be many cases where certain objects require runtime inputs to create. For example, if we want to create a PyTorch optimizer, we can give it all the arguments except the model parameters (because it's only known at runtime). ```yaml optimizer: _target_: torch.optim.AdamW params: ?? # this is only known at runtime lr: 8e-4 weight_decay: 1e-3 ``` :::{tip} In this case we can use a partial function. Instead of creating an optimizer object, we create a Python partial function that can then be used to instantiate the optimizer in code later: ::: ```yaml optimizer_fn: _target_: torch.optim.AdamW _partial_: true lr: 8e-4 weight_decay: 1e-3 ``` ```python # later in the runner optimizer = optimizer_fn(model.parameters()) ``` ## Training UMA The UMA model is completely defined [here](https://github.com/facebookresearch/fairchem/tree/main/src/fairchem/core/models/uma). It is also called "escn_md" during internal development since it was based on the eSEN architecture. Training, evaluation, and inference are all defined in the [mlip unit](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/units/mlip_unit/mlip_unit.py). To train a model, we need to initialize a [TrainRunner](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/train/train_runner.py) with a [MLIPTrainEvalUnit](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/units/mlip_unit/mlip_unit.py). Due to the complexity of UMA and training a multi-architecture, multi-dataset, multi-task model, we leverage [config groups](https://hydra.cc/docs/tutorials/basic/your_first_app/config_groups/) syntax in Hydra to organize UMA training into the [following sections](https://github.com/facebookresearch/fairchem/tree/main/configs/uma/training_release): - **backbone** - selects the specific backbone architecture (e.g., uma-sm, uma-md, uma-large) - **cluster** - quickly switch settings between different SLURM clusters or local environment - **dataset** - select the dataset to train on - **element_refs** - select the element references - **tasks** - select the task set (e.g., for direct or conservative training) :::{tip} We can switch between different combinations of configs easily this way! ::: ### Example Commands **Get training started locally using local settings and the debug dataset:** ```bash fairchem -c configs/uma/training_release/uma_sm_direct_pretrain.yaml cluster=h100_local dataset=uma_debug ``` **Train UMA conservative with 16 nodes on SLURM:** ```bash fairchem -c configs/uma/training_release/uma_sm_conserve_finetune.yaml cluster=h100 job.scheduler.num_nodes=16 run_name="uma_conserve_train" ``` --- Source: `docs/core/common_tasks/evaluation.md` # Evaluating Pretrained Models FAIRChem v2 provides a number of methods used to benchmark and evaluate the UMA models that will be helpful for apples-to-apples comparisons with the paper results. :::{danger} Security Warning **Never run YAML configuration files from untrusted sources.** FAIRChem uses [Hydra](https://hydra.cc/) to instantiate Python objects from YAML configs via the `_target_` key. A maliciously crafted config file can execute arbitrary code on your machine. Only use configs that you have written yourself or that come from trusted sources. This is analogous to the security risks of Python's `pickle` and `torch.load()`. ::: ## Running Model Evaluations To evaluate a UMA model using a pre-existing configuration file, follow these steps. Example configuration files used to evaluate UMA models are stored in `configs/uma/evaluate`. Run the evaluation script: ```bash fairchem --config evaluation_config.yaml ``` Replace `evaluation_config.yaml` with the desired config file. For example, `configs/uma/evaluate/uma_conserving.yaml` :::{tip} Results will be logged according to the specified logger. We currently only support Weights and Biases. ::: ## Evaluation Configuration File Format Evaluation configuration files are written in Hydra YAML format and specify how a model evaluation should be run. UMA evaluation configuration files, which can be used as templates to evaluate other models if needed, are located in `configs/uma/evaluate/`. ### Top-Level Keys Similar to training configuration files, the only allowed top-level keys are the `job` and `runner` keys as well interpolation keys that are resolved at runtime. - **job**: Contains all settings related to the evaluation job itself, including model, data, and logger configuration. - **runner**: Contains settings for the evaluation runner, such as which script to use and runtime options. Important configuration options are nested under these keys as follows: ### Under `job`: Specifications of how to run the actual job. The configuration options are the same here as those in a training job. Some notable flags are detailed below: - `device_type`: The device to run model inference on (e.g., CUDA or CPU) - `scheduler`: The compute scheduler specifications - `logger`: Configuration for logging results - `type`: Logger type (e.g., `wandb`) - `project`: Logging project name - `entity`: (Optional) Logger entity/user - `run_dir`: Directory where results and logs will be saved ### Under `runner`: The actual benchmark details such as model checkpoint and the dataset are specified under the runner flag. An evaluation run should use the `EvalRunner` class which relies on an `MLIPEvalUnit` to run inference using a pretrained model. - `dataloader`: Dataloader specification for the evaluation dataset - `eval_unit`: The specification of the `MLIPEvalUnit` to be used - `tasks`: The prediction task configuration. In almost all cases, these should be loaded from a model checkpoint using the `fairchem.core.units.mlip_unit.utils.load_tasks` function - `model`: Defines how to load a pretrained model. We recommend using the `fairchem.core.units.mlip_unit.mlip_unit.load_inference_model` function ### Using the `defaults` Key to Define Config Groups The `defaults` key is a Hydra feature that allows you to compose configuration files from modular config groups. Each entry under `defaults` refers to a config group (such as `model`, `data`, or other reusable components) that is merged into the final configuration at runtime. This makes it easy to swap out models, datasets, or other settings without duplicating configuration code. :::{admonition} Example Config Groups :class: dropdown For example, in the UMA evaluation configs we have set up the following config groups and defaults: ```yaml defaults: - _self_ - model: omc_conserving - data: my_eval_data ``` This will include the configuration from `configs/uma/evaluate/model/omc_conserving.yaml` and `configs/uma/evaluate/data/my_eval_data.yaml` into the main config. The `_self_` entry ensures the current file's contents are included. You can create new config groups or override existing ones by changing the entries under `defaults`: ```yaml defaults: - cluster: Configuration settings for a particular compute cluster - dataset: Configuration settings for the evaluation dataset - checkpoint: Configuration settings of the pretrained model checkpoint - _self_ ``` ::: Using config groups allows you to easily override defaults in the CLI. For example: ```bash fairchem --config evaluation_config.yaml cluster=cluster_config checkpoint=checkpoint_config ``` Where `cluster_config` and `checkpoint_config` are cluster and checkpoint configuration files written to directories under cluster and checkpoint respectively. See the files in `configs/uma/evaluate` as a full example. --- Source: `docs/core/common_tasks/fine_tuning.md` # Fine-tuning This repo provides a number of scripts to quickly fine-tune a model using a custom ASE LMDB dataset. These scripts are merely for convenience and fine-tuning uses the exact same tooling and infrastructure as our standard training (see Training section). Training in the fairchem repo uses the fairchem CLI tool and configs are in [Hydra yaml](https://hydra.cc/) format. :::{danger} Security Warning **Never run YAML configuration files from untrusted sources.** FAIRChem uses [Hydra](https://hydra.cc/) to instantiate Python objects from YAML configs via the `_target_` key. A maliciously crafted config file can execute arbitrary code on your machine. Only use configs that you have written yourself or that come from trusted sources. This is analogous to the security risks of Python's `pickle` and `torch.load()`. ::: :::{note} Training datasets must be in the [ASE-lmdb format](https://wiki.fysik.dtu.dk/ase/ase/db/db.html#ase.db.core.connect). For UMA models, we provide a simple script to help generate ASE-lmdb datasets from a variety of input formats (CIFs, traj, extxyz, etc.) as well as a fine-tuning YAML config that can be directly used for fine-tuning. ::: ## Generating Training/Fine-tuning Datasets First we need to generate a dataset in the aselmdb format for fine-tuning. :::{tip} The only requirement is that you have input files that can be read as ASE atoms objects by the `ase.io.read` routine and that they contain energy (forces, stress) in the correct format. For concrete examples, refer to the test at `tests/core/scripts/test_create_finetune_dataset.py`. ::: First you should checkout the fairchem repo and install it to access the scripts ```{code-cell} ipython3 :tags: [skip-execution] git clone git@github.com:facebookresearch/fairchem.git pip install -e fairchem/src/packages/fairchem-core[dev] ``` Run this script to create the aselmdbs as well as a set of templated yamls for finetuning, we will use a few dummy structures for demonstration purposes ```{code-cell} ipython3 import os from pathlib import Path repo_root = next( path for path in (Path.cwd(), *Path.cwd().parents) if (path / "src/fairchem").is_dir() ) os.chdir(repo_root) ! python src/fairchem/core/scripts/create_uma_finetune_dataset.py --train-dir docs/core/common_tasks/finetune_assets/train/ --val-dir docs/core/common_tasks/finetune_assets/val --output-dir /tmp/bulk --uma-task=omat --regression-task e ``` :::{warning} **Task Selection:** The `uma-task` can be one of: `omol`, `odac`, `oc20`, `oc22`, `oc25`, `omat`, `omc`. While UMA was trained in a multi-task fashion, we ONLY support fine-tuning on a single UMA task at a time. Multi-task training can become very complicated! Feel free to contact us on GitHub if you have a special use-case for multi-task fine-tuning, or refer to the training configs in `/training_release` to mimic the original UMA training configs. ::: :::{admonition} Regression Task Options :class: dropdown The `regression-task` can be one of: - **e**: Energy only - **ef**: Energy + forces - **efs**: Energy + forces + stress Choose based on the data you have available in the ASE db. For example, some aperiodic DFT codes only support energy/forces and not gradients, and some very fancy codes like QMC only produce energies. **Note:** Even if you train on just energy or energy/forces, all gradients (forces/stresses) will be computable via the model gradients. ::: This will generate a folder of LMDBs and a `uma_sm_finetune_template.yaml` that you can run directly with the fairchem CLI to start training. :::{tip} If you want to only create the ASE LMDBs, you can use `src/fairchem/core/scripts/create_finetune_dataset.py` which is called by `create_uma_finetune_dataset.py`. ::: ## Model Fine-tuning (Default Settings) The previous step should have generated some YAML files to get you started on fine-tuning. You can simply run this with the `fairchem` CLI. The default is configured to run locally on 1 GPU. ```{code-cell} ipython3 :tags: [skip-execution] ! fairchem -c /tmp/bulk/uma_sm_finetune_template.yaml ``` ## Advanced Configuration The scripts provide a simple way to get started on fine-tuning, but likely for your own use cases you will need to modify the parameters. The configuration uses [Hydra-style YAMLs](https://hydra.cc/). :::{tip} To modify the generated YAMLs, you can either edit the files directly or use [Hydra override notation](https://hydra.cc/docs/advanced/override_grammar/basic/). Changing parameters on the command line is very simple: ::: ```{code-cell} ipython3 :tags: [skip-execution] ! fairchem -c /tmp/bulk/uma_sm_finetune_template.yaml epochs=2 lr=2e-4 job.run_dir=/tmp/finetune_dir +job.timestamp_id=some_id ``` The basic YAML configuration looks like the following: ```yaml job: device_type: CUDA scheduler: mode: LOCAL ranks_per_node: 1 num_nodes: 1 debug: True run_dir: /tmp/uma_finetune_runs/ run_name: uma_finetune logger: _target_: fairchem.core.common.logger.WandBSingletonLogger.init_wandb _partial_: true entity: example project: uma_finetune base_model_name: uma-s-1p2p1 max_neighbors: 300 epochs: 1 steps: null batch_size: 2 lr: 4e-4 train_dataloader ... eval_dataloader ... runner ... ``` :::{admonition} Configuration Parameters :class: dropdown - **base_model_name**: Refers to a model name that can be retrieved from [HuggingFace](https://huggingface.co/facebook/UMA). If you want to use your custom UMA checkpoint, provide the path directly in the runner: ```yaml model: _target_: fairchem.core.units.mlip_unit.mlip_unit.initialize_finetuning_model checkpoint_location: /path/to/your/checkpoint.pt ``` - **max_neighbors**: The number of neighbors used for the equivariant SO2 convolutions. 300 is the default used in UMA training, but if you don't have a lot of memory, 100 is usually fine to ensure smoothness of the potential (see the [ESEN paper](https://arxiv.org/abs/2502.12147)). - **epochs**, **steps**: Choose to either run for an integer number of epochs or steps. Only 1 can be specified; the other must be null. - **batch_size**: In this configuration we use the batch sampler. Start with the largest batch size that can fit on your system without running out of memory. However, don't use a batch size so large that you complete training in very few steps. The optimal batch size is usually the one that minimizes the final validation loss for a fixed compute budget. - **lr**, **weight_decay**: These are standard learning parameters. The recommended values we use are the defaults. ::: ### Logging and Artifacts For logging and checkpoints, all artifacts are stored in the location specified in `job.run_dir`. The visual logger we support is [Weights and Biases](https://wandb.ai/site/). :::{warning} Tensorboard is no longer supported. You must set up your W&B account separately and `job.debug` must be set to `False` for W&B logging to work. ::: ### Distributed Training We support multi-GPU distributed training without additional infrastructure and multi-node distributed training on [SLURM](https://slurm.schedmd.com/documentation.html) only. **Multi-GPU locally:** Simply set `job.scheduler.ranks_per_node=N` where N is the number of GPUs you want to train on. **Multi-node on SLURM:** Change `job.scheduler.mode=SLURM` and set both `job.scheduler.ranks_per_node` and `job.scheduler.num_nodes` to the desired values. :::{note} The `run_dir` must be in a shared network accessible mount for multi-node training to work. ::: ### Resuming Runs To resume from a checkpoint in the middle of a run, find the checkpoint folder at the step you want and use the same fairchem command: ```{code-cell} ipython3 :tags: [skip-execution] ! fairchem -c /tmp/finetune_dir/some_id/checkpoints/final/resume.yaml ``` ### Running Inference on the Fine-tuned Model Inference is run in the same way as the UMA models, except you need to load the checkpoint from a local path. :::{warning} You must use the same task that you used for fine-tuning! ::: ```{code-cell} ipython3 :tags: [skip-execution] from fairchem.core.units.mlip_unit import load_predict_unit from fairchem.core import FAIRChemCalculator predictor = load_predict_unit("/tmp/finetune_dir/some_id/checkpoints/final/inference_ckpt.pt") calc = FAIRChemCalculator(predictor, task_name="omat") ``` --- Source: `docs/core/common_tasks/workflows.md` # Calculation Workflows with FAIRChem Models This repo is integrated with workflow tools like [QuAcc](https://github.com/Quantum-Accelerators/quacc) to make complex molecular simulation workflows easy. You can use any MLP recipe (relaxations, single-points, elastic calculations, etc.) and simply specify the `fairchem` model type. :::{tip} One of the nice things about QuAcc is that you can use plugins for whatever your favorite workflow engine is (Fireworks, Parsl, Prefect, etc.). Some of these methods can scale to hundreds of thousands of parallel calculations and are used by the FAIR chemistry team regularly! ::: Below is an example that uses the default `elastic_tensor_flow` flow: ```{code-cell} ipython3 from __future__ import annotations from ase.build import bulk from quacc.recipes.mlip.elastic import elastic_tensor_flow # Make an Atoms object of a bulk Cu structure atoms = bulk("Cu") # Run an elastic property calculation with our favorite MLP potential result = elastic_tensor_flow( atoms, job_params={ "all": dict( library="fairchem", name_or_path="uma-s-1p2p1", task_name="omat", ), }, ) ``` --- Source: `docs/core/generative_models.md` # Generative Models The FAIR chemistry team has released and published four generative models for inorganic materials and molecules. Close collaborators have also released generative models for catalysts. These releases currently live outside of the main FAIR chemistry repo (noted where appropriate). Much of this work has been driven by an incredible group of Meta/FAIR summer PhD interns! :::{tip} Generative models can create novel materials and molecules by learning patterns from training data, enabling rapid exploration of chemical space. ::: ::::{grid} 1 :::{card} **All-atom Diffusion Transformers (ADiT)** **Core Idea:** Generative models for molecules and generative models for materials tended to be two separate tasks in the AI/ML community, and we developed a transformer-based latent diffusion approach that was able to encode both in the same latent space, leading to synergistic learning. **Paper:** **Code:** ::: :::{card} **FlowLLM** **Core Idea:** We noticed that the Crystal-text-llm did much better at generating compositions than the actual crystal structures, so we took a best-of-both-worlds approach using an LLM to generate compositions we should study, and Flow Matching to generate the actual crystal structures. **Paper:** **Code:** ::: :::{card} **FlowMM** **Core Idea:** Flow Matching, an emerging method in the broader AI/ML generative model space, could be used to more quickly and efficiently generate inorganic crystal structures than some prior diffusion-based methods. **Paper:** **Code:** ::: :::{card} **Crystal-text-llm** **Core Idea:** Building on others in the community who had suggested LLMs could generate molecules or materials as text, we showed that fine-tuned LLaMA models could work quite well for this task, and enable text conditioning. **Paper:** **Code:** ::: :::: --- Source: `docs/leaderboard/leaderboard.md` # Leaderboard :::{tip} Community Leaderboard Submit your model predictions for evaluation on the [fairchem_leaderboard](https://huggingface.co/spaces/facebook/fairchem_leaderboard) hosted on HuggingFace. ::: The FAIR Chemistry Leaderboard provides a centralized platform for evaluating machine learning interatomic potentials (MLIPs) across different chemical domains. Researchers can submit predictions and compare their models against existing benchmarks. The leaderboard currently supports the following datasets: - **[OMol25](omol.md)**: Organic molecules - S2EF evaluation and downstream chemistry tasks (ligand pocket, conformers, spin gap, etc.) - **[OC20](oc20.md)**: Heterogeneous catalysis - S2EF and IS2RE evaluation across in-domain and out-of-domain test splits. ## How to Submit 1. Generate prediction files for the appropriate task (see dataset-specific pages for details). 2. Go to the [fairchem_leaderboard](https://huggingface.co/spaces/facebook/fairchem_leaderboard) on HuggingFace. 3. Sign in with your HuggingFace account. 4. Fill in the submission metadata (model name, organization, contact info, etc.). 5. Select the evaluation type that matches your prediction file. 6. Upload your file and click Submit. 7. Wait for the evaluation to complete and see the success message. :::{warning} Submission limits: Users are limited to **5 successful submissions per month** for each evaluation type. ::: ### Need Help? - [GitHub Issues](https://github.com/facebookresearch/fairchem) - [Discussion Forum](https://huggingface.co/spaces/facebook/fairchem_leaderboard/discussions) --- Source: `docs/leaderboard/omol.md` # OMol25 Leaderboard This leaderboard evaluates performance on the **Open Molecules 2025 (OMol25)** dataset - a diverse, high-quality collection that uniquely combines elemental, chemical, and structural diversity. See the [OMol25 paper](https://arxiv.org/pdf/2505.08762) for more details. The leaderboard is broken into two different sections - "S2EF" and "Evaluations". Structure to Energy and Forces (S2EF) is the most straightforward evaluation for MLIPs - given a structure, how well can you predict the total energy and per-atom forces. Evaluations correspond to several chemistry relevant tasks (spin gap, ligand-strain, etc.) introduced in OMol25 to evaluate MLIPs beyond simple energy and force metrics (see the [paper](https://arxiv.org/pdf/2505.08762) for more details). The simplest way to get started is to have an ASE-compatible MLIP calculator that can make energy and force predictions. Input data for the different benchmarks can be downloaded below. ## Download | Benchmarks | URL | |----------|----------| | S2EF (Val/Test) | [HuggingFace Dataset Splits](https://huggingface.co/facebook/OMol25/blob/main/DATASET.md#dataset-splits) | | Evaluations | [HuggingFace Evaluation Data](https://huggingface.co/facebook/OMol25/blob/main/DATASET.md#evaluation-data) | ## Install the necessary packages ```bash pip install "fairchem-core>=2.5.0" pip install "fairchem-data-omol>=0.1.1" ``` ## S2EF The leaderboard supports S2EF evaluations for both the OMol25 "Validation" and "Test" sets. Validation labels are already accessible in the released dataset for local benchmarking and debugging, so we highly encourage users to make Test submissions to fairly and accurately compare models. The size of each split is as follows: | Split | Size | |----------|----------| | Val | 2,762,021 | | Test | 2,805,046 | Predictions must be saved as ".npz" files and shall contain the following information: ``` ids energy forces natoms ``` Where, - `ids` corresponds to the unique identifier, `atoms.info["source"]` - `energy` is the predicted energy - `forces` is the predicted forces, concatenated across all systems - `natoms` is the number of atoms corresponding to each prediction As an example: ```python from fairchem.core.datasets import AseDBDataset from fairchem.core import pretrained_mlip, FAIRChemCalculator ### Define your MLIP calculator predictor = pretrained_mlip.get_predict_unit(args.checkpoint, device="cuda") calc = FAIRChemCalculator(predictor, task_name="omol") ### Read in the dataset you wish to submit predictions to dataset = AseDBDataset({"src": "path/to/omol/test_data"}) ids = [] energy = [] forces = [] natoms = [] for idx in range(len(dataset)): atoms = dataset.get_atoms(idx) atoms.calc = calc ids.append(atoms.info["source"]) natoms.append(len(atoms)) energy.append(atoms.get_potential_energy()) forces.append(atoms.get_forces()) ### Do not forget this! Your submission will fail. forces = np.concatenate(forces) np.savez_compressed( "test_predictions.npz", ids=ids, energy=energy, forces=forces, natoms=natoms, ) ``` :::{warning} The above example can be very slow on a single GPU and we encourage users to parallelize this however they like. We provide the example as a means to understand the expected format for the leaderboard. ::: Once a prediction file is generated, proceed to the leaderboard, fill in the submission form, upload your file, select "Validation" or "Test" and hit submit. Stay on the page until you see the success message. ## Evaluations The following evaluations are currently available on the OMol25 leaderboard: * Ligand pocket: Protein-ligand interaction energy as a proxy to the binding energy, central to many biological processes. * Ligand strain: Ligand-strain energy is an important task to understanding protein-ligand binding. * Conformers: Identifying the lowest energy conformer is a crucial part of many biological and pharmaceutical tasks. * Protonation: As a proxy to pKa prediction, we evaluate energy differences of structures differing by one proton. * Distance scaling: Short range and long range intermolecular interactions are essential for observable properties like phase changes, density, etc. * IE/EA: The addition, removal, and transfer of electrons is central to many redox processes. * Spin gap: Differences between spin states can play a critical role of molecular optic devices and photactive catalysts. For a detailed descripion of each task we refer people to the original [manuscript](https://arxiv.org/pdf/2505.08762). To generate prediction files for the different tasks, we have released a set of [recipes](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py) to be used with ASE-compatible calculators. Each evaluation task has its own unique structure, a detailed description of the expected output is provided in the recipe docstrings. The following recipes should be used to evaluate the corresponding task: * [Ligand pocket](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L323) * [Ligand strain](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L372) * [Conformers](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L140) * [Protonation](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L188) * [Distance scaling](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L439) * [IE/EA](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L237) * [Spin gap](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/omol.py#L284) As an example, to run the `ligand_pocket` evaluation: ```python import json import pickle from fairchem.core import pretrained_mlip, FAIRChemCalculator from fairchem.core.components.calculate.recipes.omol import ligand_pocket ### Define your MLIP calculator predictor = pretrained_mlip.get_predict_unit(args.checkpoint, device="cuda") calc = FAIRChemCalculator(predictor, task_name="omol") ### Load the desired evaluation task input data with open("path/to/ligand_pocket_inputs.pkl", "rb") as f: ligand_pocket_data = pickle.load(f) results = ligand_pocket(ligand_pocket_data, calc) with open("ligand_pocket_results.json") as f: json.dump(results, f) ``` :::{warning} Conformers, Protonation, Ligand strain, and Distance scaling can be quite slow on a single GPU and we encourage users to parallelize this however they like. ::: Once a prediction file is generated, proceed to the leaderboard, fill in the submission form, upload your file, select the corresponding evaluation task and hit submit. Stay on the page until you see the success message. --- Source: `docs/leaderboard/oc20.md` # OC20 Leaderboard This leaderboard evaluates performance on the **Open Catalyst 2020 (OC20)** dataset - a large-scale dataset for catalyst discovery containing DFT relaxations across a wide variety of adsorbate-catalyst combinations. See the [OC20 paper](https://arxiv.org/abs/2010.09990) for more details. The leaderboard supports two tasks: - **S2EF (Structure to Energy and Forces)**: Predict energy and per-atom forces given an atomic structure. - **IS2RE (Initial Structure to Relaxed Energy)**: Predict the relaxed energy given the initial structure. Both tasks are evaluated across four test splits: - **ID**: In-domain test set - **OOD-Ads**: Out-of-domain adsorbates - **OOD-Cat**: Out-of-domain catalysts - **OOD-Both**: Out-of-domain adsorbates and catalysts ## Download | Benchmarks | URL | |----------|----------| | S2EF | [Test](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_test_lmdbs.tar.gz) | | IS2RE | [Train+Val+Test](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_train_val_test_lmdbs.tar.gz) | ## Install the necessary packages ```bash pip install "fairchem-core>=2.5.0" ``` ## S2EF Predictions must be saved as ".npz" files containing the following keys for each split (`id`, `ood_ads`, `ood_cat`, `ood_both`): ``` {split}_ids {split}_energy {split}_forces {split}_chunk_idx ``` Where, - `{split}_ids` corresponds to the unique system identifiers - `{split}_energy` is the predicted energy for each system - `{split}_forces` is the predicted forces, concatenated across all systems - `{split}_chunk_idx` is the cumulative atom count used to split the concatenated forces back into per-system arrays As an example: ```python import numpy as np from fairchem.core.datasets import AseDBDataset from fairchem.core import pretrained_mlip, FAIRChemCalculator ### Define your MLIP calculator predictor = pretrained_mlip.get_predict_unit(args.checkpoint, device="cuda") calc = FAIRChemCalculator(predictor, task_name="oc20") results = {} for split in ["id", "ood_ads", "ood_cat", "ood_both"]: dataset = AseDBDataset({"src": f"path/to/oc20/s2ef/test/{split}"}) ids = [] energy = [] forces = [] natoms = [] for idx in range(len(dataset)): atoms = dataset.get_atoms(idx) atoms.calc = calc ids.append(atoms.info["sid"]) natoms.append(len(atoms)) energy.append(atoms.get_potential_energy()) forces.append(atoms.get_forces()) forces = np.concatenate(forces) chunk_idx = np.cumsum(natoms)[:-1] results[f"{split}_ids"] = np.array(ids) results[f"{split}_energy"] = np.array(energy) results[f"{split}_forces"] = forces results[f"{split}_chunk_idx"] = chunk_idx np.savez_compressed("oc20_s2ef_predictions.npz", **results) ``` :::{warning} The above example can be very slow on a single GPU and we encourage users to parallelize this however they like. We provide the example as a means to understand the expected format for the leaderboard. ::: ## IS2RE Predictions must be saved as ".npz" files containing the following keys for each split (`id`, `ood_ads`, `ood_cat`, `ood_both`): ``` {split}_ids {split}_energy ``` Where, - `{split}_ids` corresponds to the unique system identifiers - `{split}_energy` is the predicted relaxed energy for each system Once a prediction file is generated, proceed to the [leaderboard](https://huggingface.co/spaces/facebook/fairchem_leaderboard), fill in the submission form, upload your file, select the corresponding evaluation type ("OC20 S2EF Test" or "OC20 IS2RE Test") and hit submit. Stay on the page until you see the success message. --- Source: `docs/molecules/datasets/summary.md` # Datasets :::{margin} ```{image} ../../assets/icons/molecules.svg :alt: Molecules :width: 100px ``` ::: Explore the molecular datasets available for training and benchmarking machine learning interatomic potentials. ::::{grid} 1 2 2 2 :::{card} OMol25 :link: omol25 Over 100 million DFT calculations of organic and inorganic molecules including transition metal complexes and electrolytes. ::: :::{card} OMol25 Electronic Structures :link: omol25_elec Raw ORCA outputs, GBW files, and density matrices from the OMol25 calculations for advanced electronic structure analysis. ::: :::{card} OMC25 :link: omc25 25+ million structures of organic molecular crystals from relaxation trajectories. ::: :::: --- Source: `docs/molecules/datasets/omol25.md` # OMol25 :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 100+ million DFT calculations | | **Domain** | Organic/inorganic molecules, transition metal complexes, electrolytes | | **Labels** | Total energy (eV), forces (eV/A) | | **Level of Theory** | wB97M-V/def2-TZVPD (ORCA6) | | **License** | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | | **Download** | [HuggingFace](https://huggingface.co/facebook/OMol25) | ::: The Open Molecules 2025 (OMol25) dataset contains over 100 million single point calculations of non-equilibrium structures and structural relaxations across a wide swath of organic and inorganic molecular space, including things like transition metal complexes and electrolytes. The dataset contains structures labeled with total energy (eV) and forces (eV/A) via ORCA6. A much larger amount of electronic structure data were also stored during generation and we hope to make these available to the community (reach out via github issue). All information about the dataset is available at the [OMol25 HuggingFace site](https://huggingface.co/facebook/OMol25). If you have issues with the gated model request form, please reach out via a github issue on this repository. ## Dataset format The dataset is provided in ASE DB compatible lmdb files (*.aselmdb). The dataset contains labels of the total charge and spin multiplicity, saved in the `atoms.info` dictionary because ASE does not support these as default properties. ## Level of theory OMol25 was calculated at the wB97M-V/def2-TZVPD level, including non-local dispersion, as defined in ORCA6. To reproduce the calculations, please `fairchem.data.omol.orca.calc` to write compatible ORCA inputs. ### Citing OMol25 The OMol25 dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any publications that uses this dataset: ```bib @misc{levine2025openmolecules2025omol25, title={The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models}, author={Daniel S. Levine and Muhammed Shuaibi and Evan Walter Clark Spotte-Smith and Michael G. Taylor and Muhammad R. Hasyim and Kyle Michel and Ilyes Batatia and Gábor Csányi and Misko Dzamba and Peter Eastman and Nathan C. Frey and Xiang Fu and Vahe Gharakhanyan and Aditi S. Krishnapriyan and Joshua A. Rackers and Sanjeev Raja and Ammar Rizvi and Andrew S. Rosen and Zachary Ulissi and Santiago Vargas and C. Lawrence Zitnick and Samuel M. Blau and Brandon M. Wood}, year={2025}, eprint={2505.08762}, archivePrefix={arXiv}, primaryClass={physics.chem-ph}, url={https://arxiv.org/abs/2505.08762}, } ``` --- Source: `docs/molecules/datasets/omol25_elec.md` # OMol25 Electronic Structures :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 4M subset (full dataset available on request) | | **Contents** | ORCA outputs, GBW files, density matrices | | **Domains** | Small molecules, biomolecules, metal complexes, electrolytes | | **Level of Theory** | wB97M-V/def2-TZVPD | | **Access** | [Argonne National Laboratory via Globus](https://www.materialsdatafacility.org/spotlight/omol25) | ::: The Open Molecules 2025 (OMol25) dataset represents the largest dataset of its kind, with more than 100 million density functional theory (DFT) calculations at the wB97M-V/def2-TZVPD level of theory, spanning several chemical domains including small molecules, biomolecules, metal complexes, and electrolytes. At release, the OMol25 dataset provided structure energies, per-atom forces, and Lowdin/Mulliken charges and spins, where available. These properties were sufficient to train state-of-the-art machine learning interatomic potentials (MLIPs) and are already demonstrating incredible performance across a wide range of applications. However, to maximize the community benefit of these calculations, we have partnered with the [Department of Energy’s Argonne National Laboratory](https://www.anl.gov/) to provide access to the raw DFT outputs and additional files for the OMol25 dataset. By releasing the [ORCA](https://www.faccts.de/docs/orca/6.0/manual/) output files, users will be able to parse NBO orbital/bonding information, reduced orbital populations, Fock matrices, and more. By releasing the ORCA GBW files, users will be able to run electronic structure post-processing in order to obtain higher quality partial charges and partial spins and a variety of more advanced electronic features that could be extremely valuable for physics-informed ML models. Finally, the release will provide critical high quality data for nascent ML models that train directly on electron densities. ## Data Description The OMol25 dataset is broken into several training splits - All and 4M. The 4M split corresponds to a randomly sampled 4M subset of the full OMol25 dataset. Given the size of the full dataset, O(petabytes), we are first releasing all electronic structure and ORCA output data for the 4M split. Based on community interest, we will work to provide the full dataset. For each calculation, the following data is available: * **orca.tar.zst**: Bundle of the raw [ORCA](https://www.faccts.de/docs/orca/6.0/manual/) outputs - including (orca.out, orca.inp orca.engrad, orca_property.txt, orca.xyz). To open: ``` >> tar --zstd -xvf orca.tar.zst orca.engrad orca.inp orca.inp.orig orca.out orca.xyz orca_property.txt orca_stderr ``` * **orca.gbw.zstd0**: Geometry-Basis-Wavefunction file - containing molecular orbitals and wavefunction information for the converged SCF. ``` >> zstd -d orca.gbw.zstd0 -o orca.gbw orca.gbw.zstd0 : 9462880 bytes ``` * **density_mat.npz**: The upper-triangle of the density matrix ("orca.scfp") (two in the case of unrestricted systems with the addition of the spin density ("orca.scfr")). This vectorized form of the density can be inflated into a symmetric matrix with the following code: ```python import numpy as np # Load the NPZ file with np.load('density_mat.npz') as loaded_data: dens_vector = loaded_data['orca.scfp'] # Re-inflate the symmetric matrix n = (np.sqrt(8 * len(dens_vector) + 1) - 1) // 2 mat = np.zeros((n,n)) mat[np.triu_indices(n)] = dens_vector mat = mat + mat.T - np.diag(mat.diagonal()) ``` The dataset is organized on the Argonne cluster based on how we organized it internally for generation. The easiest way to find systems that you may be interested in is by using the ASE-DB format of the dataset that can be downloaded at [train_4M.tar.gz](https://huggingface.co/facebook/OMol25/blob/main/DATASET.md#dataset-splits): ```python # pip install fairchem-core if not already installed from fairchem.core.datasets import AseDBDataset dataset = AseDBDataset({"src": "path/to/train_4M/"}) indices = range(len(dataset)) argonne_paths = [] for idx in indices: # ASE Atoms object that can be visualized/examined atoms = dataset.get_atoms(idx) # Check if this is a system you care about. is_relevant = is_atoms_object_relevant(atoms) if is_relevant: # Extract the relative path that matches the Argonne cluster relative_dir = os.path.dirname(atoms.info["source"]) argonne_paths.append(relative_dir) ``` ## How to Access the Data The data are stored and accessible via storage on the Eagle cluster at Argonne National Laboratory. For free access to the data, you will need to do the following: 1. Follow the access instructions here: [OMol25 Electronic Structure Spotlight Dataset](https://www.materialsdatafacility.org/spotlight/omol25). After group approval (this step requires human validation, so may take some time), you will be able to access the data via the same page. 2. To download the data, you can access via HTTPS calls (slower) or by downloading the [Globus Connect Personal client](https://www.globus.org/globus-connect-personal) (preferred method) and creating a local endpoint. ## Contact Us [General Issues](https://github.com/facebookresearch/fairchem) Dataset questions? * [Muhammed Shuaibi](mshuaibi@meta.com) * [Daniel Levine](levineds@meta.com) Cluster/Access questions? * [Ben Blaiszik](blaiszik@uchicago.edu) --- Source: `docs/molecules/datasets/omc25.md` # OMC25 :::{margin} ```{image} ../../assets/icons/molecular-crystals.svg :alt: Molecular Crystals :width: 100px ``` ::: :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 25+ million structures | | **Domain** | Organic molecular crystals | | **Labels** | Total energy (eV), forces (eV/A), stress (eV/A^3) | | **Level of Theory** | PBE+D3 (VASP) | | **License** | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | | **Download** | [HuggingFace](https://huggingface.co/facebook/OMC25) | ::: The Open Molecular Crystals 2025 (OMC25) dataset comprises >25 million structures of organic molecular crystals from relaxation trajectories of random packings of OE62 molecules into various 3D unit cells using Genarris 3.0 package. The dataset contains structures labeled with total energy (eV), forces (eV/A), and stress (eV/A^3) via VASP. The training and validation splits of the OMC25 dataset are available for download from HuggingFace at https://huggingface.co/facebook/OMC25, under the CC BY 4.0 license, after applying for the repository access on HuggingFace. ## Dataset format The dataset is provided in ASE DB compatible lmdb files (*.aselmdb). ## Level of theory OMC25 was calculated at the PBE+D3 level via VASP. To reproduce the calculations, please use `fairchem.data.omc.scripts.create_vasp_inputs.py` to write compatible VASP inputs. ## Citing We encourage users to cite this paper when using the OMC25 dataset or pretrained models for molecular crystals in their research. ```bibtex @misc{gharakhanyan2025openmolecularcrystals2025omc25dataset, title={Open Molecular Crystals 2025 (OMC25) Dataset and Models}, author={Vahe Gharakhanyan and Luis Barroso-Luque and Yi Yang and Muhammed Shuaibi and Kyle Michel and Daniel S. Levine and Misko Dzamba and Xiang Fu and Meng Gao and Xingyu Liu and Haoran Ni and Keian Noori and Brandon M. Wood and Matt Uyttendaele and Arman Boromand and C. Lawrence Zitnick and Noa Marom and Zachary W. Ulissi and Anuroop Sriram}, year={2025}, eprint={2508.02651}, archivePrefix={arXiv}, primaryClass={physics.chem-ph}, url={https://arxiv.org/abs/2508.02651}, } ``` --- Source: `docs/molecules/models.md` # Pretrained models :::{margin} ```{image} ../assets/icons/molecules.svg :alt: Molecules :width: 100px ``` ::: :::{tip} Recommended Model **2025 recommendation:** We suggest using the [UMA model](../core/uma), trained on all of the FAIR chemistry datasets. UMA provides state-of-the-art accuracy, energy conservation, and will continue to receive updates. ::: The UMA model has a number of nice features over the previous checkpoints: 1. It is state-of-the-art in out-of-domain prediction accuracy 2. The UMA small model is an energy conserving and smooth checkpoint, so should work much better for vibrational calculations, molecular dynamics, etc. 3. The UMA model is most likely to be updated in the future. ## Baseline models in the OMol25 paper As part of the OMol25 release, we released two sets of models: 1. [preferred] UMA models trained on a range of FAIR chemistry datasets, available at [HuggingFace](https://huggingface.co/facebook/UMA) 2. eSEN models trained only on OMol25, available at [HuggingFace](https://huggingface.co/facebook/OMol25/tree/main) The UMA models will continue to be updated regularly and we expect those to remain the default and performant option for the forseeable future. The OMol25-only eSEN models are provided mostly as a base-line for models trained only on OMol25. ## Citing If you use the OMol25-trained eSEN models, please cite the following paper. ```bib @misc{levine2025openmolecules2025omol25, title={The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models}, author={Daniel S. Levine and Muhammed Shuaibi and Evan Walter Clark Spotte-Smith and Michael G. Taylor and Muhammad R. Hasyim and Kyle Michel and Ilyes Batatia and Gábor Csányi and Misko Dzamba and Peter Eastman and Nathan C. Frey and Xiang Fu and Vahe Gharakhanyan and Aditi S. Krishnapriyan and Joshua A. Rackers and Sanjeev Raja and Ammar Rizvi and Andrew S. Rosen and Zachary Ulissi and Santiago Vargas and C. Lawrence Zitnick and Samuel M. Blau and Brandon M. Wood}, year={2025}, eprint={2505.08762}, archivePrefix={arXiv}, primaryClass={physics.chem-ph}, url={https://arxiv.org/abs/2505.08762}, } ``` ## Baseline models in the OMC25 paper As part of the OMC25 release, we released eSEN model trained only on OMC25, available at [HuggingFace](https://huggingface.co/facebook/OMC25). [preferred] UMA models trained on a range of FAIR chemistry datasets are available at [HuggingFace](https://huggingface.co/facebook/UMA). ## Citing We encourage users to cite this paper when using the OMC25 dataset or pretrained models for molecular crystals in their research. ```bibtex @misc{gharakhanyan2025openmolecularcrystals2025omc25dataset, title={Open Molecular Crystals 2025 (OMC25) Dataset and Models}, author={Vahe Gharakhanyan and Luis Barroso-Luque and Yi Yang and Muhammed Shuaibi and Kyle Michel and Daniel S. Levine and Misko Dzamba and Xiang Fu and Meng Gao and Xingyu Liu and Haoran Ni and Keian Noori and Brandon M. Wood and Matt Uyttendaele and Arman Boromand and C. Lawrence Zitnick and Noa Marom and Zachary W. Ulissi and Anuroop Sriram}, year={2025}, eprint={2508.02651}, archivePrefix={arXiv}, primaryClass={physics.chem-ph}, url={https://arxiv.org/abs/2508.02651}, } ``` ## License All models require users to agree to the FAIR Chemistry License as part of the HuggingFace model gating process. --- Source: `docs/inorganic_materials/datasets/summary.md` # Datasets :::{margin} ```{image} ../../assets/icons/inorganic.svg :alt: Inorganic Materials :width: 100px ``` ::: Explore the inorganic materials datasets for training and benchmarking machine learning interatomic potentials. ::::{grid} 1 2 2 2 :::{card} OMat24 :link: omat24 Over 1 million structures from rattled and AIMD calculations for inorganic materials, fully compatible with Matbench-Discovery. ::: :::: --- Source: `docs/inorganic_materials/datasets/omat24.md` # OMat24 :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 1.07M structures (train) + 1.02M (val) | | **Domain** | Inorganic bulk materials | | **Labels** | Total energy (eV), forces (eV/A), stress (eV/A^3) | | **Level of Theory** | DFT (PBE/PBE+U) | | **License** | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | | **Benchmark** | Matbench-Discovery compatible | ::: The Open Materials 2024 (OMat24) dataset contains a mix of single point calculations of non-equilibrium structures and structural relaxations. The dataset contains structures labeled with total energy (eV), forces (eV/A) and stress (eV/A^3). The dataset is provided in ASE DB compatible lmdb files. The OMat24 train and val splits are fully compatible with the Matbench-Discovery benchmark test set. 1. The splits do not contain any structure that has a protostructure label present in the initial or relaxed structures of the WBM dataset. 2. The splits do not include any structure that was generated starting from an Alexandria relaxed structure with protostructure lable in the intitial or relaxed structures of the WBM datset. ## Subdatasets OMat24 is made up of X subdatasets based on how the structures were generated. The subdatasets included are: 1. rattled-1000-subsampled & rattled-1000 2. rattled-500-subsampled & rattled-300 3. rattled-300-subsampled & rattled-500 4. aimd-from-PBE-1000-npt 5. aimd-from-PBE-1000-nvt 6. aimd-from-PBE-3000-npt 7. aimd-from-PBE-3000-nvt 8. rattled-relax **Note** There are two subdatasets for the rattled-< T > datasets. Both subdatasets in each pair were generated with the same procedure as described in our manuscript. ## File contents and downloads ### OMat24 train split | Sub-dataset | No. structures | File size | Download | |:------------------------:|:--------------:|:---------:|:-----------------------------------------------------------------------------------------------------------------------------------------------:| | rattled-1000 | 122,937 | 21 GB | [rattled-1000.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-1000.tar.gz) | | rattled-1000-subsampled | 41,786 | 7.1 GB | [rattled-1000-subsampled.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-1000-subsampled.tar.gz) | | rattled-500 | 75,167 | 13 GB | [rattled-500.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-500.tar.gz) | | rattled-500-subsampled | 43,068 | 7.3 GB | [rattled-500-subsampled.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-500-subsampled.tar.gz) | | rattled-300 | 68,593 | 12 GB | [rattled-300.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-300.tar.gz) | | rattled-300-subsampled | 37,393 | 6.4 GB | [rattled-300-subsampled.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-300-subsampled.tar.gz) | | aimd-from-PBE-1000-npt | 223,574 | 26 GB | [aimd-from-PBE-1000-npt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/aimd-from-PBE-1000-npt.tar.gz) | | aimd-from-PBE-1000-nvt | 215,589 | 24 GB | [aimd-from-PBE-1000-nvt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/aimd-from-PBE-1000-nvt.tar.gz) | | aimd-from-PBE-3000-npt | 65,244 | 25 GB | [aimd-from-PBE-3000-npt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/aimd-from-PBE-3000-npt.tar.gz) | | aimd-from-PBE-3000-nvt | 84,063 | 32 GB | [aimd-from-PBE-3000-nvt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/aimd-from-PBE-3000-nvt.tar.gz) | | rattled-relax | 99,968 | 12 GB | [rattled-relax.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/omat/train/rattled-relax.tar.gz) | | Total | 1,077,382 | 185.8 GB | ### OMat24 val split (this is a 1M subset used to train eqV2 models from the 5M val split) **_NOTE:_** The original validation sets contained a duplicated structures. Corrected validation sets were uploaded on 20/12/24. Please see this [issue](https://github.com/facebookresearch/fairchem/issues/942) for more details, an re-download the correct version of the validation sets if needed. | Sub-dataset | Size | File Size | Download | |:-----------------------:|:---------:|:---------:|----------------------------------------------------------------------------------------------------------------------------------------------:| | rattled-1000 | 117,004 | 218 MB | [rattled-1000.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-1000.tar.gz) | | rattled-1000-subsampled | 39,785 | 77 MB | [rattled-1000-subsampled.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-1000-subsampled.tar.gz) | | rattled-500 | 71,522 | 135 MB | [rattled-500.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-500.tar.gz) | | rattled-500-subsampled | 41,021 | 79 MB | [rattled-500-subsampled.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-500-subsampled.tar.gz) | | rattled-300 | 65,235 | 122 MB | [rattled-300.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-300.tar.gz) | | rattled-300-subsampled | 35,579 | 69 MB | [rattled-300-subsampled.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-300-subsampled.tar.gz) | | aimd-from-PBE-1000-npt | 212,737 | 261 MB | [aimd-from-PBE-1000-npt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/aimd-from-PBE-1000-npt.tar.gz) | | aimd-from-PBE-1000-nvt | 205,165 | 251 MB | [aimd-from-PBE-1000-nvt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/aimd-from-PBE-1000-nvt.tar.gz) | | aimd-from-PBE-3000-npt | 62,130 | 282 MB | [aimd-from-PBE-3000-npt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/aimd-from-PBE-3000-npt.tar.gz) | | aimd-from-PBE-3000-nvt | 79,977 | 364 MB | [aimd-from-PBE-3000-nvt.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/aimd-from-PBE-3000-nvt.tar.gz) | | rattled-relax | 95,206 | 118 MB | [rattled-relax.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241220/omat/val/rattled-relax.tar.gz) | | Total | 1,025,361 | 1.98 GB | ### sAlex Dataset We also provide the sAlex dataset used for fine-tuning of our OMat models. sAlex is a subsampled, Matbench-Discovery compliant, version of the original [Alexandria](https://alexandria.icams.rub.de/). sAlex was created by removing structures matched in WBM and only sampling structure along a trajectory with an energy difference greater than 10 meV/atom. For full details, please see the manuscript. | Dataset | Split | No. Structures | File Size | Download | |:-------:|:-----:|:--------------:|:---------:|-------------------------------------------------------------------------------------------------------:| | sAlex | train | 10,447,765 | 7.6 GB | [train.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/sAlex/train.tar.gz) | | sAlex | val | 553,218 | 408 MB | [val.tar.gz](https://dl.fbaipublicfiles.com/opencatalystproject/data/omat/241018/sAlex/val.tar.gz) | ## Getting ASE atoms objects Dataset files are written as `AseLMDBDatabase` objects which are an implementation of an [ASE Database](https://wiki.fysik.dtu.dk/ase/ase/db/db.html), in LMDB format. A single **.aselmdb* file can be read and queried like any other ASE DB. You can also read many DB files at once and access atoms objects using the `AseDBDataset` class. For example to read the **rattled-relax** subdataset, ```python from fairchem.core.datasets import AseDBDataset dataset_path = "/path/to/omat24/train/rattled-relax" config_kwargs = {} # see tutorial on additiona configuration dataset = AseDBDataset(config=dict(src=dataset_path, **config_kwargs)) # atoms objects can be retrieved by index atoms = dataset.get_atoms(0) ``` To read more than one subdataset you can simply pass a list of subdataset paths, ```python from fairchem.core.datasets import AseDBDataset config_kwargs = {} # see tutorial on additiona configuration dataset_paths = [ "/path/to/omat24/train/rattled-relax", "/path/to/omat24/train/rattled-1000-subsampled", "/path/to/omat24/train/rattled-1000", ] dataset = AseDBDataset(config=dict(src=dataset_paths, **config_kwargs)) ``` To read all of the OMat24 training or validations splits simply pass the paths to all subdatasets. ### Citing OMat24 The OMat24 dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any publications that uses this dataset: ```bib @article{barroso_omat24, title={Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models}, author={Barroso-Luque, Luis and Muhammed, Shuaibi and Fu, Xiang and Wood, Brandon, Dzamba, Misko, and Gao, Meng and Rizvi, Ammar and Zitnick, C. Lawrence and Ulissi, Zachary W.}, journal={arXiv preprint arXiv:2410.12771}, year={2024} } @article{schmidt_2023_machine, title={Machine-Learning-Assisted Determination of the Global Zero-Temperature Phase Diagram of Materials}, author={Schmidt, Jonathan and Hoffmann, Noah and Wang, Hai-Chen and Borlido, Pedro and Carri{\c{c}}o, Pedro JMA and Cerqueira, Tiago FT and Botti, Silvana and Marques, Miguel AL}, journal={Advanced Materials}, volume={35}, number={22}, pages={2210788}, year={2023}, url={https://onlinelibrary.wiley.com/doi/full/10.1002/adma.202210788}, publisher={Wiley Online Library} } ``` --- Source: `docs/inorganic_materials/models.md` # Pretrained models :::{margin} ```{image} ../assets/icons/inorganic.svg :alt: Inorganic Materials :width: 100px ``` ::: :::{tip} Recommended Model **2025 recommendation:** We suggest using the [UMA model](../core/uma), trained on all of the FAIR chemistry datasets. UMA provides state-of-the-art accuracy, energy conservation, and will continue to receive updates. ::: The UMA model has a number of nice features over the previous checkpoints: 1. It is state-of-the-art in out-of-domain prediction accuracy 2. The UMA small model is an energy conserving and smooth checkpoint, so should work much better for vibrational calculations, molecular dynamics, etc. 3. The UMA model is most likely to be updated in the future. ## Legacy OMat pretrained models :::{note} These checkpoints are included here for baselining and model reproducibility. For new projects, we recommend using UMA. ::: * All config files for the OMat24 models are available in the [`configs/omat24`](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24) directory. * All models are equiformerV2 S2EFS models **Note** in order to download any of the model checkpoints from the links below, you will need to first request access through the [OMAT24 Hugging Face page](https://huggingface.co/fairchem/OMAT24). These checkpoints are trained on OMat24 only. Note that predictions are *not* Materials Project compatible. | Model Name | Checkpoint | Config | |-----------------------|--------------|---------------------------------------------------------------------------------------------| | EquiformerV2-31M-OMat | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_31M_omat.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/all/eqV2_31M.yml) | | EquiformerV2-86M-OMat | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_86M_omat.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/all/eqV2_86M.yml) | | EquiformerV2-153M-OMat | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_153M_omat.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/all/eqV2_153M.yml) | ## MPTrj only models These models are trained only on the [MPTrj](https://figshare.com/articles/dataset/Materials_Project_Trjectory_MPtrj_Dataset/23713842) dataset. | Model Name | Checkpoint | Config | |---------------------------|--------------|---------------------------------------------------------------------------------| | EquiformerV2-31M-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_31M_mp.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/mptrj/eqV2_31M_mptrj.yml) | | EquiformerV2-31M-DeNS-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_dens_31M_mp.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/mptrj/eqV2_31M_dens_mptrj.yml) | | EquiformerV2-86M-DeNS-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_dens_86M_mp.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/mptrj/eqV2_86M_dens_mptrj.yml) | | EquiformerV2-153M-DeNS-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_dens_153M_mp.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/mptrj/eqV2_153M_dens_mptrj.yml) | ## Finetuned OMat models These models are finetuned from the OMat pretrained checkpoints using MPTrj or MPTrj and sub-sampled trajectories from the 3D PBE Alexandria dataset, which we call Alex. | Model Name | Checkpoint | Config | |--------------------------------|--------------|------------------------------------------------------------------------------------| | EquiformerV2-31M-OMat-Alex-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_31M_omat_mp_salex.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/finetune/eqV2_31M_ft_salexmptrj.yml) | | EquiformerV2-86M-OMat-Alex-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24/blob/main/eqV2_86M_omat_mp_salex.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/finetune/eqV2_86M_ft_salexmptrj.yml) | | EquiformerV2-153M-OMat-Alex-MP | [checkpoint](https://huggingface.co/fairchem/OMAT24) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/omat24/finetune/eqV2_153M_ft_salexmptrj.yml) | Please consider citing the following work if you use OMat24 models in your work, ```bibtex @article{barroso-luqueOpenMaterials20242024, title = {Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models}, author = {Barroso-Luque, Luis and Shuaibi, Muhammed and Fu, Xiang and Wood, Brandon M. and Dzamba, Misko and Gao, Meng and Rizvi, Ammar and Zitnick, C. Lawrence and Ulissi, Zachary W.}, date = {2024-10-16}, eprint = {2410.12771}, eprinttype = {arXiv}, doi = {10.48550/arXiv.2410.12771}, url = {http://arxiv.org/abs/2410.12771}, } ``` --- Source: `docs/inorganic_materials/examples_tutorials/summary.md` # Examples & Tutorials :::{margin} ```{image} ../../assets/icons/inorganic.svg :alt: Inorganic Materials :width: 100px ``` ::: Learn how to use FAIRChem models for inorganic materials property prediction and simulation. ::::{grid} 1 2 2 3 :::{card} Formation Energy :link: formation-energy Calculate formation energies for inorganic materials using UMA and MP-compatible corrections. ::: :::{card} Phonons :link: phonons Run phonon calculations to predict thermal conductivity, vibrational modes, and finite-temperature stability. ::: :::{card} Elastic Tensors :link: elastic Calculate elastic constants and bulk modulus using automated strain-based workflows. ::: :::: --- Source: `docs/inorganic_materials/examples_tutorials/formation_energy.md` # Formation Energy :::{tip} What You Will Learn Calculate formation energies for inorganic materials using UMA with Materials Project-compatible corrections. ::: We're going to start simple here - let's run a local relaxation (optimize the unit cell and positions) using a pre-trained UMA model to compute formation energies for inorganic materials. Note predicting formation energy using models that models trained solely on OMat24 must use OMat24 compatible references and corrections for mixing PBE and PBE+U calculations. We use MP2020-style corrections fitted to OMat24 DFT calculations. For more information see the [documentation](https://docs.materialsproject.org/methodology/materials-methodology/thermodynamic-stability/thermodynamic-stability/anion-and-gga-gga+u-mixing) at the Materials Project. The necessary references can be found using the `fairchem.data.omat` package! ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-omat ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ```{code-cell} ipython3 from __future__ import annotations import pprint from ase.build import bulk from ase.optimize import FIRE from quacc.recipes.mlp.core import relax_job from quacc import flow from fairchem.core.calculate import FAIRChemCalculator, FormationEnergyCalculator # Make an Atoms object of a bulk Cu structure atoms = bulk("Cu") # Run a structure relaxation @flow def relax_flow(*args, **kwargs): return relax_job(*args, **kwargs) result = relax_flow( atoms, method="fairchem", name_or_path="uma-s-1p2p1", task_name="omat", relax_cell=True, opt_params={"fmax": 1e-3, "optimizer": FIRE}, ) # Get the realxed atoms! atoms = result["atoms"] # Create an calculator using uma-s-1p2p1 calculator = FAIRChemCalculator.from_model_checkpoint("uma-s-1p2p1", task_name="omat") # Now use the FormationEnergyCalculator to calculate the formation energy # This will now return MP-style corrected formation energies # For the omat task, this defaults to apply MP2020 style corrections with OMat24 compatibility form_e_calc = FormationEnergyCalculator(calculator, apply_corrections=True) atoms.calc = form_e_calc form_energy = atoms.get_potential_energy() ``` ```{code-cell} ipython3 pprint.pprint(f"Total energy: {result['results']['energy']} eV \n Formation energy {form_energy} eV") ``` Compare the results to the value of [-3.038 eV/atom reported](https://next-gen.materialsproject.org/materials/mp-1265?chemsys=Mg-O#thermodynamic_stability) in the Materials Project! *Note that we expect differences due to the different DFT settings used to calculate the OMat24 training data.* Congratulations; you ran your first relaxation and predicted the formation energy of MgO using UMA and `quacc`! --- Source: `docs/inorganic_materials/examples_tutorials/phonons.md` # Phonon Calculations :::{tip} What You Will Learn Run phonon calculations to predict thermal conductivity, vibrational modes, entropy, and finite-temperature stability. ::: Phonon calculations are very important for inorganic materials science to * Calculate thermal conductivity * Understand the vibrational modes, and thus entropy and free energy, of a material * Predict the stability of a material at finite temperature (e.g. 300 K) among many others! We can run a similarly straightforward calculation that 1. Runs a relaxation on the unit cell and atoms 2. Repeats the unit cell a number of times to make it sufficiently large to capture many interesting vibrational models 3. Generatives a number of finite displacement structures by moving each atom of the unit cell a little bit in each direction 4. Running single point calculations on each of (3) 5. Gathering all of the calculations and calculating second derivatives (the hessian matrix!) 6. Calculating the eigenvalues/eigenvectors of the hessian matrix to find the vibrational modes of the material 7. Analyzing the thermodynamic properties of the vibrational modes. Note that this analysis assumes that all vibrational modes are harmonic, which is a pretty reasonable approximately for low/moderate temperature materials, but becomes less realistic at high temperatures. ```{code-cell} ipython3 from __future__ import annotations from ase.build import bulk from quacc.recipes.mlp.phonons import phonon_flow # Make an Atoms object of a bulk Cu structure atoms = bulk("Cu") # Run a phonon (hessian) calculation with our favorite MLP potential result = phonon_flow( atoms, method="fairchem", job_params={ "all": dict( name_or_path="uma-s-1p2p1", task_name="omat", ), }, min_lengths=10.0, # set the minimum unit cell size smaller to be compatible with limited github runner ram ) ``` ```{code-cell} ipython3 print( f'The entropy at { result["results"]["thermal_properties"]["temperatures"][-1]:.0f} K is { result["results"]["thermal_properties"]["entropy"][-1]:.2f} kJ/mol' ) ``` Congratulations, you ran your first phonon calculation! ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` --- Source: `docs/inorganic_materials/examples_tutorials/elastic.md` # Elastic Tensors :::{tip} What You Will Learn Calculate elastic constants and bulk modulus using automated strain-based workflows with quacc. ::: Let's do something more interesting that normally takes quite a bit of work in DFT: calculating an elastic constant! Elastic properties are important to understand how strong or easy to deform a material is, or how a material might change if compressed or expanded in specific directions (i.e. the Poisson ratio!). We don't have to change much code from above, we just use a built-in recipe to calculate the elastic tensor from `quacc`. This recipe 1. (optionally) Relaxes the unit cell using the MLIP 2. Generates a number of deformed unit cells by applying strains 3. For each deformation, a relaxation using the MLIP and (optionally) a single point calculation is run 4. Finally, all of the above calculations are used to calculate the elastic properties of the material For more documentation, see the quacc docs for [quacc.recipes.mlp.elastic_tensor_flow](https://quantum-accelerators.github.io/quacc/reference/quacc/recipes/mlp/elastic.html#quacc.recipes.mlp.elastic.elastic_tensor_flow) ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ```{code-cell} ipython3 from __future__ import annotations from ase.build import bulk from quacc.recipes.mlp.elastic import elastic_tensor_flow # Make an Atoms object of a bulk Cu structure atoms = bulk("Cu") # Run an elastic property calculation with our favorite MLP potential result = elastic_tensor_flow( atoms, job_params={ "all": dict( method="fairchem", name_or_path="uma-s-1p2p1", task_name="omat", ), }, ) ``` ```{code-cell} ipython3 result["elasticity_doc"].bulk_modulus ``` Congratulations, you ran your first elastic tensor calculation! --- Source: `docs/catalysts/datasets/summary.md` # Datasets :::{margin} ```{image} ../../assets/icons/catalysis.svg :alt: Catalysis :width: 100px ``` ::: This section provides documentation for all catalyst-related datasets in the Open Catalyst Project. These datasets are used to train and evaluate machine learning models for heterogeneous catalysis applications. :::{tip} For most new users, we recommend starting with the [UMA model](../../core/uma), which has been trained on all of these datasets and provides state-of-the-art performance. ::: ::::{grid} 2 2 3 3 :::{grid-item-card} OC20 :link: oc20 The foundational Open Catalyst 2020 dataset with 133M+ DFT calculations for adsorbate-surface systems. ::: :::{grid-item-card} OC20-mAds :link: oc20_mads Multi-adsorbate extension of OC20 including coverage effects on catalyst surfaces. ::: :::{grid-item-card} OC20Dense :link: oc20dense Dense sampling of adsorbate configurations for adsorption energy calculations. ::: :::{grid-item-card} OC20NEB :link: oc20neb NEB trajectories for transition state calculations including desorptions, dissociations, and transfers. ::: :::{grid-item-card} OC22 :link: oc22 Open Catalyst 2022 dataset focusing on oxide electrocatalysts with total energy predictions. ::: :::{grid-item-card} OC25 :link: oc25 Solid-liquid interface dataset with 8M DFT calculations for electrocatalysis applications. ::: :::{grid-item-card} OCx24 :link: ocx24 Experimental validation dataset bridging computational and experimental catalysis research. ::: :::: --- Source: `docs/catalysts/datasets/oc20.md` # Open Catalyst 2020 (OC20) :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 133M+ DFT calculations | | **Systems** | ~460K adsorbate-catalyst relaxations | | **Tasks** | S2EF, IS2RE, IS2RS | | **Elements** | 55 elements from periodic table | | **Adsorbates** | 82 adsorbates | | **Paper** | [ACS Catalysis 2021](https://doi.org/10.1021/acscatal.0c04525) | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: :::{warning} Data files for all tasks/splits were updated on Feb 10, 2021 due to minor bugs (affecting < 1% of the data) in earlier versions. If you downloaded data before Feb 10, 2021, please re-download the data. ::: ## Download and preprocess the dataset IS2* datasets are stored as LMDB files and are ready to be used upon download. S2EF train+val datasets require an additional preprocessing step. For convenience, a self-contained script can be found [here](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/scripts/download_data.py) to download, preprocess, and organize the data directories to be readily usable by the existing [configs](https://github.com/facebookresearch/fairchem/tree/main/configs/oc20). For IS2*, run the script as: ``` python scripts/download_data.py --task is2re ``` For S2EF train/val, run the script as: ``` python scripts/download_data.py --task s2ef --split SPLIT_SIZE --get-edges --num-workers WORKERS --ref-energy ``` * `--split`: split size to download: `"200k", "2M", "20M", "all", "val_id", "val_ood_ads", "val_ood_cat", or "val_ood_both"`. * `--get-edges`: includes edge information in LMDBs (~10x storage requirement, ~3-5x slowdown), otherwise, compute edges on the fly (larger GPU memory requirement). * `--num-workers`: number of workers to parallelize preprocessing across. * `--ref-energy`: uses referenced energies instead of raw energies. For S2EF test, run the script as: ``` python scripts/download_data.py --task s2ef --split test ``` To download and process the dataset in a directory other than your local `fairchem/data` folder, add the following command line argument `--data-path`. Note that the baseline [configs](https://github.com/facebookresearch/fairchem/tree/main/configs/oc20). expect the data to be found in `fairchem/data`, make sure you symlink your directory or modify the paths in the configs accordingly. The following sections list dataset download links and sizes for various S2EF and IS2RE/IS2RS task splits. If you used the above `download_data.py` script to download and preprocess the data, you are good to go and can stop reading here! ## Structure to Energy and Forces (S2EF) task For this task’s train and validation sets, we provide compressed trajectory files with the input structures and output energies and forces. We provide precomputed LMDBs for the test sets. To use the train and validation datasets, first download the files and uncompress them. The uncompressed files are used to generate LMDBs, which are in turn used by the dataloaders to train the ML models. Code for the dataloaders and generating the LMDBs may be found in the Github repository. Four training datasets are provided with different sizes. Each is a subset of the other, i.e., the 2M dataset is contained in the 20M and all datasets. Four datasets are provided for validation set. Each dataset corresponds to a subsplit used to evaluate different types of extrapolation, in domain (id, same distribution as the training dataset), out of domain adsorbate (ood_ads, unseen adsorbate), out of domain catalyst (ood_cat, unseen catalyst composition), and out of domain both (ood_both, unseen adsorbate and catalyst composition). For the test sets, we provide precomputed LMDBs for each of the 4 subsplits (In Domain, OOD Adsorbate, OOD Catalyst, OOD Both). Each tarball has a README file containing details about file formats, number of structures / trajectories, etc. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train | | | | | |all |225G |1.1T | [12a7087bfd189a06ccbec9bc7add2bcd](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_train_all.tar) | |20M |34G |165G | [863bc983245ffc0285305a1850e19cf7](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_train_20M.tar) | |2M |3.4G |17G | [953474cb93f0b08cdc523399f03f7c36](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_train_2M.tar) | |200K |344M |1.7G | [f8d0909c2623a393148435dede7d3a46](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_train_200K.tar) | | | | | | | |Validation | | | | | |val_id |1.7G |8.3G | [f57f7f5c1302637940f2cc858e789410](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_val_id.tar) | |val_ood_ads |1.7G |8.2G | [431ab0d7557a4639605ba8b67793f053](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_val_ood_ads.tar) | |val_ood_cat |1.7G |8.3G | [532d6cd1fe541a0ddb0aa0f99962b7db](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_val_ood_cat.tar) | |val_ood_both |1.9G |9.5G | [5731862978d80502bbf7017d68c2c729](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_val_ood_both.tar) | | | | | | | |Test (LMDBs for all splits) |30G |415G | [bcada432482f6e87b24e14b6b744992a](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_test_lmdbs.tar.gz) | | | | | | | |Rattled data |29G |136G | [40431149b27b64ce1fb40cac4e2e064b](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_rattled.tar) | | | | | | | |MD data |42G |306G | [9fed845aaab8fb4bf85e3a8db57796e0](https://dl.fbaipublicfiles.com/opencatalystproject/data/s2ef_md.tar) | | | | | | ## Initial Structure to Relaxed Structure (IS2RS) and Initial Structure to Relaxed Energy (IS2RE) tasks For the IS2RS and IS2RE tasks, we are providing: * One `.tar.gz` file with precomputed LMDBs which once downloaded and uncompressed, can be used directly to train ML models. The LMDBs contain the input initial structures and the output relaxed structures and energies. Training datasets are split by size, with each being a subset of the larger splits, similar to S2EF. The validation and test datasets are broken into subsplits based on different extrapolation evaluations (In Domain, OOD Adsorbate, OOD Catalyst, OOD Both). * underlying ASE relaxation trajectories for the adsorbate+catalyst in the entire training and validation sets for the IS2RE and IS2RS tasks. These are **not** required to download for training ML models, but are available for interested users. Each tarball has README file containing details about file formats, number of structures / trajectories, etc. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train (all splits) + Validation (all splits) + test (all splits) |8.1G |97G | [cfc04dd2f87b4102ab2f607240d25fb1](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_train_val_test_lmdbs.tar.gz) | |Test-challenge 2021 ([challenge details](https://opencatalystproject.org/challenge.html)) |1.3G |17G | [aed414cdd240fbb5670b5de6887a138b](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2re_test_challenge_2021.tar.gz) | | | | | | ## Relaxation Trajectories ### Adsorbate+catalyst system trajectories (optional download) |Split |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |All IS2RE/S training (~466k trajectories) |109G |841G | [9e3ed4d1e497bfdce4472ee70455edef](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_train_trajectories.tar) | | | | | | |IS2RE/S Validation | | | | |val_id (~25K trajectories) |5.9G |46G | [fcb71363018fb1e7127db2500e39e11a](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_val_id_trajectories.tar) | |val_ood_ads (~25K trajectories) |5.7G |44G | [5ced8ea84584aa229d31e693e0fb090f](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_val_ood_ads_trajectories.tar) | |val_ood_cat (~25K trajectories) |6.0G |46G | [88dcc02fd8c174a72d2c416878fc44ff](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_val_ood_cat_trajectories.tar) | |val_ood_both (~25K trajectories) |4.4G |35G | [bc74b6474a13542cc56eaa97bd51adfc](https://dl.fbaipublicfiles.com/opencatalystproject/data/is2res_val_ood_both_trajectories.tar) | #### Per-adsorbate trajectories (optional download) Download links are in the table below: |Adsorbate symbol |Downloadable path |size |MD5 checksum | |--- |--- |--- |--- | |*O |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/0.tar |1006M |d4151542856b4b6405f276808f75358a | |*H |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/1.tar |850M |3697f04faf04251a23da8b88a78209f7 | |*OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/2.tar |1.6G |a21081f3f55eb0c98a91021bbe3dac44 | |*OH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/3.tar |1.8G |b12b706854f5d899e02a9ae6578b5d45 | |*C |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/4.tar |1.1G |e4fe9890764fcf59e01e3ceab089b978 | |*CH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/6.tar |1.4G |ec9aa2c4c4bd4419359438ba7fbb881d | |*CHO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/7.tar |1.4G |d32200f74ad5c3bfd42e8835f36d57ab | |*COH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/8.tar |1.6G |5418a1b331f6c7689a5405cca4cc8d15 | |*CH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/9.tar |1.6G |8ee1066149c305d7c17c219b369c5a73 | |*CH2*O |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/10.tar |1.7G |960c2450814024b66f3c79121179ac60 | |*CHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/11.tar |1.8G |60ac9f965f9589a3389483e3d1e58144 | |*CH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/12.tar |1.7G |7e123e6f4fb10d6897be3f47721dfd4a | |*OCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/13.tar |1.8G |0823047bbbe05fa0e63f9d83ec601487 | |*CH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/14.tar |1.9G |9ac71e198d75b1427182cd34abb73e4d | |*CH4 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/15.tar |1.9G |a405ce403018bf8afbd4425d5c0b34d5 | |*OHCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/16.tar |2.1G |d3c829f1952db6e4f428273ee05f59b1 | |*C*C |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/17.tar |1.5G |d687a151345305897b9245af4b0f9967 | |*CCO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/18.tar |1.7G |214ca96e620c5ec6e8a6ff8144a22a04 | |*CCH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/19.tar |1.6G |da2268545e80ca1664026449dd2fdd24 | |*CHCO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/20.tar |1.7G |386c99407fe63080d26cda525dfdd8cd | |*CCHO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/21.tar |1.8G |918b20960438494ab160a9dbd9668157 | |*COCHO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/22.tar |1.8G |84424aa2ad30301e23ece1438ea39923 | |*CCHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/23.tar |2.0G |3cc90425ec042a70085ba7eb2916a79a | |*CCH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/24.tar |1.8G |9dbcf7566e40965dd7f8a186a75a718e | |*CH*CH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/25.tar |1.7G |a193b4c72f915ba0b21a41790696b23c | |CH2*CO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/26.tar |1.8G |de83cf50247f5556fa4f9f64beff1eeb | |*CHCHO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/27.tar |1.9G |1d140aaa2e7b287124ab38911a711d70 | |*CH*COH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/28.tar |1.3G |682d8a6b05ca5948b34dc5e5f6bbcd61 | |*COCH2O |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/29.tar |1.9G |c8742faa8ca40e8edb4110069817fa70 | |*CHO*CHO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/30.tar |2.0G |8cfbb67beb312b98c40fcb891dfa480a | |*COHCHO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/31.tar |1.9G |6ffa903a62d8ec3319ecec6a03b06276 | |*COHCOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/32.tar |2.0G |caca0058b641bfdc9f8de4527e60feb7 | |*CCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/33.tar |1.8G |906543aaefc171edab388ff4f0fe8a20 | |*CHCH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/34.tar |1.8G |4dfab479495f76179749c1956046fbd8 | |*COCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/35.tar |1.9G |29d1b992715054e920e8bb2afe97b393 | |*CHCHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/38.tar |2.0G |9e5912df6f7b11706d1046cdb9e3087e | |*CCH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/39.tar |2.1G |7bcae43cee451306e34ec416588a7f09 | |*CHOCHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/40.tar |2.0G |f98866d08fe3451ae7ebc47bb51599aa | |*COCH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/41.tar |1.4G |bfaf689e5827fcf26c51e567bb8dd1be | |*COHCHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/42.tar |2.0G |236fe4e950aa2fbdde94ef2821fb48d2 | |*OCHCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/44.tar |2.1G |66acc5460a999625c3364f0f3bcca871 | |*COHCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/45.tar |2.1G |bb4a01956736399c8cee5e219f8c1229 | |*CHOHCH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/46.tar |2.1G |e836de4ec146b1b611533f1ef682cace | |*CHCH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/47.tar |2.0G |66df44121806debef6dc038df7115d1d | |*OCH2CHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/48.tar |2.2G |ff6981fdbcd2e65d351505c15d218d76 | |*CHOCH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/49.tar |2.1G |448f7d352ab6e32f754e24de64ca302a | |*COHCH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/50.tar |2.1G |8bff6bf3e10cc84acc4a283a375fcc23 | |*CHOHCHOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/51.tar |2.0G |9c9e4d617d306751760a80f1453e71f1 | |*CH2CH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/52.tar |2.0G |ec1e964d2ee6f468fa5773743e3994a4 | |*OCH2CH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/53.tar |2.1G |d297b27b02822f9b6af80bdb64aee819 | |*CHOHCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/54.tar |2.1G |368de083dafdc3bbdb560d35e2a102c0 | |*CH2CH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/55.tar |2.1G |3c1aaf790659f7ff89bf1eed8b396b63 | |*CHOHCH2OH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/56.tar |2.2G |2d71adb9e305e6f3bca49e5df9b5a86a | |*OHCH2CH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/57.tar |2.3G |cf51128f8522b7b66fc68d79980d6def | |*NH2N(CH3)2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/58.tar |1.6G |36ba974d80c20ff636431f7c0ad225da | |*ONN(CH3)2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/59.tar |2.3G |fdc4cd19977496909d61be4aee61c4f1 | |*OHNNCH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/60.tar |2.1G |50a6ff098f9ba7adbba9ac115726cc5a | |*ONH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/62.tar |1.8G |47573199c545afe46c554ff756c3e38f | |*NHNH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/63.tar |1.7G |dd456b7e19ef592d9f0308d911b91d7c | |*N*NH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/65.tar |1.6G |c05289fd56d64c74306ebf57f1061318 | |*NO2NO2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/67.tar |2.1G |4822a06f6c5f41bdefd3cbbd8856c11f | |*N*NO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/68.tar |1.6G |2a27de122d32917cc5b6ac0a21c63c1c | |*N2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/69.tar |1.5G |cc668fecf679b6edaac8fd8fb9cdd404 | |*ONNH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/70.tar |2.1G |dff880f1a5baa7f67b52fd3ed745443d | |*NH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/71.tar |1.6G |c7f383b50faa6244e265c9611466cb8f | |*NH3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/72.tar |1.9G |2b355741f9300445703270e0e4b8c01c | |*NONH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/73.tar |1.8G |48877a0c6f2994baac82cb722711aaa2 | |*NH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/74.tar |1.4G |7979b9e7ab557d6979b33e352486f0ef | |*NO2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/75.tar |1.7G |9f352fbc32bb2b8caf4788aba28b2eb7 | |*NO |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/76.tar |1.4G |482ee306a5ae2eee78cac40d10059ebc | |*N |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/77.tar |1.1G |bfb6e03d4a687987ff68976f0793cc46 | |*NO3 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/78.tar |1.8G |700834326e789a6e38bf3922d9fcb792 | |*OHNH2 |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/79.tar |2.1G |fa24472e0c02c34d91f3ffe6b77bfb11 | |*ONOH |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/80.tar |1.4G |4ddcccd62a834a76fe6167461f512529 | |*CN |https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/81.tar |1.5G |bc7c55330ece006d09496a5ff01d5d50 | Note - A few adsorbates are intentionally left out for the test splits. Downloading any of the above and extracting will result in a folder : `/` * `system.txt` Text file containing information about the different adsorbate+catalyst system names. In total there are N systems. More details described below. * `/` * This contains N compressed trajectory files of the format `.extxyz.xz`. * Files are named as `.extxyz.xz` (where `system_id` is defined below). where, `` can be 0 to 81. N is dependent on which adsorbate index is chosen. The file `system.txt` has information in the following format: `system_id,reference_energy` where: * `system_id `- Internal random ID corresponding to an adsorbate+catalyst system. * `reference_energy` - Energy used to reference system energies to bare catalyst+gas reference energies. Used for adsorption energy calculations. The `.extxyz.xz` files are LZMA compressed `.extxyz` trajectory files. Each trajectory corresponds to a relaxation trajectory of a different adsorbate+catalyst system. Information about the `.extxyz` trajectory file format may be found at https://wiki.fysik.dtu.dk/ase/dev/ase/io/formatoptions.html#extxyz . In order to uncompress the files, `uncompress.py` provides a multi-core implementation which could be used. ### Catalyst system trajectories (optional download) |Number |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) |MD5 checksum (download link) | |--- |--- |--- |--- | |294k systems |20G |151G | [347f4183465810e9b384e7a033baefc7](https://dl.fbaipublicfiles.com/opencatalystproject/data/slab_trajectories.tar) | ## Bader charge data We provide Bader charge data for all final frames of our train + validation systems in OC20 (for both S2EF and IS2RE/RS tasks). A `.tar.gz` file, when downloaded and uncompressed will contain several directories with unique system-ids (of the format `random` where `XYZ` is an integer). Each directory will contain raw Bader charge analysis outputs. For more details on the Bader charge analysis, see https://theory.cm.utexas.edu/henkelman/research/bader/. Downloadable link: https://dl.fbaipublicfiles.com/opencatalystproject/data/oc20_bader_data.tar (MD5 checksum: `aecc5e23542de49beceb4b7e44c153b9`) ## OC20 mappings ### Data mapping information We provide a Python pickle file containing information about the slab and adsorbates for each of the systems in OC20 dataset. Loading the pickle file will load a Python dictionary. The keys of this dictionary are the adsorbate+catalyst system-ids (of the format `random` where `XYZ` is an integer), and the corresponding value of each key is a dictionary with information about: * `bulk_mpid` : Materials Project ID of the bulk system used corresponding to the catalyst surface * `bulk_symbols` Chemical composition of the bulk counterpart * `ads_symbols` Chemical composition of the adsorbate counterpart * `ads_id` : internal unique identifier, one for each of the 82 adsorbates used in the dataset * `bulk_id` : internal unique identifier one for each of the 11500 bulks used in the dataset * `miller_index`: 3-tuple of integers indicating the Miller indices of the surface * `shift`: c-direction shift used to determine cutoff for the surface (c-direction is following the nomenclature from Pymatgen) * `top`: boolean indicating whether the chosen surface was at the top or bottom of the originally enumerated surface * `adsorption_site`: A tuple of 3-tuples containing the Cartesian coordinates of each binding adsorbate atom * `class`: integer indicating the class of materials the system's slab is part of, where: * 0 - intermetallics * 1 - metalloids * 2 - non-metals * 3 - halides * `anomaly`: integer indicating possible anomalies (based off general heuristics, not to be taken as perfect classifications), where: * 0 - no anomaly * 1 - adsorbate dissociation * 2 - adsorbate desorption * 3 - surface reconstruction * 4 - incorrect CHCOH placement, appears to be CHCO with a lone, uninteracting, H far off in the unit cell Downloadable link: https://dl.fbaipublicfiles.com/opencatalystproject/data/oc20_data_mapping.pkl (MD5 checksum: `6b5d485019861f6e7efca38338375b61`) An example entry is ``` 'random2181546': {'bulk_id': 6510, 'ads_id': 69, 'bulk_mpid': 'mp-22179', 'bulk_symbols': 'Si2Ti2Y2', 'ads_symbols': '*N2', 'miller_index': (2, 0, 1), 'shift': 0.145, 'top': True, 'adsorption_site': ((4.5, 12.85, 16.13),), 'class': 1, 'anomaly': 0} ``` ## Adsorbate-catalyst system to catalyst system mapping information We provide a Python pickle file containing information about the mapping from adsorbate-catalyst systems to their corresponding catalyst systems. Loading the pickle file will load a Python dictionary. The keys of this dictionary are the adsorbate+catalyst system-ids (of the format `random` where `XYZ` is an integer), and values will be the catalyst system-ids (of the format `random` where `PQR` is an integer). Downloadable link: https://dl.fbaipublicfiles.com/opencatalystproject/data/mapping_adslab_slab.pkl (MD5 checksum: `079041076c3f15d18ecb5d17c509cdfe`) An example entry is ``` 'random1981709': 'random533137' ``` ## Dataset changelog ### September 2021 * Released IS2RE `test-challenge` data for the [Open Catalyst Challenge 2021](https://opencatalystproject.org/challenge.html) ### March 2021 * Modified the pickle corresponding to data mapping information. Now the pickle includes extra information about `miller_index`, `shift`, `top` and `adsorption_site`. * Added Molecular Dynamics (MD) and rattled data for S2EF task. ### Version 2, Feb 2021 Modifications: * Removed slab systems which had single frame checkpoints, this led to modifications of reference frame energies of 350k frames out of 130M. * Fixed stitching of checkpoints in adsorbate+catalyst trajectories. * Added release of slab trajectories. Below are actual updates numbers, of the form `old` → `new` Total S2EF frames: * train: 133953162 → 133934018 * validation: * val_id : 1000000 → 999866 * val_ood_ads: 1000000 → 999838 * val_ood_cat: 1000000 → 999809 * val_ood_both: 1000000 → 999944 * test: * test_id: 1000000 → 999736 * test_ood_ads: 1000000 → 999859 * test_ood_cat: 1000000 → 999826 * test_ood_both: 1000000 → 999973 Total IS2RE and IS2RS systems: * train: 461313 → 460328 * validation: * val_id : 24946 → 24943 * val_ood_ads: 24966 → 24961 * val_ood_cat: 24988 → 24963 * val_ood_both: 24963 → 24987 * test: * test_id: 24951 → 24948 * test_ood_ads: 24931 → 24930 * test_ood_cat: 24967 → 24965 * test_ood_both: 24986 → 24985 ### Version 1, Oct 2020 Total S2EF frames: * train: 133953162 * validation: * val_id : 1000000 * val_ood_ads: 1000000 * val_ood_cat: 1000000 * val_ood_both: 1000000 * test: * test_id: 1000000 * test_ood_ads: 1000000 * test_ood_cat: 1000000 * test_ood_both: 1000000 Total IS2RE and IS2RS systems: * train: 461313 * validation: * val_id : 24936 * val_ood_ads: 24966 * val_ood_cat: 24988 * val_ood_both: 24963 * test: * test_id: 24951 * test_ood_ads: 24931 * test_ood_cat: 24967 * test_ood_both: 24986 ## Citing OC20 The Open Catalyst 2020 (OC20) dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any research manuscript using the OC20 dataset: ```bibtex @article{ocp_dataset, author = {Chanussot*, Lowik and Das*, Abhishek and Goyal*, Siddharth and Lavril*, Thibaut and Shuaibi*, Muhammed and Riviere, Morgane and Tran, Kevin and Heras-Domingo, Javier and Ho, Caleb and Hu, Weihua and Palizhati, Aini and Sriram, Anuroop and Wood, Brandon and Yoon, Junwoong and Parikh, Devi and Zitnick, C. Lawrence and Ulissi, Zachary}, title = {Open Catalyst 2020 (OC20) Dataset and Community Challenges}, journal = {ACS Catalysis}, year = {2021}, doi = {10.1021/acscatal.0c04525}, } ``` # Per-adsorbate trajectories |Adsorbate symbol |Size |MD5 checksum (download link) | |--- |--- |--- | |*O |1006M |[d4151542856b4b6405f276808f75358a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/0.tar) | |*H |850M |[3697f04faf04251a23da8b88a78209f7](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/1.tar) | |*OH |1.6G |[a21081f3f55eb0c98a91021bbe3dac44](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/2.tar) | |*OH2 |1.8G |[b12b706854f5d899e02a9ae6578b5d45](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/3.tar) | |*C |1.1G |[e4fe9890764fcf59e01e3ceab089b978](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/4.tar) | |*CH |1.4G |[ec9aa2c4c4bd4419359438ba7fbb881d](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/6.tar) | |*CHO |1.4G |[d32200f74ad5c3bfd42e8835f36d57ab](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/7.tar) | |*COH |1.6G |[5418a1b331f6c7689a5405cca4cc8d15](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/8.tar) | |*CH2 |1.6G |[8ee1066149c305d7c17c219b369c5a73](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/9.tar) | |*CH2*O |1.7G |[960c2450814024b66f3c79121179ac60](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/10.tar) | |*CHOH |1.8G |[60ac9f965f9589a3389483e3d1e58144](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/11.tar) | |*CH3 |1.7G |[7e123e6f4fb10d6897be3f47721dfd4a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/12.tar) | |*OCH3 |1.8G |[0823047bbbe05fa0e63f9d83ec601487](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/13.tar) | |*CH2OH |1.9G |[9ac71e198d75b1427182cd34abb73e4d](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/14.tar) | |*CH4 |1.9G |[a405ce403018bf8afbd4425d5c0b34d5](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/15.tar) | |*OHCH3 |2.1G |[d3c829f1952db6e4f428273ee05f59b1](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/16.tar) | |*C*C |1.5G |[d687a151345305897b9245af4b0f9967](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/17.tar) | |*CCO |1.7G |[214ca96e620c5ec6e8a6ff8144a22a04](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/18.tar) | |*CCH |1.6G |[da2268545e80ca1664026449dd2fdd24](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/19.tar) | |*CHCO |1.7G |[386c99407fe63080d26cda525dfdd8cd](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/20.tar) | |*CCHO |1.8G |[918b20960438494ab160a9dbd9668157](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/21.tar) | |*COCHO |1.8G |[84424aa2ad30301e23ece1438ea39923](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/22.tar) | |*CCHOH |2.0G |[3cc90425ec042a70085ba7eb2916a79a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/23.tar) | |*CCH2 |1.8G |[9dbcf7566e40965dd7f8a186a75a718e](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/24.tar) | |*CH*CH |1.7G |[a193b4c72f915ba0b21a41790696b23c](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/25.tar) | |CH2*CO |1.8G |[de83cf50247f5556fa4f9f64beff1eeb](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/26.tar) | |*CHCHO |1.9G |[1d140aaa2e7b287124ab38911a711d70](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/27.tar) | |*CH*COH |1.3G |[682d8a6b05ca5948b34dc5e5f6bbcd61](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/28.tar) | |*COCH2O |1.9G |[c8742faa8ca40e8edb4110069817fa70](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/29.tar) | |*CHO*CHO |2.0G |[8cfbb67beb312b98c40fcb891dfa480a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/30.tar) | |*COHCHO |1.9G |[6ffa903a62d8ec3319ecec6a03b06276](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/31.tar) | |*COHCOH |2.0G |[caca0058b641bfdc9f8de4527e60feb7](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/32.tar) | |*CCH3 |1.8G |[906543aaefc171edab388ff4f0fe8a20](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/33.tar) | |*CHCH2 |1.8G |[4dfab479495f76179749c1956046fbd8](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/34.tar) | |*COCH3 |1.9G |[29d1b992715054e920e8bb2afe97b393](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/35.tar) | |*CHCHOH |2.0G |[9e5912df6f7b11706d1046cdb9e3087e](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/38.tar) | |*CCH2OH |2.1G |[7bcae43cee451306e34ec416588a7f09](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/38.tar) | |*CHOCHOH |2.0G |[f98866d08fe3451ae7ebc47bb51599aa](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/40.tar) | |*COCH2OH |1.4G |[bfaf689e5827fcf26c51e567bb8dd1be](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/41.tar) | |*COHCHOH |2.0G |[236fe4e950aa2fbdde94ef2821fb48d2](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/42.tar) | |*OCHCH3 |2.1G |[66acc5460a999625c3364f0f3bcca871](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/44.tar) | |*COHCH3 |2.1G |[bb4a01956736399c8cee5e219f8c1229](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/45.tar) | |*CHOHCH2 |2.1G |[e836de4ec146b1b611533f1ef682cace](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/46.tar) | |*CHCH2OH |2.0G |[66df44121806debef6dc038df7115d1d](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/47.tar) | |*OCH2CHOH |2.2G |[ff6981fdbcd2e65d351505c15d218d76](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/48.tar) | |*CHOCH2OH |2.1G |[448f7d352ab6e32f754e24de64ca302a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/49.tar) | |*COHCH2OH |2.1G |[8bff6bf3e10cc84acc4a283a375fcc23](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/50.tar) | |*CHOHCHOH |2.0G |[9c9e4d617d306751760a80f1453e71f1](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/51.tar) | |*CH2CH3 |2.0G |[ec1e964d2ee6f468fa5773743e3994a4](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/52.tar) | |*OCH2CH3 |2.1G |[d297b27b02822f9b6af80bdb64aee819](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/53.tar) | |*CHOHCH3 |2.1G |[368de083dafdc3bbdb560d35e2a102c0](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/54.tar) | |*CH2CH2OH |2.1G |[3c1aaf790659f7ff89bf1eed8b396b63](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/55.tar) | |*CHOHCH2OH |2.2G |[2d71adb9e305e6f3bca49e5df9b5a86a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/56.tar) | |*OHCH2CH3 |2.3G |[cf51128f8522b7b66fc68d79980d6def](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/57.tar) | |*NH2N(CH3)2 |1.6G |[36ba974d80c20ff636431f7c0ad225da](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/58.tar) | |*ONN(CH3)2 |2.3G |[fdc4cd19977496909d61be4aee61c4f1](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/59.tar) | |*OHNNCH3 |2.1G |[50a6ff098f9ba7adbba9ac115726cc5a](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/60.tar) | |*ONH |1.8G |[47573199c545afe46c554ff756c3e38f](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/62.tar) | |*NHNH |1.7G |[dd456b7e19ef592d9f0308d911b91d7c](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/63.tar) | |*N*NH |1.6G |[c05289fd56d64c74306ebf57f1061318](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/65.tar) | |*NO2NO2 |2.1G |[4822a06f6c5f41bdefd3cbbd8856c11f](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/67.tar) | |*N*NO |1.6G |[2a27de122d32917cc5b6ac0a21c63c1c](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/68.tar) | |*N2 |1.5G |[cc668fecf679b6edaac8fd8fb9cdd404](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/69.tar) | |*ONNH2 |2.1G |[dff880f1a5baa7f67b52fd3ed745443d](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/70.tar) | |*NH2 |1.6G |[c7f383b50faa6244e265c9611466cb8f](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/71.tar) | |*NH3 |1.9G |[2b355741f9300445703270e0e4b8c01c](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/72.tar) | |*NONH |1.8G |[48877a0c6f2994baac82cb722711aaa2](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/73.tar) | |*NH |1.4G |[7979b9e7ab557d6979b33e352486f0ef](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/74.tar) | |*NO2 |1.7G |[9f352fbc32bb2b8caf4788aba28b2eb7](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/75.tar) | |*NO |1.4G |[482ee306a5ae2eee78cac40d10059ebc](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/76.tar) | |*N |1.1G |[bfb6e03d4a687987ff68976f0793cc46](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/77.tar) | |*NO3 |1.8G |[700834326e789a6e38bf3922d9fcb792](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/78.tar) | |*OHNH2 |2.1G |[fa24472e0c02c34d91f3ffe6b77bfb11](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/79.tar) | |*ONOH |1.4G |[4ddcccd62a834a76fe6167461f512529](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/80.tar) | |*CN |1.5G |[bc7c55330ece006d09496a5ff01d5d50](https://dl.fbaipublicfiles.com/opencatalystproject/data/per_adsorbate_is2res/81.tar) | Note - A few adsorbates are intentionally left out for the test splits. Downloading any of the above and extracting will result in a folder: `/` * `system.txt` Text file containing information about the different adsorbate+catalyst system names. In total there are N systems. More details described below. * `/` * This contains N compressed trajectory files of the format `.extxyz.xz`. * Files are named as `.extxyz.xz` (where `system_id` is defined below). where, `` can be 0 to 81. N is dependent on which adsorbate index is chosen. The file `system.txt` has information in the following format: `system_id,reference_energy` where: * `system_id `- Internal random ID corresponding to an adsorbate+catalyst system. * `reference_energy` - Energy used to reference system energies to bare catalyst+gas reference energies. Used for adsorption energy calculations. The `.extxyz.xz` files are LZMA compressed `.extxyz` trajectory files. Each trajectory corresponds to a relaxation trajectory of a different adsorbate+catalyst system. Information about the `.extxyz` trajectory file format may be found at https://wiki.fysik.dtu.dk/ase/ase/io/formatoptions.html#extxyz. In order to uncompress the files, `uncompress.py` provides a multi-core implementation which could be used. --- Source: `docs/catalysts/datasets/oc20_mads.md` # Open Catalyst 2020 Multi-Adsorbate (mAds) Dataset :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 21.8M structures | | **Max Adsorbates** | Up to 5 adsorbates per surface | | **Purpose** | Multi-adsorbate and coverage effects | | **Paper** | [UMA Paper](https://arxiv.org/pdf/2506.23971) | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: ## Overview The OC20-mAds dataset is a training set expanding the original OC20 dataset to include multi-adsorbate and coverage effects on catalyst surfaces. Adsorbates are randomly sampled from the list of OC20 adsorbates, up to 5 maximum adsorbates. For a small fraction of the dataset, all adsorbates on the surface may be identical. OC20-mAds is introduced in the [UMA paper](https://arxiv.org/pdf/2506.23971). ## File Contents and Download |Splits |Size | MD5 checksum (download link) | |--- |--- |--- | |Train | 21,804,758 | [6435960ba5ad1a7c949bd2f2b51825bc](https://dl.fbaipublicfiles.com/opencatalystproject/data/oc20mAds/oc20_multiads_train.tar.gz) | The following metadata can be accessed in the respective `atoms.info` entry: - `bulk_id`: Bulk identifier - `millers`: 3-tuple of integers indicating the Miller indices of the surface. - `shift`: C-direction shift used to determine cutoff for the surface (c-direction is following the nomenclature from Pymatgen). - `top`: Boolean indicating whether the chosen surface was at the top or bottom of the originally enumerated surface. - `adsorbates`: List of adsorbates sampled and their respective placements. - `sid`: Unique system identifier. - `fid`: Frame index along the relaxation/AIMD trajectory. - `results_path`: Internal results location. - `fmax`: Max per-atom force. --- Source: `docs/catalysts/datasets/oc22.md` # Open Catalyst 2022 (OC22) :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | ~62K oxide systems | | **Tasks** | S2EF-Total, IS2RE-Total, IS2RS | | **Focus** | Oxide electrocatalysts | | **Energy Type** | DFT total energies | | **Paper** | [ACS Catalysis 2023](https://doi.org/10.1021/acscatal.2c05426) | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: :::{note} OC22 models are trained on DFT total energies, in contrast to OC20 models which are trained on adsorption energies. ::: ## Structure to Total Energy and Forces (S2EF-Total) task For this task’s train, validation and test sets, we provide precomputed LMDBs that can be directly used with dataloaders provided in our code. The LMDBs contain input structures from all points in relaxation trajectories along with the energy of the structure and the atomic forces. The validation and test datasets are broken into subsplits based on in-distribution and out-of-distribution materials relative to the training dataset. All LMDBs are compressed into a single `.tar.gz` file. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train (all splits) + Validation (all splits) + test (all splits) | 20G | 71G | [ebea523c6f8d61248a37b4dd660b11e6](https://dl.fbaipublicfiles.com/opencatalystproject/data/oc22/s2ef_total_train_val_test_lmdbs.tar.gz) | | | | | ## Initial Structure to Relaxed Structure (IS2RS) and Initial Structure to Relaxed Total Energy (IS2RE-Total) tasks For IS2RE-Total / IS2RS training, validation and test sets, we provide precomputed LMDBs that can be directly used with dataloaders provided in our code. The LMDBs contain input initial structures and the output relaxed structures and energies. The validation and test datasets are broken into subsplits based on in-distribution and out-of-distribution materials relative to the training dataset. All LMDBs are compressed into a single `.tar.gz` file. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train (all splits) + Validation (all splits) + test (all splits) | 109M | 424M | [b35dc24e99ef3aeaee6c5c949903de94](https://dl.fbaipublicfiles.com/opencatalystproject/data/oc22/is2res_total_train_val_test_lmdbs.tar.gz) | | | | | | ## Relaxation Trajectories ### System trajectories (optional download) We provide relaxation trajectories for all systems used in train and validation sets of S2EF-Total and IS2RE-Total/RS task: |Number |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | | S2EF and IS2RE (both train and validation) | 34G | 80G | [977b6be1cbac6864e63c4c7fbf8a3fce](https://dl.fbaipublicfiles.com/opencatalystproject/data/oc22/oc22_trajectories.tar.gz) | | | | | | ## OC22 Mappings ### Data mapping information We provide a Python pickle file containing information about the slab and adsorbates for each of the systems in OC22 dataset. Loading the pickle file will load a Python dictionary. The keys of this dictionary are the system-ids (of the format `XYZ` where `XYZ` is an integer, corresponding to the `sid` in the LMDB Data object), and the corresponding value of each key is a dictionary with information about: * `bulk_id`: Materials Project ID of the bulk system used corresponding to the catalyst surface * `bulk_symbols`: Chemical composition of the bulk counterpart * `miller_index`: 3-tuple of integers indicating the Miller indices of the surface * `traj_id`: Identifier associated with the accompanying raw trajectory (if available) * `slab_sid`: Identifier associated with the corresponding slab (if available) * `ads_symbols`: Chemical composition of the adsorbate counterpart (adosrbate+slabs only) * `nads`: Number of adsorbates present Downloadable link: https://dl.fbaipublicfiles.com/opencatalystproject/data/oc22/oc22_metadata.pkl (MD5 checksum: `13dc06c6510346d8a7f614d5b26c8ffa` ) An example adsorbate+slab entry: ``` 6877: {'bulk_id': 'mp-559112', 'miller_index': (1, 0, 0), 'nads': 1, 'traj_id': 'K2Zn6O7_mp-559112_RyQXa0N0uc_ohyUKozY3G', 'bulk_symbols': 'K4Zn12O14', 'slab_sid': 30859, 'ads_symbols': 'O2'}, ``` An example slab entry: ``` 34815: {'bulk_id': 'mp-18793', 'miller_index': (1, 2, 1), 'nads': 0, 'traj_id': 'LiCrO2_mp-18793_clean_3HDHBg6TIz', 'bulk_symbols': 'Li2Cr2O4'}, ``` ### ### OC20 reference information In order to train models on OC20 total energy, we provide a Python pickle file containing the energy necessary to convert adsorption energy values to total energy. Loading the pickle file will load a Python dictionary. The keys of this dictionary are the system-ids (of the format `random` where `XYZ` is an integer, corresponding to the `sid` in the LMDB Data object), and the corresponding value of each key is the energy to be added to OC20 energy values. To train on total energies for OC20, specify the path to this pickle file in your training configs. Downloadable link: https://dl.fbaipublicfiles.com/opencatalystproject/data/oc22/oc20_ref.pkl (MD5 checksum: `043e1e0b0cce64c62f01a8563dbc3178`) ### ## Citing OC22 The Open Catalyst 2022 (OC22) dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any research manuscript using the OC22 dataset: ```bibtex @article{oc22_dataset, author = {Tran*, Richard and Lan*, Janice and Shuaibi*, Muhammed and Wood*, Brandon and Goyal*, Siddharth and Das, Abhishek and Heras-Domingo, Javier and Kolluru, Adeesh and Rizvi, Ammar and Shoghi, Nima and Sriram, Anuroop and Ulissi, Zachary and Zitnick, C. Lawrence}, title = {The Open Catalyst 2022 (OC22) dataset and challenges for oxide electrocatalysts}, journal = {ACS Catalysis}, year={2023}, } ``` --- Source: `docs/catalysts/datasets/oc20dense.md` # Open Catalyst 2020 Dense (OC20Dense) :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 85,658 unique configurations | | **Systems** | ~1,000 adsorbate+surface materials | | **Purpose** | Global minimum adsorption energy evaluation | | **Paper** | [AdsorbML (arXiv)](https://arxiv.org/abs/2211.16486) | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: ## Overview The OC20Dense dataset is a validation dataset which was used to assess model performance in [AdsorbML: A Leap in Efficiency for Adsorption Energy Calculations using Generalizable Machine Learning Potentials](https://arxiv.org/abs/2211.16486). OC20-Dense contains a dense sampling of adsorbate configurations on ~1,000 randomly selected adsorbate+surface materials from the [OC20](https://arxiv.org/abs/2010.09990) dataset. It comprises a total of 85,658 unique input configurations. This dataset, and the paper written for it, supports the determination of global minimum adsorbate-surface energies (the adsorption energy). This differs from OC20, which contains local adsorbate relaxations. Under low coverage conditions, the global minimum energy site is the most likely to be occupied. For computational catalysis research, we correlate the adsorption energy with important figures of merit, so aquisition of it is an important task. ## File Contents and Download |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |LMDB |654M |9.8G | [0163b0e8c4df6d9c426b875a28d9178a](https://dl.fbaipublicfiles.com/opencatalystproject/data/adsorbml/oc20_dense_data.tar.gz) | |ASE Trajectories |29G |112G | [ee937e5290f8f720c914dc9a56e0281f](https://dl.fbaipublicfiles.com/opencatalystproject/data/adsorbml/oc20_dense_trajectories.tar.gz) | The following files are also provided to be used for evaluation and general information: * `oc20dense_mapping.pkl` : Mapping of the LMDB `sid` to general metadata information. If this file is not present, run the command `python src/fairchem/core/scripts/download_large_files.py adsorbml` from the root of the fairchem repo to download it. - * `system_id`: Unique system identifier for an adsorbate, bulk, surface combination. * `config_id`: Unique configuration identifier, where `rand` and `heur` correspond to random and heuristic initial configurations, respectively. * `mpid`: Materials Project bulk identifier. * `miller_idx`: 3-tuple of integers indicating the Miller indices of the surface. * `shift`: C-direction shift used to determine cutoff for the surface (c-direction is following the nomenclature from Pymatgen). * `top`: Boolean indicating whether the chosen surface was at the top or bottom of the originally enumerated surface. * `adsorbate`: Chemical composition of the adsorbate. * `adsorption_site`: A tuple of 3-tuples containing the Cartesian coordinates of each binding adsorbate atom * `oc20dense_targets.pkl` : DFT adsorption energies across different system and placement ids. * `oc20dense_compute.pkl` : DFT compute as measured in the number of ionic and scf steps for each evaluated relaxation. * `oc20dense_ref_energies.pkl` : Reference energy used for a specified `system_id`. This energy includes the relaxed clean surface and the gas phase adsorbate energy to ensure consistency across calculations. * `oc20dense_tags.pkl` : Tag information used for a specified `system_id`. Where 0 = subsurface, 1 = surface, 2 = adsorbate. All mappings can be obtained at the following downloadable link: https://dl.fbaipublicfiles.com/opencatalystproject/data/adsorbml/oc20_dense_mappings.tar.gz MD5 checksums: ``` c18735c405ce6ce5761432b07287d8d9 oc20_dense_mappings.tar.gz 3e26c3bcef01ccfc9b001931065ea6e6 oc20dense_mapping.pkl fd589b013b72e62e11a6b2a5bd1d323c oc20dense_targets.pkl 78d25997e0aaf754df526ab37276bb89 oc20dense_compute.pkl b07c64158e4bfa5f7b9bf6263753ecc5 oc20dense_ref_energies.pkl 1ba0bc266130f186850f5faa547b6a02 oc20dense_tags.pkl ``` --- Source: `docs/catalysts/datasets/oc20neb.md` # Open Catalyst 2020 Nudged Elastic Band (OC20NEB) :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | 932 NEB relaxation trajectories | | **Reaction Types** | Desorptions, Dissociations, Transfers | | **Purpose** | Transition state energy calculations | | **Paper** | [CatTSunami (arXiv)](https://arxiv.org/abs/2405.02078) | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: ## Overview This is a validation dataset which was used to assess model performance in [CatTSunami: Accelerating Transition State Energy Calculations with Pre-trained Graph Neural Networks](https://arxiv.org/abs/2405.02078). It is comprised of 932 NEB relaxation trajectories. There are three different types of reactions represented: desorptions, dissociations, and transfers. NEB calculations allow us to find transition states. The rate of reaction is determined by the transition state energy, so access to transition states is very important for catalysis research. For more information, check out the paper. ## File Structure and Contents The tar file contains 3 subdirectories: dissociations, desorptions, and transfers. As the names imply, these directories contain the converged DFT trajectories for each of the reaction classes. Within these directories, the trajectories are named to identify the contents of the file. Here is an example and the anatomy of the name: ```desorption_id_83_2409_9_111-4_neb1.0.traj``` 1. `desorption` indicates the reaction type (dissociation and transfer are the other possibilities) 2. `id` identifies that the material belongs to the validation in domain split (ood - out of domain is th e other possibility) 3. `83` is the task id. This does not provide relavent information 4. `2409` is the bulk index of the bulk used in the ocdata bulk pickle file 5. `9` is the reaction index. for each reaction type there is a reaction pickle file in the repository. In this case it is the 9th entry to that pickle file 6. `111-4` the first 3 numbers are the miller indices (i.e. the (1,1,1) surface), and the last number cooresponds to the shift value. In this case the 4th shift enumerated was the one used. 7. `neb1.0` the number here indicates the k value used. For the full dataset, 1.0 was used so this does not distiguish any of the trajectories from one another. The content of these trajectory files is the repeating frame sets. Despite the initial and final frames not being optimized during the NEB, the initial and final frames are saved for every iteration in the trajectory. For the dataset, 10 frames were used - 8 which were optimized over the neb. So the length of the trajectory is the number of iterations (N) * 10. If you wanted to look at the frame set prior to optimization and the optimized frame set, you could get them like this: ```{code-cell} ipython3 from __future__ import annotations from pathlib import Path from urllib.request import urlretrieve trajectory_path = Path("desorption_id_83_2409_9_111-4_neb1.0.traj") if not trajectory_path.exists(): urlretrieve( "https://dl.fbaipublicfiles.com/opencatalystproject/data/large_files/" "desorption_id_83_2409_9_111-4_neb1.0.traj", trajectory_path, ) from ase.io import read traj = read(trajectory_path, ":") unrelaxed_frames = traj[0:10] relaxed_frames = traj[-10:] ``` ## Download |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |ASE Trajectories |1.5G |6.3G | [52af34a93758c82fae951e52af445089](https://dl.fbaipublicfiles.com/opencatalystproject/data/oc20neb/oc20neb_dft_trajectories_04_23_24.tar.gz) | ## Use One more note: We have not prepared an lmdb for this dataset. This is because it is NEB calculations are not supported directly in ocp. You must use the ase native OCP class along with ase infrastructure to run NEB calculations. Here is an example of a use: ```{code-cell} ipython3 import os from ase.io import read from ase.mep import DyNEB from ase.optimize import BFGS from fairchem.core import FAIRChemCalculator, pretrained_mlip traj = read("desorption_id_83_2409_9_111-4_neb1.0.traj", ":") images = traj[0:10] predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") neb = DyNEB(images, k=1) for image in images: image.calc = FAIRChemCalculator(predictor, task_name="oc20") optimizer = BFGS( neb, trajectory="neb.traj", ) # Use a small number of steps here to keep the docs fast during CI, but otherwise do quite reasonable settings. fast_docs = os.environ.get("FAST_DOCS", "false").lower() == "true" if fast_docs: optimization_steps = 20 else: optimization_steps = 300 conv = optimizer.run(fmax=0.45, steps=optimization_steps) if conv: neb.climb = True conv = optimizer.run(fmax=0.05, steps=optimization_steps) ``` --- Source: `docs/catalysts/datasets/ocx24.md` # Open Catalyst Experiments 2024 (OCx24): Bridging Experiments and Computational Models :::{card} Dataset Overview | Property | Value | |----------|-------| | **Type** | Experimental + Computational | | **Reactions** | HER, CO2 electrochemical reduction | | **Data** | XRF, XRD, electrochemical testing | | **Paper** | [arXiv:2411.11783](http://arxiv.org/abs/2411.11783) | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: In this work, we seek to directly bridge the gap between computational descriptors and experimental outcomes for heterogeneous catalysis. We consider two important green chemistries: the hydrogen evolution reaction and the electrochemical reduction of carbon dioxide. To do this, we created a curated dataset of experimental results with materials synthesized and tested in a reproducible manner under industrially relevant conditions. We used this data to build models to directly predict experimental outcomes using computational features. For more information, please read the manuscript [paper](http://arxiv.org/abs/2411.11783). ## Experimental datasets To support this work, we performed X-ray fluorescence (XRF), X-ray diffraction (XRD), and electrochemical testing. Summaries of this data is all publically available [here](https://github.com/facebookresearch/fairchem/tree/main/src/fairchem/applications/ocx/data/experimental_data). ## Computational datasets |Splits |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- | |Computational screening data |1.5G | [9e75b95bb1a2ae691f07cf630eac3378](https://dl.fbaipublicfiles.com/opencatalystproject/data/ocx24/comp_df_241022.csv) | ## Citing this work If you use this codebase in your work, please consider citing: ```bibtex @misc{abed2024opencatalystexperiments2024, title={Open Catalyst Experiments 2024 (OCx24): Bridging Experiments and Computational Models}, author={Jehad Abed and Jiheon Kim and Muhammed Shuaibi and Brook Wander and Boris Duijf and Suhas Mahesh and Hyeonseok Lee and Vahe Gharakhanyan and Sjoerd Hoogland and Erdem Irtem and Janice Lan and Niels Schouten and Anagha Usha Vijayakumar and Jason Hattrick-Simpers and John R. Kitchin and Zachary W. Ulissi and Aaike van Vugt and Edward H. Sargent and David Sinton and C. Lawrence Zitnick}, year={2024}, eprint={2411.11783}, archivePrefix={arXiv}, primaryClass={cond-mat.mtrl-sci}, url={https://arxiv.org/abs/2411.11783}, } ``` --- Source: `docs/catalysts/datasets/oc25.md` # OC25 :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | ~8 million DFT calculations | | **Systems** | 1.5 million explicit solvent environments | | **Avg. System Size** | 144 atoms | | **Elements** | 88 elements | | **DFT Level** | VASP with RPBE+D3 functional | | **License** | [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | ::: The Open Catalyst 2025 (OC25) dataset consists of nearly 8 million DFT calculations across 1.5 million unique explicit solvent environments with system sizes of 144 atoms on average. This dataset represents the largest and most diverse solid-liquid interface dataset that is currently available and provides configurational and elemental diversity: spanning 88 elements, commonly used solvents/ions, varying solvent layers, and off-equilibrium sampling. The dataset enables training of state-of-the-art machine-learned interatomic potentials for applications in electrocatalysis. All structures are labeled with total energies (eV) and forces (eV/Angstrom) computed using VASP with the RPBE+D3 functional. :::{tip} All information about the dataset is available at the [OC25 Huggingface site](https://huggingface.co/facebook/OC25). For questions or issues, please open a GitHub issue in this repository. ::: ## Dataset format The dataset is provided in ASE DB compatible lmdb files (`*.aselmdb`). ### Citing OC25 The OC25 dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any publications that uses this dataset: ```bib @misc{oc25, title={The Open Catalyst 2025 (OC25) Dataset and Models for Solid-Liquid Interfaces}, author={Sushree Jagriti Sahoo and Mikael Maraschin and Daniel S. Levine and Zachary Ulissi and C. Lawrence Zitnick and Joel B Varley and Joseph A. Gauthier and Nitish Govindarajan and Muhammed Shuaibi}, year={2025}, eprint={}, archivePrefix={arXiv}, primaryClass={}, url={}, } ``` --- Source: `docs/catalysts/models.md` # Pretrained Models :::{margin} ```{image} ../assets/icons/catalysis.svg :alt: Catalysis :width: 100px ``` ::: :::{tip} **2025 Recommendation:** We now suggest using the [UMA model](../core/uma), trained on all of the FAIR chemistry datasets before using one of the checkpoints below. ::: :::{card} Why UMA? 1. **State-of-the-art accuracy** in out-of-domain prediction 2. **Total energy predictions** which are helpful for properties beyond adsorption energies and removes ambiguity when catalyst surfaces may reconstruct 3. **Energy conserving and smooth** (UMA small model) - works much better for vibrational calculations, molecular dynamics, etc. 4. **Most likely to be updated** in the future ::: --- ## Legacy FAIRChemV1 Models Trained on Open Catalyst 2020 (OC20) :::{warning} These models are only available if you install `fairchem<=2.0`. For new projects, we recommend using the UMA model instead. ::: This page summarizes all the pretrained models released as part of the [Open Catalyst Project](https://opencatalystproject.org/). * All configurations for these models are available in the [`configs/`](https://github.com/facebookresearch/fairchem/tree/main/configs) directory. * All of these models are trained on various splits of the OC20 S2EF / IS2RE datasets. For details, see [https://arxiv.org/abs/2010.09990](https://arxiv.org/abs/2010.09990) and https://github.com/facebookresearch/fairchem/blob/main/DATASET.md. * All OC20 models are trained on adsorption energies, i.e. the DFT total energies minus the clean surface and gas phase adsorbate energies. For details on how to train models on OC20 total energies, please read the [referencing section here](https://facebookresearch.github.io/fairchem/oc20). ### S2EF models: optimized for EFwT | Model Name | Model | Split |Download |val ID force MAE (eV / Å) |val ID EFwT | |-------------------------------------|---------------------|--------|--- |--- |--- | | CGCNN-S2EF-OC20-200k | CGCNN | 200k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/cgcnn_200k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/200k/cgcnn/cgcnn.yml) |0.08 |0% | | CGCNN-S2EF-OC20-2M | CGCNN | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/cgcnn_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/cgcnn/cgcnn.yml) |0.0673 |0.01% | | CGCNN-S2EF-OC20-20M | CGCNN | 20M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/cgcnn_20M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/20M/cgcnn/cgcnn.yml) |0.065 |0% | | CGCNN-S2EF-OC20-All | CGCNN | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/cgcnn_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/cgcnn/cgcnn.yml) |0.0684 |0.01% | | DimeNet-S2EF-OC20-200k | DimeNet | 200k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/dimenet_200k.pt) |0.0693 |0.01% | | DimeNet-S2EF-OC20-2M | DimeNet | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/dimenet_2M.pt) |0.0576 |0.02% | | SchNet-S2EF-OC20-200k | SchNet | 200k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/schnet_200k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/200k/schnet/schnet.yml) |0.0743 |0% | | SchNet-S2EF-OC20-2M | SchNet | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/schnet_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/schnet/schnet.yml) |0.0737 |0% | | SchNet-S2EF-OC20-20M | SchNet | 20M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/schnet_20M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/20M/schnet/schnet.yml) |0.0568 |0.03% | | SchNet-S2EF-OC20-All | SchNet | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/schnet_all_large.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/schnet/schnet.yml) |0.0494 |0.12% | | DimeNet++-S2EF-OC20-200k | DimeNet++ | 200k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_200k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/200k/dimenet_plus_plus/dpp.yml) |0.0741 |0% | | DimeNet++-S2EF-OC20-2M | DimeNet++ | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/dimenet_plus_plus/dpp.yml) |0.0595 |0.01% | | DimeNet++-S2EF-OC20-20M | DimeNet++ | 20M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_20M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/20M/dimenet_plus_plus/dpp.yml) |0.0511 |0.06% | | DimeNet++-S2EF-OC20-All | DimeNet++ | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/dimenet_plus_plus/dpp.yml) |0.0444 |0.12% | | SpinConv-S2EF-OC20-2M | SpinConv | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_12/s2ef/spinconv_force_centric_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/spinconv/spinconv_force.yml) |0.0329 |0.18% | | SpinConv-S2EF-OC20-All | SpinConv | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_08/s2ef/spinconv_force_centric_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/spinconv/spinconv_force.yml) |0.0267 |1.02% | | GemNet-dT-S2EF-OC20-2M | GemNet-dT | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_12/s2ef/gemnet_t_direct_h512_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/gemnet/gemnet-dT.yml) |0.0257 |1.10% | | GemNet-dT-S2EF-OC20-All | GemNet-dT | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_08/s2ef/gemnet_t_direct_h512_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/gemnet/gemnet-dT.yml) |0.0211 |2.21% | | PaiNN-S2EF-OC20-All | PaiNN | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_05/s2ef/painn_h512_s2ef_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/painn/painn_h512.yml) \| [scale file](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/painn/painn_nb6_scaling_factors.pt) |0.0294 |0.91% | | GemNet-OC-S2EF-OC20-2M | GemNet-OC | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_07/s2ef/gemnet_oc_base_s2ef_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/gemnet/gemnet-oc.yml) \| [scale file](https://github.com/facebookresearch/fairchem/blob/481f3a5a92dc787384ddae9fe3f50f5d932712fd/configs/oc20/s2ef/all/gemnet/scaling_factors/gemnet-oc.pt) |0.0225 |2.12% | | GemNet-OC-S2EF-OC20-All | GemNet-OC | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_07/s2ef/gemnet_oc_base_s2ef_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/gemnet/gemnet-oc.yml) \| [scale file](https://github.com/facebookresearch/fairchem/blob/481f3a5a92dc787384ddae9fe3f50f5d932712fd/configs/oc20/s2ef/all/gemnet/scaling_factors/gemnet-oc.pt) |0.0179 |4.56% | | GemNet-OC-S2EF-OC20-All+MD | GemNet-OC | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/gemnet_oc_base_s2ef_all_md.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/gemnet/gemnet-oc.yml) \| [scale file](https://github.com/facebookresearch/fairchem/blob/481f3a5a92dc787384ddae9fe3f50f5d932712fd/configs/oc20/s2ef/all/gemnet/scaling_factors/gemnet-oc.pt) |0.0173 |4.72% | | GemNet-OC-Large-S2EF-OC20-All+MD | GemNet-OC-Large | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_07/s2ef/gemnet_oc_large_s2ef_all_md.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/gemnet/gemnet-oc-large.yml) \| [scale file](https://github.com/facebookresearch/fairchem/blob/481f3a5a92dc787384ddae9fe3f50f5d932712fd/configs/oc20/s2ef/all/gemnet/scaling_factors/gemnet-oc-large.pt) |0.0164 |5.34% | | SCN-S2EF-OC20-2M | SCN | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/scn_t1_b1_s2ef_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/scn/scn-t1-b1.yml) |0.0216 |1.68% | | SCN-t4-b2-S2EF-OC20-2M | SCN-t4-b2 | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/scn_t4_b2_s2ef_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/scn/scn-t4-b2.yml) |0.0193 |2.68% | | SCN-S2EF-OC20-All+MD | SCN | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/scn_all_md_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/scn/scn-all-md.yml) |0.0160 |5.08% | | eSCN-L4-M2-Lay12-S2EF-OC20-2M | eSCN-L4-M2-Lay12 | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/escn_l4_m2_lay12_2M_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/escn/eSCN-L4-M2-Lay12.yml) |0.0191 |2.55% | | eSCN-L6-M2-Lay12-S2EF-OC20-2M | eSCN-L6-M2-Lay12 | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/escn_l6_m2_lay12_2M_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/escn/eSCN-L6-M2-Lay12.yml) \| [exported](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/escn_l6_m2_lay12_2M_s2ef_export_cuda_9182024.pt2) |0.0186 |2.66% | | eSCN-L6-M2-Lay12-S2EF-OC20-All+MD | eSCN-L6-M2-Lay12 | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/escn_l6_m2_lay12_all_md_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/escn/eSCN-L6-M2-Lay12-All-MD.yml) |0.0161 |4.28% | | eSCN-L6-M3-Lay20-S2EF-OC20-All+MD | eSCN-L6-M3-Lay20 | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_03/s2ef/escn_l6_m3_lay20_all_md_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/escn/eSCN-L6-M3-Lay20-All-MD.yml) |0.0139 |6.64% | | EquiformerV2-83M-S2EF-OC20-2M | EquiformerV2 (83M) | 2M |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_06/oc20/s2ef/eq2_83M_2M.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/2M/equiformer_v2/equiformer_v2_N@12_L@6_M@2.yml) |0.0167 |4.26% | | EquiformerV2-31M-S2EF-OC20-All+MD | EquiformerV2 (31M) | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_06/oc20/s2ef/eq2_31M_ec4_allmd.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/equiformer_v2/equiformer_v2_N@8_L@4_M@2_31M.yml) |0.0142 |6.20% | | EquiformerV2-153M-S2EF-OC20-All+MD | EquiformerV2 (153M) | All+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_06/oc20/s2ef/eq2_153M_ec4_allmd.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/equiformer_v2/equiformer_v2_N@20_L@6_M@3_153M.yml) |0.0126 |8.90% | ### S2EF models: optimized for force only | Model Name |Model |Split |Download |val ID force MAE | |--------------------------------------|--- |--- |--- |--- | | SchNet-S2EF-force-only-OC20-All |SchNet |All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/schnet_all_forceonly.pt) |0.0443 | | DimeNet++-force-only-OC20-All | DimeNet++ |All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/s2ef/dimenetpp_all_forceonly.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/dimenet_plus_plus/dpp_forceonly.yml) |0.0334 | | DimeNet++-Large-S2EF-force-only-OC20-All | DimeNet++-Large |All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_large_all_forceonly.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/dimenet_plus_plus/dpp10.7M_forceonly.yml) |0.02825 | | DimeNet++-S2EF-force-only-OC20-20M+Rattled | DimeNet++ |20M+Rattled |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_20M_rattled_forceonly.pt) |0.0614 | | DimeNet++-S2EF-force-only-OC20-20M+MD | DimeNet++ |20M+MD |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/s2ef/dimenetpp_20M_md_forceonly.pt) |0.0594 | ### IS2RE models | Model Name | Model | Split |Download |val ID energy MAE | |---------------------------|------------|--------|--- |--- | | CGCNN-IS2RE-OC20-10k | CGCNN | 10k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/cgcnn_10k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/10k/cgcnn/cgcnn.yml) |0.9881 | | CGCNN-IS2RE-OC20-100k | CGCNN | 100k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/cgcnn_100k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/100k/cgcnn/cgcnn.yml) |0.682 | | CGCNN-IS2RE-OC20-All | CGCNN | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/cgcnn_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/all/cgcnn/cgcnn.yml) |0.6199 | | DimeNet-IS2RE-OC20-10k | DimeNet | 10k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/is2re/dimenet_10k.pt) |1.0117 | | DimeNet-IS2RE-OC20-100k | DimeNet | 100k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/is2re/dimenet_100k.pt) |0.6658 | | DimeNet-IS2RE-OC20-all | DimeNet | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2020_11/is2re/dimenet_all.pt) |0.5999 | | SchNet-IS2RE-OC20-10k | SchNet | 10k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/schnet_10k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/10k/schnet/schnet.yml) |1.059 | | SchNet-IS2RE-OC20-100k | SchNet | 100k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/schnet_100k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/100k/schnet/schnet.yml) |0.7137 | | SchNet-IS2RE-OC20-All | SchNet | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/schnet_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/all/schnet/schnet.yml) |0.6458 | | DimeNet++-IS2RE-OC20-10k | DimeNet++ | 10k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/dimenetpp_10k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/10k/dimenet_plus_plus/dpp.yml) |0.8837 | | DimeNet++-IS2RE-OC20-100k | DimeNet++ | 100k |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/dimenetpp_100k.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/100k/dimenet_plus_plus/dpp.yml) |0.6388 | | DimeNet++-IS2RE-OC20-All | DimeNet++ | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2021_02/is2re/dimenetpp_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/all/dimenet_plus_plus/dpp.yml) |0.5639 | | PaiNN-IS2RE-OC20-All | PaiNN | All |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_05/is2re/painn_h1024_bs4x8_is2re_all.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/is2re/all/painn/painn_h1024_bs8x4.yml) \| [scale file](https://github.com/facebookresearch/fairchem/blob/main/configs/oc20/s2ef/all/painn/painn_nb6_scaling_factors.pt) |0.5728 | The Open Catalyst 2020 (OC20) dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any research manuscript using the OC20 dataset or pretrained models, as well as the original paper for each model: ```bibtex @article{ocp_dataset, author = {Chanussot*, Lowik and Das*, Abhishek and Goyal*, Siddharth and Lavril*, Thibaut and Shuaibi*, Muhammed and Riviere, Morgane and Tran, Kevin and Heras-Domingo, Javier and Ho, Caleb and Hu, Weihua and Palizhati, Aini and Sriram, Anuroop and Wood, Brandon and Yoon, Junwoong and Parikh, Devi and Zitnick, C. Lawrence and Ulissi, Zachary}, title = {Open Catalyst 2020 (OC20) Dataset and Community Challenges}, journal = {ACS Catalysis}, year = {2021}, doi = {10.1021/acscatal.0c04525}, } ``` ## Open Catalyst 2022 (OC22) * All configurations for these models are available in the [`configs/oc22`](https://github.com/facebookresearch/fairchem/tree/main/configs/oc22) directory. * All of these models are trained on various splits of the OC22 S2EF / IS2RE datasets. For details, see [https://arxiv.org/abs/2206.08917](https://arxiv.org/abs/2206.08917) and https://github.com/facebookresearch/fairchem/blob/main/DATASET.md. * All OC22 models released here are trained on DFT total energies, in contrast to the OC20 models listed above, which are trained on adsorption energies. ### S2EF-Total models | Model Name |Model |Training |Download |val ID force MAE |val ID energy MAE | |-----------------------------------|--- |--- |--- |--- |--- | | GemNet-dT-S2EFS-OC22 |GemNet-dT | OC22 |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_09/oc22/s2ef/gndt_oc22_all_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc22/s2ef/gemnet-dt/gemnet-dT.yml) |0.032 |1.127 | | GemNet-OC-S2EFS-OC22 | GemNet-OC | OC22 |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_09/oc22/s2ef/gnoc_oc22_all_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc22/s2ef/gemnet-oc/gemnet_oc.yml) |0.030 |0.563 | | GemNet-OC-S2EFS-OC20+OC22 | GemNet-OC | OC20+OC22 |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_09/oc22/s2ef/gnoc_oc22_oc20_all_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc22/s2ef/gemnet-oc/gemnet_oc_oc20_oc22.yml) |0.027 |0.483 | | GemNet-OC-S2EFS-nsn-OC20+OC22 | GemNet-OC
(trained with `enforce_max_neighbors_strictly=False`, [#467](https://github.com/facebookresearch/fairchem/pull/467)) | OC20+OC22 |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_05/oc22/s2ef/gnoc_oc22_oc20_all_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc22/s2ef/gemnet-oc/gemnet_oc_oc20_oc22_degen_edges.yml) |0.027 |0.458 | | GemNet-OC-S2EFS-OC20->OC22 | GemNet-OC | OC20->OC22 |[checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2022_09/oc22/s2ef/gnoc_finetune_all_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc22/s2ef/gemnet-oc/gemnet_oc_finetune.yml) |0.030 |0.417 | | EquiformerV2-lE4-lF100-S2EFS-OC22 | EquiformerV2 ($\lambda_E$=4, $\lambda_F$=100) | OC22 | [checkpoint](https://dl.fbaipublicfiles.com/opencatalystproject/models/2023_10/oc22/s2ef/eq2_121M_e4_f100_oc22_s2ef.pt) \| [config](https://github.com/facebookresearch/fairchem/blob/main/configs/oc22/s2ef/equiformer_v2/equiformer_v2_N@18_L@6_M@2_e4_f100_121M.yml) | 0.023 | 0.447 The Open Catalyst 2022 (OC22) dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any research manuscript using the OC22 dataset or pretrained models, as well as the original paper for each model: ```bibtex @article{oc22_dataset, author = {Tran*, Richard and Lan*, Janice and Shuaibi*, Muhammed and Wood*, Brandon and Goyal*, Siddharth and Das, Abhishek and Heras-Domingo, Javier and Kolluru, Adeesh and Rizvi, Ammar and Shoghi, Nima and Sriram, Anuroop and Ulissi, Zachary and Zitnick, C. Lawrence}, title = {The Open Catalyst 2022 (OC22) dataset and challenges for oxide electrocatalysts}, journal = {ACS Catalysis}, year={2023}, } ``` --- Source: `docs/catalysts/examples_tutorials/summary.md` # Examples & Tutorials :::{margin} ```{image} ../../assets/icons/catalysis.svg :alt: Catalysis :width: 100px ``` ::: This section provides hands-on tutorials for using FAIRChem models in catalysis research. From basic adsorption energy calculations to advanced transition state searches, these tutorials will help you get started with computational catalysis. :::{tip} New to FAIRChem? Start with the [Intro to Adsorption Energies](OCP-introduction.md) tutorial to learn the basics. ::: ::::{grid} 1 2 2 2 :::{grid-item-card} Intro to Adsorption Energies :link: OCP-introduction Learn the fundamentals of calculating adsorption energies using UMA models and compare trends across different metals. +++ **Difficulty:** Beginner ::: :::{grid-item-card} Expert Adsorption Energies :link: adsorption_energies/adsorption_energies Advanced tutorial on reproducing literature results for NRR/HER selectivity using automated adsorbate placement. +++ **Difficulty:** Advanced ::: :::{grid-item-card} AdsorbML Walkthrough :link: adsorbml_walkthrough Use the AdsorbML workflow to automatically find optimal adsorption sites with ML-accelerated relaxations. +++ **Difficulty:** Intermediate ::: :::{grid-item-card} Transition State Search (NEBs) :link: cattsunami_tutorial Find reaction transition states using CatTsunami tools for NEB calculations on catalyst surfaces. +++ **Difficulty:** Advanced ::: :::{grid-item-card} OCP API :link: ocpapi Programmatically access the Open Catalyst Demo for adsorbate binding site searches. +++ **Difficulty:** Intermediate ::: :::: --- Source: `docs/catalysts/examples_tutorials/OCP-introduction.md` # Intro to Adsorption Energies :::{card} Tutorial Overview | Property | Value | |----------|-------| | **Difficulty** | Beginner | | **Time** | 15-30 minutes | | **Prerequisites** | Basic Python, familiarity with ASE | | **Goal** | Calculate adsorption energies using UMA models | ::: To introduce OCP we start with using it to calculate adsorption energies for a simple, atomic adsorbate where we specify the site we want to the adsorption energy for. Conceptually, you do this like you would do it with density functional theory. You create a slab model for the surface, place an adsorbate on it as an initial guess, run a relaxation to get the lowest energy geometry, and then compute the adsorption energy using reference states for the adsorbate. :::{important} Some OCP model/checkpoint combinations return a total energy like density functional theory would, but some return an "adsorption energy" directly. You have to know which one you are using. In this example, the model we use returns an "adsorption energy". ::: +++ (intro-to-adsorption-energies)= ## Intro to Adsorption energies Adsorption energies are always a reaction energy (an adsorbed species relative to some implied combination of reactants). There are many common schemes in the catalysis literature. For example, you may want the adsorption energy of oxygen, and you might compute that from this reaction: 1/2 O2 + slab -> slab-O DFT has known errors with the energy of a gas-phase O2 molecule, so it's more common to compute this energy relative to a linear combination of H2O and H2. The suggested reference scheme for consistency with OC20 is a reaction x CO + (x + y/2 - z) H2 + (z-x) H2O + w/2 N2 + * -> CxHyOzNw* Here, x=y=w=0, z=1, so the reaction ends up as -H2 + H2O + * -> O* or alternatively, H2O + * -> O* + H2 It is possible through thermodynamic cycles to compute other reactions. If we can look up rH1 below and compute rH2 H2 + 1/2 O2 -> H2O re1 = -3.03 eV, from exp H2O + * -> O* + H2 re2 # Get from UMA Then, the adsorption energy for 1/2O2 + * -> O* is just re1 + re2. Based on https://atct.anl.gov/Thermochemical%20Data/version%201.118/species/?species_number=986, the formation energy of water is about -3.03 eV at standard state experimentally. You could also compute this using DFT, but you would probably get the wrong answer for this. The first step is getting a checkpoint for the model we want to use. UMA is currently the state-of-the-art model and will provide total energy estimates at the RPBE level of theory if you use the "OC20" task. ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` If you find your kernel is crashing, it probably means you have exceeded the allowed amount of memory. This checkpoint works fine in this example, but it may crash your kernel if you use it in the NRR example. This next cell will automatically download the checkpoint from huggingface and load it. ```{code-cell} from __future__ import annotations from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") ``` Next we can build a slab with an adsorbate on it. Here we use the ASE module to build a Pt slab. We use the experimental lattice constant that is the default. This can introduce some small errors with DFT since the lattice constant can differ by a few percent, and it is common to use DFT lattice constants. In this example, we do not constrain any layers. ```{code-cell} from ase.build import add_adsorbate, fcc111 from ase.optimize import BFGS ``` ```{code-cell} # reference energies from a linear combination of H2O/N2/CO/H2! atomic_reference_energies = { "H": -3.477, "N": -8.083, "O": -7.204, "C": -7.282, } re1 = -3.03 slab = fcc111("Pt", size=(2, 2, 5), vacuum=20.0) slab.pbc = True adslab = slab.copy() add_adsorbate(adslab, "O", height=1.2, position="fcc") slab.set_calculator(calc) opt = BFGS(slab) opt.run(fmax=0.05, steps=100) slab_e = slab.get_potential_energy() adslab.set_calculator(calc) opt = BFGS(adslab) opt.run(fmax=0.05, steps=100) adslab_e = adslab.get_potential_energy() # Energy for ((H2O-H2) + * -> *O) + (H2 + 1/2O2 -> H2) leads to 1/2O2 + * -> *O! adslab_e - slab_e - atomic_reference_energies["O"] + re1 ``` It is good practice to look at your geometries to make sure they are what you expect. ```{code-cell} import matplotlib.pyplot as plt from ase.visualize.plot import plot_atoms fig, axs = plt.subplots(1, 2) plot_atoms(slab, axs[0]) plot_atoms(slab, axs[1], rotation=("-90x")) axs[0].set_axis_off() axs[1].set_axis_off() ``` ```{code-cell} import matplotlib.pyplot as plt from ase.visualize.plot import plot_atoms fig, axs = plt.subplots(1, 2) plot_atoms(adslab, axs[0]) plot_atoms(adslab, axs[1], rotation=("-90x")) axs[0].set_axis_off() axs[1].set_axis_off() ``` How did we do? We need a reference point. In the paper below, there is an atomic adsorption energy for O on Pt(111) of about -4.264 eV. This is for the reaction O + * -> O*. To convert this to the dissociative adsorption energy, we have to add the reaction: 1/2 O2 -> O D = 2.58 eV (expt) to get a comparable energy of about -1.68 eV. There is about ~0.2 eV difference (we predicted -1.47 eV above, and the reference comparison is -1.68 eV) to account for. The biggest difference is likely due to the differences in exchange-correlation functional. The reference data used the PBE functional, and eSCN was trained on *RPBE* data. To additional places where there are differences include: 1. Difference in lattice constant 2. The reference energy used for the experiment references. These can differ by up to 0.5 eV from comparable DFT calculations. 3. How many layers are relaxed in the calculation Some of these differences tend to be systematic, and you can calibrate and correct these, especially if you can augment these with your own DFT calculations. See [convergence study](#convergence-study) for some additional studies of factors that influence this number. +++ ### Exercises 1. Explore the effect of the lattice constant on the adsorption energy. 2. Try different sites, including the bridge and top sites. Compare the energies, and inspect the resulting geometries. +++ ## Trends in adsorption energies across metals. Xu, Z., & Kitchin, J. R. (2014). Probing the coverage dependence of site and adsorbate configurational correlations on (111) surfaces of late transition metals. J. Phys. Chem. C, 118(44), 25597–25602. http://dx.doi.org/10.1021/jp508805h [Supporting information](https://pubs.acs.org/doi/suppl/10.1021/jp508805h/suppl_file/jp508805h_si_001.pdf). These are atomic adsorption energies: O + * -> O* We have to do some work to get comparable numbers from OCP H2 + 1/2 O2 -> H2O re1 = -3.03 eV H2O + * -> O* + H2 re2 # Get from UMA O -> 1/2 O2 re3 = -2.58 eV Then, the adsorption energy for O + * -> O* is just re1 + re2 + re3. Here we just look at the fcc site on Pt. First, we get the data stored in the paper. Next we get the structures and compute their energies. Some subtle points are that we have to account for stoichiometry, and normalize the adsorption energy by the number of oxygens. +++ First we get a reference energy from the paper (PBE, 0.25 ML O on Pt(111)). ```{code-cell} import json with open("energies.json") as f: edata = json.load(f) with open("structures.json") as f: sdata = json.load(f) edata["Pt"]["O"]["fcc"]["0.25"] ``` Next, we load data from the SI to get the geometry to start from. ```{code-cell} with open("structures.json") as f: s = json.load(f) sfcc = s["Pt"]["O"]["fcc"]["0.25"] ``` Next, we construct the atomic geometry, run the geometry optimization, and compute the energy. ```{code-cell} re3 = -2.58 # O -> 1/2 O2 re3 = -2.58 eV from ase import Atoms adslab = Atoms(sfcc["symbols"], positions=sfcc["pos"], cell=sfcc["cell"], pbc=True) # Grab just the metal surface atoms slab = adslab[adslab.arrays["numbers"] == adslab.arrays["numbers"][0]] adsorbates = adslab[~(adslab.arrays["numbers"] == adslab.arrays["numbers"][0])] slab.set_calculator(calc) opt = BFGS(slab) opt.run(fmax=0.05, steps=100) adslab.set_calculator(calc) opt = BFGS(adslab) opt.run(fmax=0.05, steps=100) re2 = ( adslab.get_potential_energy() - slab.get_potential_energy() - sum([atomic_reference_energies[x] for x in adsorbates.get_chemical_symbols()]) ) nO = 0 for atom in adslab: if atom.symbol == "O": nO += 1 re2 += re1 + re3 print(re2 / nO) ``` ### Site correlations This cell reproduces a portion of a figure in the paper. We compare oxygen adsorption energies in the fcc and hcp sites across metals and coverages. These adsorption energies are highly correlated with each other because the adsorption sites are so similar. At higher coverages, the agreement is not as good. This is likely because the model is extrapolating and needs to be fine-tuned. ```{code-cell} import time from tqdm import tqdm t0 = time.time() data = {"fcc": [], "hcp": []} refdata = {"fcc": [], "hcp": []} for metal in ["Cu", "Ag", "Pd", "Pt", "Rh", "Ir"]: print(metal) for site in ["fcc", "hcp"]: for adsorbate in ["O"]: for coverage in tqdm(["0.25"]): entry = s[metal][adsorbate][site][coverage] adslab = Atoms( entry["symbols"], positions=entry["pos"], cell=entry["cell"], pbc=True, ) # Grab just the metal surface atoms adsorbates = adslab[ ~(adslab.arrays["numbers"] == adslab.arrays["numbers"][0]) ] slab = adslab[adslab.arrays["numbers"] == adslab.arrays["numbers"][0]] slab.set_calculator(calc) opt = BFGS(slab) opt.run(fmax=0.05, steps=100) adslab.set_calculator(calc) opt = BFGS(adslab) opt.run(fmax=0.05, steps=100) re2 = ( adslab.get_potential_energy() - slab.get_potential_energy() - sum( [ atomic_reference_energies[x] for x in adsorbates.get_chemical_symbols() ] ) ) nO = 0 for atom in adslab: if atom.symbol == "O": nO += 1 re2 += re1 + re3 data[site] += [re2 / nO] refdata[site] += [edata[metal][adsorbate][site][coverage]] f"Elapsed time = {time.time() - t0} seconds" ``` First, we compare the computed data and reference data. There is a systematic difference of about 0.5 eV due to the difference between RPBE and PBE functionals, and other subtle differences like lattice constant differences and reference energy differences. This is pretty typical, and an expected deviation. ```{code-cell} plt.plot(refdata["fcc"], data["fcc"], "r.", label="fcc") plt.plot(refdata["hcp"], data["hcp"], "b.", label="hcp") plt.plot([-5.5, -3.5], [-5.5, -3.5], "k-") plt.xlabel("Ref. data (DFT)") plt.ylabel("UMA-OC20 prediction"); ``` Next we compare the correlation between the hcp and fcc sites. Here we see the same trends. The data falls below the parity line because the hcp sites tend to be a little weaker binding than the fcc sites. ```{code-cell} plt.plot(refdata["hcp"], refdata["fcc"], "r.") plt.plot(data["hcp"], data["fcc"], ".") plt.plot([-6, -1], [-6, -1], "k-") plt.xlabel("$H_{ads, hcp}$") plt.ylabel("$H_{ads, fcc}$") plt.legend(["DFT (PBE)", "UMA-OC20"]); ``` ### Exercises 1. You can also explore a few other adsorbates: C, H, N. 2. Explore the higher coverages. The deviations from the reference data are expected to be higher, but relative differences tend to be better. You probably need fine tuning to improve this performance. This data set doesn't have forces though, so it isn't practical to do it here. +++ ## Next steps In the next step, we consider some more complex adsorbates in nitrogen reduction, and how we can leverage OCP to automate the search for the most stable adsorbate geometry. See [the next step](./NRR/NRR_example-gemnet). +++ (convergence-study)= ### Convergence study In [the adsorption energies section](#intro-to-adsorption-energies) we discussed some possible reasons we might see a discrepancy. Here we investigate some factors that impact the computed energies. In this section, the energies refer to the reaction 1/2 O2 -> O*. +++ ### Effects of number of layers Slab thickness could be a factor. Here we relax the whole slab, and see by about 4 layers the energy is converged to ~0.02 eV. ```{code-cell} for nlayers in [3, 4, 5, 6, 7, 8]: slab = fcc111("Pt", size=(2, 2, nlayers), vacuum=10.0) slab.pbc = True slab.set_calculator(calc) opt_slab = BFGS(slab, logfile=None) opt_slab.run(fmax=0.05, steps=100) slab_e = slab.get_potential_energy() adslab = slab.copy() add_adsorbate(adslab, "O", height=1.2, position="fcc") adslab.pbc = True adslab.set_calculator(calc) opt_adslab = BFGS(adslab, logfile=None) opt_adslab.run(fmax=0.05, steps=100) adslab_e = adslab.get_potential_energy() print( f"nlayers = {nlayers}: {adslab_e - slab_e - atomic_reference_energies['O'] + re1:1.2f} eV" ) ``` ### Effects of relaxation It is common to only relax a few layers, and constrain lower layers to bulk coordinates. We do that here. We only relax the adsorbate and the top layer. This has a small effect (0.1 eV). ```{code-cell} from ase.constraints import FixAtoms for nlayers in [3, 4, 5, 6, 7, 8]: slab = fcc111("Pt", size=(2, 2, nlayers), vacuum=10.0) slab.set_constraint(FixAtoms(mask=[atom.tag > 1 for atom in slab])) slab.pbc = True slab.set_calculator(calc) opt_slab = BFGS(slab, logfile=None) opt_slab.run(fmax=0.05, steps=100) slab_e = slab.get_potential_energy() adslab = slab.copy() add_adsorbate(adslab, "O", height=1.2, position="fcc") adslab.set_constraint(FixAtoms(mask=[atom.tag > 1 for atom in adslab])) adslab.pbc = True adslab.set_calculator(calc) opt_adslab = BFGS(adslab, logfile=None) opt_adslab.run(fmax=0.05, steps=100) adslab_e = adslab.get_potential_energy() print( f"nlayers = {nlayers}: {adslab_e - slab_e - atomic_reference_energies['O'] + re1:1.2f} eV" ) ``` ### Unit cell size Coverage effects are quite noticeable with oxygen. Here we consider larger unit cells. This effect is large, and the results don't look right, usually adsorption energies get more favorable at lower coverage, not less. This suggests fine-tuning could be important even at low coverages. ```{code-cell} for size in [1, 2, 3, 4, 5]: slab = fcc111("Pt", size=(size, size, 5), vacuum=10.0) slab.set_constraint(FixAtoms(mask=[atom.tag > 1 for atom in slab])) slab.pbc = True slab.set_calculator(calc) opt_slab = BFGS(slab, logfile=None) opt_slab.run(fmax=0.05, steps=100) slab_e = slab.get_potential_energy() adslab = slab.copy() add_adsorbate(adslab, "O", height=1.2, position="fcc") adslab.set_constraint(FixAtoms(mask=[atom.tag > 1 for atom in adslab])) adslab.pbc = True adslab.set_calculator(calc) opt_adslab = BFGS(adslab, logfile=None) opt_adslab.run(fmax=0.05, steps=100) adslab_e = adslab.get_potential_energy() print( f"({size}x{size}): {adslab_e - slab_e - atomic_reference_energies['O'] + re1:1.2f} eV" ) ``` ## Summary As with DFT, you should take care to see how these kinds of decisions affect your results, and determine if they would change any interpretations or not. --- Source: `docs/catalysts/examples_tutorials/adsorption_energies/adsorption_energies.md` # Expert Adsorption Energies :::{card} Tutorial Overview | Property | Value | |----------|-------| | **Difficulty** | Advanced | | **Time** | 45-60 minutes | | **Prerequisites** | Basic catalysis knowledge, Python, ASE | | **Goal** | Reproduce NRR/HER selectivity literature results | ::: One of the most common tasks in computational catalysis is calculating the binding energies or adsorption energies of small molecules on catalyst surfaces. ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ```{code-cell} ipython3 from __future__ import annotations import os import pickle import time from glob import glob import ase.io import matplotlib.pyplot as plt import numpy as np import pandas as pd from ase.optimize import QuasiNewton from fairchem.core import FAIRChemCalculator, pretrained_mlip from fairchem.data.oc.core import Adsorbate, AdsorbateSlabConfig, Bulk, Slab from fairchem.data.oc.utils import DetectTrajAnomaly from scipy.stats import linregress # Set random seed to ensure adsorbate enumeration yields a valid candidate # If using a larger number of random samples this wouldn't be necessary np.random.seed(22) ``` # Introduction We will reproduce Fig 6b from the following paper: Zhou, Jing, et al. "Enhanced Catalytic Activity of Bimetallic Ordered Catalysts for Nitrogen Reduction Reaction by Perturbation of Scaling Relations." ACS Catalysis 134 (2023): 2190-2201 (https://doi.org/10.1021/acscatal.2c05877). The gist of this figure is a correlation between H* and NNH* adsorbates across many different alloy surfaces. Then, they identify a dividing line between these that separates surfaces known for HER and those known for NRR. To do this, we will enumerate adsorbate-slab configurations and run ML relaxations on them to find the lowest energy configuration. We will assess parity between the model predicted values and those reported in the paper. Finally we will make the figure and assess separability of the NRR favored and HER favored domains. +++ # Enumerate the adsorbate-slab configurations to run relaxations on +++ Be sure to set the path in `fairchem/data/oc/configs/paths.py` to point to the correct place or pass the paths as an argument. The database pickles can be found in `fairchem/data/oc/databases/pkls` (some pkl files are only downloaded by running the command `python src/fairchem/core/scripts/download_large_files.py oc` from the root of the fairchem repo). We will show one explicitly here as an example and then run all of them in an automated fashion for brevity. ```{code-cell} ipython3 from pathlib import Path import fairchem.data.oc db = Path(fairchem.data.oc.__file__).parent / Path("databases/pkls/adsorbates.pkl") db ``` ## Work out a single example We load one bulk id, create a bulk reference structure from it, then generate the surfaces we want to compute. ```{code-cell} ipython3 bulk_src_id = "oqmd-343039" adsorbate_smiles_nnh = "*N*NH" adsorbate_smiles_h = "*H" bulk = Bulk(bulk_src_id_from_db=bulk_src_id, bulk_db_path="NRR_example_bulks.pkl") adsorbate_H = Adsorbate( adsorbate_smiles_from_db=adsorbate_smiles_h, adsorbate_db_path=db ) adsorbate_NNH = Adsorbate( adsorbate_smiles_from_db=adsorbate_smiles_nnh, adsorbate_db_path=db ) slab = Slab.from_bulk_get_specific_millers(bulk=bulk, specific_millers=(1, 1, 1)) slab ``` We now need to generate potential placements. We use two kinds of guesses, a heuristic and a random approach. This cell generates 13 potential adsorption geometries. ```{code-cell} ipython3 # Perform heuristic placements heuristic_adslabs = AdsorbateSlabConfig(slab[0], adsorbate_H, mode="heuristic") # Perform random placements # (for AdsorbML we use `num_sites = 100` but we will use 4 for brevity here) random_adslabs = AdsorbateSlabConfig( slab[0], adsorbate_H, mode="random_site_heuristic_placement", num_sites=4 ) adslabs = [*heuristic_adslabs.atoms_list, *random_adslabs.atoms_list] len(adslabs) ``` Let's see what we are looking at. It is a little tricky to see the tiny H atom in these figures, but with some inspection you can see there are ontop, bridge, and hollow sites in different places. This is not an exhaustive search; you can increase the number of random placements to check more possibilities. The main idea here is to *increase* the probability you find the most relevant sites. ```{code-cell} ipython3 from ase.visualize.plot import plot_atoms fig, axs = plt.subplots(4, 4) for i, slab in enumerate(adslabs): plot_atoms(slab, axs[i % 4, i // 4]) axs[i % 4, i // 4].set_axis_off() for i in range(16): axs[i % 4, i // 4].set_axis_off() plt.tight_layout() ``` ### Run an ML relaxation We will use an ASE compatible calculator to run these. +++ Running the model with QuasiNewton prints at each relaxation step which is a lot to print. So we will just run one to demonstrate what happens on each iteration. ```{code-cell} ipython3 os.makedirs(f"data/{bulk_src_id}_{adsorbate_smiles_h}", exist_ok=True) # Define the predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") ``` Now we setup and run the relaxation. ```{code-cell} ipython3 t0 = time.time() os.makedirs(f"data/{bulk_src_id}_H", exist_ok=True) adslab = adslabs[0] adslab.calc = calc adslab.pbc = True opt = QuasiNewton(adslab, trajectory=f"data/{bulk_src_id}_H/test.traj") opt.run(fmax=0.05, steps=100) print(f"Elapsed time {time.time() - t0:1.1f} seconds") ``` With a GPU this runs pretty quickly. It is much slower on a CPU. +++ # Run all the systems In principle you can run all the systems now. It takes about an hour though, and we leave that for a later exercise if you want. For now we will run the first two, and for later analysis we provide a results file of all the runs. Let's read in our reference file and take a look at what is in it. ```{code-cell} ipython3 with open("NRR_example_bulks.pkl", "rb") as f: bulks = pickle.load(f) bulks ``` We have 19 bulk materials we will consider. Next we extract the `src-id` for each one. ```{code-cell} ipython3 bulk_ids = [row["src_id"] for row in bulks] ``` In theory you would run all of these, but it takes about an hour with a GPU. We provide the relaxation logs and trajectories in the repo for the next step. These steps are embarrassingly parallel, and can be launched that way to speed things up. The only thing you need to watch is that you don't exceed the available RAM, which will cause the Jupyter kernel to crash. +++ The goal here is to relax each candidate adsorption geometry and save the results in a trajectory file we will analyze later. Each trajectory file will have the geometry and final energy of the relaxed structure. It is somewhat time consuming to run this. We're going to use a small number of bulks for the testing of this documentation, but otherwise run all of the results for the actual documentation. ```{code-cell} ipython3 import os fast_docs = os.environ.get("FAST_DOCS", "false").lower() == "true" if fast_docs: num_bulks = 1 num_sites = 5 relaxation_steps = 20 else: num_bulks = -1 num_sites = 20 relaxation_steps = 300 ``` ```{code-cell} ipython3 import random import time from tqdm import tqdm tinit = time.time() random.seed(42) random.shuffle(bulk_ids) # Note we're just doing the first bulk_id! for bulk_src_id in tqdm(bulk_ids[:num_bulks]): # Set up data directories os.makedirs("data/slabs/", exist_ok=True) os.makedirs(f"data/adslabs/{bulk_src_id}_H", exist_ok=True) os.makedirs(f"data/adslabs/{bulk_src_id}_NNH", exist_ok=True) # Enumerate slabs and establish adsorbates bulk = Bulk(bulk_src_id_from_db=bulk_src_id, bulk_db_path="NRR_example_bulks.pkl") slab = Slab.from_bulk_get_specific_millers(bulk=bulk, specific_millers=(1, 1, 1)) slab_atoms = slab[0].atoms.copy() slab_atoms.calc = calc slab_atoms.pbc = True opt = QuasiNewton( slab_atoms, trajectory=f"data/slabs/{bulk_src_id}.traj", logfile=f"data/slabs/{bulk_src_id}.log", ) opt.run(fmax=0.05, steps=relaxation_steps) print( f" Elapsed time: {time.time() - t0:1.1f} seconds for data/slabs/{bulk_src_id} slab relaxation" ) # Perform heuristic placements heuristic_adslabs_H = AdsorbateSlabConfig( slab[0], adsorbate_H, mode="random_site_heuristic_placement", num_sites=num_sites, ) heuristic_adslabs_NNH = AdsorbateSlabConfig( slab[0], adsorbate_NNH, mode="random_site_heuristic_placement", num_sites=num_sites, ) print(f"{len(heuristic_adslabs_H.atoms_list)} H slabs to compute for {bulk_src_id}") print( f"{len(heuristic_adslabs_NNH.atoms_list)} NNH slabs to compute for {bulk_src_id}" ) for idx, adslab in enumerate(heuristic_adslabs_H.atoms_list): t0 = time.time() adslab.calc = calc adslab.pbc = True print(f"Running data/adslabs/{bulk_src_id}_H/{idx}") opt = QuasiNewton( adslab, trajectory=f"data/adslabs/{bulk_src_id}_H/{idx}.traj", logfile=f"data/adslabs/{bulk_src_id}_H/{idx}.log", ) opt.run(fmax=0.05, steps=200) print( f" Elapsed time: {time.time() - t0:1.1f} seconds for data/adslabs/{bulk_src_id}_H/{idx}" ) for idx, adslab in enumerate(heuristic_adslabs_NNH.atoms_list): t0 = time.time() adslab.calc = calc adslab.pbc = True print(f"Running data/adslabs/{bulk_src_id}_NNH/{idx}") opt = QuasiNewton( adslab, trajectory=f"data/adslabs/{bulk_src_id}_NNH/{idx}.traj", logfile=f"data/adslabs/{bulk_src_id}_NNH/{idx}.log", ) opt.run(fmax=0.05, steps=relaxation_steps) print( f" Elapsed time: {time.time() - t0:1.1f} seconds for data/adslabs/{bulk_src_id}_NNH/{idx}" ) print(f"Elapsed time: {time.time() - tinit:1.1f} seconds") ``` # Parse the trajectories and post-process As a post-processing step we check to see if: 1. the adsorbate desorbed 2. the adsorbate disassociated 3. the adsorbate intercalated 4. the surface has changed We check these because they affect our referencing scheme and may result in energies that don't mean what we think, e.g. they aren't just adsorption, but include contributions from other things like desorption, dissociation or reconstruction. For (4), the relaxed surface should really be supplied as well. It will be necessary when correcting the SP / RX energies later. Since we don't have it here, we will ommit supplying it, and the detector will instead compare the initial and final slab from the adsorbate-slab relaxation trajectory. If a relaxed slab is provided, the detector will compare it and the slab after the adsorbate-slab relaxation. The latter is more correct! To compute the adsorption energies using the total energy UMA-OC20 model, we'll need the gas-phase reference energies from OC20 (see the original paper!). You could also calculate these quickly in DFT using a linear combination of H2O, H2, N2, and CO. ```{code-cell} ipython3 # reference energies from a linear combination of H2O/N2/CO/H2! atomic_reference_energies = { "H": -3.477, "N": -8.083, "O": -7.204, "C": -7.282, } ``` In this loop we find the most stable (most negative) adsorption energy for each adsorbate on each surface and save them in a DataFrame. ```{code-cell} ipython3 # Iterate over trajs to extract results min_E = [] for file_outer in glob("data/adslabs/*"): ads = file_outer.split("_")[1] bulk = file_outer.split("/")[-1].split("_")[0] slab = ase.io.read(f"data/slabs/{bulk}.traj") results = [] for file in glob(f"{file_outer}/*.traj"): rx_id = file.split("/")[-1].split(".")[0] traj = ase.io.read(file, ":") # Check to see if the trajectory is anomolous detector = DetectTrajAnomaly(traj[0], traj[-1], traj[0].get_tags()) anom = ( detector.is_adsorbate_dissociated() or detector.is_adsorbate_desorbed() or detector.has_surface_changed() or detector.is_adsorbate_intercalated() ) rx_energy = ( traj[-1].get_potential_energy() - slab.get_potential_energy() - sum( [ atomic_reference_energies[x] for x in traj[0][traj[0].get_tags() == 2].get_chemical_symbols() ] ) ) results.append( { "relaxation_idx": rx_id, "relaxed_atoms": traj[-1], "relaxed_energy_ml": rx_energy, "anomolous": anom, } ) df = pd.DataFrame(results) df = df[~df.anomolous].copy().reset_index() min_e = min(df.relaxed_energy_ml.tolist()) min_E.append({"adsorbate": ads, "bulk_id": bulk, "min_E_ml": min_e}) df = pd.DataFrame(min_E) df_h = df[df.adsorbate == "H"] df_nnh = df[df.adsorbate == "NNH"] df_flat = df_h.merge(df_nnh, on="bulk_id") ``` # Make parity plots for values obtained by ML v. reported in the paper ```{code-cell} ipython3 # Add literature data to the dataframe with open("literature_data.pkl", "rb") as f: literature_data = pickle.load(f) df_all = df_flat.merge(pd.DataFrame(literature_data), on="bulk_id") ``` ```{code-cell} ipython3 f, (ax1, ax2) = plt.subplots(1, 2, sharey=True) f.set_figheight(15) x = df_all.min_E_ml_x.tolist() y = df_all.E_lit_H.tolist() ax1.set_title("*H parity") ax1.plot([-3.5, 2], [-3.5, 2], "k-", linewidth=3) slope, intercept, r, p, se = linregress(x, y) ax1.plot( [-3.5, 2], [ -3.5 * slope + intercept, 2 * slope + intercept, ], "k--", linewidth=2, ) ax1.legend( [ "y = x", f"y = {slope:1.2f} x + {intercept:1.2f}, R-sq = {r**2:1.2f}", ], loc="upper left", ) ax1.scatter(x, y) ax1.axis("square") ax1.set_xlim([-3.5, 2]) ax1.set_ylim([-3.5, 2]) ax1.set_xlabel("dE predicted UMA [eV]") ax1.set_ylabel("dE NRR paper [eV]") x = df_all.min_E_ml_y.tolist() y = df_all.E_lit_NNH.tolist() ax2.set_title("*N*NH parity") ax2.plot([-3.5, 2], [-3.5, 2], "k-", linewidth=3) slope, intercept, r, p, se = linregress(x, y) ax2.plot( [-3.5, 2], [ -3.5 * slope + intercept, 2 * slope + intercept, ], "k--", linewidth=2, ) ax2.legend( [ "y = x", f"y = {slope:1.2f} x + {intercept:1.2f}, R-sq = {r**2:1.2f}", ], loc="upper left", ) ax2.scatter(x, y) ax2.axis("square") ax2.set_xlim([-3.5, 2]) ax2.set_ylim([-3.5, 2]) ax2.set_xlabel("dE predicted UMA [eV]") ax2.set_ylabel("dE NRR paper [eV]") f.set_figwidth(15) f.set_figheight(7) ``` # Make figure 6b and compare to literature results ```{code-cell} ipython3 f, (ax1, ax2) = plt.subplots(1, 2, sharey=True) x = df_all[df_all.reaction == "HER"].min_E_ml_y.tolist() y = df_all[df_all.reaction == "HER"].min_E_ml_x.tolist() comp = df_all[df_all.reaction == "HER"].composition.tolist() ax1.scatter(x, y, c="r", label="HER") for i, txt in enumerate(comp): ax1.annotate(txt, (x[i], y[i])) x = df_all[df_all.reaction == "NRR"].min_E_ml_y.tolist() y = df_all[df_all.reaction == "NRR"].min_E_ml_x.tolist() comp = df_all[df_all.reaction == "NRR"].composition.tolist() ax1.scatter(x, y, c="b", label="NRR") for i, txt in enumerate(comp): ax1.annotate(txt, (x[i], y[i])) ax1.legend() ax1.set_xlabel("dE *N*NH predicted UMA [eV]") ax1.set_ylabel("dE *H predicted UMA [eV]") x = df_all[df_all.reaction == "HER"].E_lit_NNH.tolist() y = df_all[df_all.reaction == "HER"].E_lit_H.tolist() comp = df_all[df_all.reaction == "HER"].composition.tolist() ax2.scatter(x, y, c="r", label="HER") for i, txt in enumerate(comp): ax2.annotate(txt, (x[i], y[i])) x = df_all[df_all.reaction == "NRR"].E_lit_NNH.tolist() y = df_all[df_all.reaction == "NRR"].E_lit_H.tolist() comp = df_all[df_all.reaction == "NRR"].composition.tolist() ax2.scatter(x, y, c="b", label="NRR") for i, txt in enumerate(comp): ax2.annotate(txt, (x[i], y[i])) ax2.legend() ax2.set_xlabel("dE *N*NH literature [eV]") ax2.set_ylabel("dE *H literature [eV]") f.set_figwidth(15) f.set_figheight(7) ``` --- Source: `docs/catalysts/examples_tutorials/adsorbml_walkthrough.md` # AdsorbML Tutorial :::{card} Tutorial Overview | Property | Value | |----------|-------| | **Difficulty** | Intermediate | | **Time** | 20-30 minutes | | **Prerequisites** | Basic Python, ASE | | **Goal** | Find optimal adsorption sites using ML-accelerated relaxations | ::: The [AdsorbML](https://arxiv.org/abs/2211.16486) paper showed that pre-trained machine learning potentials were now viable to find and prioritize the best adsorption sites for a given surface. The results were quite impressive, especially if you were willing to do a DFT single-point calculation on the best calculations. The latest UMA models are now total-energy models, and the results for the adsorption energy are even more impressive ([see the paper for details and benchmarks](https://ai.meta.com/research/publications/uma-a-family-of-universal-models-for-atoms/)). The AdsorbML package helps you with automated multi-adsorbate placement, and will automatically run calculations using the ML models to find the best sites to sample. ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ## Define desired adsorbate+slab system ```{code-cell} ipython3 from __future__ import annotations import pandas as pd from fairchem.data.oc.core import Bulk, Slab, Adsorbate from ase.build import molecule co_molecule = molecule("CO") adsorbate = Adsorbate(adsorbate_atoms=co_molecule, adsorbate_binding_indices=[1]) # 1 corresponds to the carbon atom # adsorbate = [Adsorbate(adsorbate_atoms=co_molecule, adsorbate_binding_indices=[1]) for _ in range(2)] # 2 COs bulk_src_id = "mp-30" bulk = Bulk(bulk_src_id_from_db=bulk_src_id) slabs = Slab.from_bulk_get_specific_millers(bulk=bulk, specific_millers=(1, 1, 1)) # There may be multiple slabs with this miller index. # For demonstrative purposes we will take the first entry. slab = slabs[0] ``` ## Run heuristic/random adsorbate placement and ML relaxations Now that we've defined the bulk, slab, and adsorbates of interest, we can quickly use the pre-trained UMA model as a calculator and the helper script `fairchem.core.components.calculate.recipes.adsorbml.run_adsorbml`. More details on the automated pipeline can be found at https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/components/calculate/recipes/adsorbml.py#L316. ```{code-cell} ipython3 from ase.optimize import LBFGS from fairchem.core import FAIRChemCalculator, pretrained_mlip from fairchem.core.components.calculate.recipes.adsorbml import run_adsorbml predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") outputs = run_adsorbml( slab=slab, adsorbate=adsorbate, calculator=calc, optimizer_cls=LBFGS, fmax=0.02, steps=20, # Increase to 200 for practical application, 20 is used for demonstrations num_placements=10, # Increase to 100 for practical application, 10 is used for demonstrations reference_ml_energies=True, # True if using a total energy model (i.e. UMA) relaxed_slab_atoms=None, place_on_relaxed_slab=False, ) ``` ```{code-cell} ipython3 top_candidates = outputs["adslabs"] global_min_candidate = top_candidates[0] ``` ```{code-cell} ipython3 top_candidates = outputs["adslabs"] pd.DataFrame(top_candidates) ``` ## Write VASP input files If you want to verify the results, you should run VASP. This assumes you have access to VASP pseudopotentials. The default VASP flags (which are equivalent to those used to make OC20) are located in `ocdata.utils.vasp`. Alternatively, you may pass your own vasp flags to the `write_vasp_input_files` function as `vasp_flags`. Note that to run this you need access to the VASP pseudopotentials and need to have those set up in ASE. ```{code-cell} ipython3 :tags: [skip-execution] import os from fairchem.data.oc.utils.vasp import write_vasp_input_files # Grab the 5 systems with the lowest energy top_5_candidates = top_candidates[:5] # Write the inputs for idx, config in enumerate(top_5_candidates): os.makedirs(f"data/{idx}", exist_ok=True) write_vasp_input_files(config["atoms"], outdir=f"data/{idx}/") ``` --- Source: `docs/catalysts/examples_tutorials/cattsunami_tutorial.md` # Transition State Search (NEBs) :::{card} Tutorial Overview | Property | Value | |----------|-------| | **Difficulty** | Advanced | | **Time** | 30-45 minutes | | **Prerequisites** | Understanding of NEB, ASE, catalysis | | **Goal** | Find transition states using CatTsunami tools | ::: FAIR chemistry models can be used to enumerate and study reaction pathways via transition state search tools built into ASE or in packages like Sella via the ASE interface. :::{note} The first section of this tutorial walks through how to use the CatTsunami tools to automatically enumerate a number of hypothetical initial/final configurations for various types of reactions on a heterogeneous catalyst surface. If you already have a NEB you're looking to optimize, you can jump straight to the last section (Run NEBs). ::: Since the NEB calculations here can be a bit time consuming, we'll use a small number of steps during the documentation testing, and otherwise use a reasonable guess. ```{code-cell} ipython3 import os # Use a small number of steps here to keep the docs fast during CI, but otherwise do quite reasonable settings. fast_docs = os.environ.get("FAST_DOCS", "false").lower() == "true" if fast_docs: optimization_steps = 20 else: optimization_steps = 300 ``` ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ## Do enumerations in an AdsorbML style ```{code-cell} ipython3 from __future__ import annotations import matplotlib.pyplot as plt from ase.io import read from ase.mep import DyNEB from ase.optimize import BFGS from fairchem.applications.cattsunami.core import Reaction from fairchem.applications.cattsunami.core.autoframe import AutoFrameDissociation from fairchem.applications.cattsunami.databases import DISSOCIATION_REACTION_DB_PATH from fairchem.core import FAIRChemCalculator, pretrained_mlip from fairchem.data.oc.core import Adsorbate, AdsorbateSlabConfig, Bulk, Slab from fairchem.data.oc.databases.pkls import ADSORBATE_PKL_PATH, BULK_PKL_PATH from x3dase.x3d import X3D # Instantiate the reaction class for the reaction of interest reaction = Reaction( reaction_str_from_db="*CH -> *C + *H", reaction_db_path=DISSOCIATION_REACTION_DB_PATH, adsorbate_db_path=ADSORBATE_PKL_PATH, ) ``` ```{code-cell} ipython3 # Instantiate our adsorbate class for the reactant and product reactant = Adsorbate( adsorbate_id_from_db=reaction.reactant1_idx, adsorbate_db_path=ADSORBATE_PKL_PATH ) product1 = Adsorbate( adsorbate_id_from_db=reaction.product1_idx, adsorbate_db_path=ADSORBATE_PKL_PATH ) product2 = Adsorbate( adsorbate_id_from_db=reaction.product2_idx, adsorbate_db_path=ADSORBATE_PKL_PATH ) ``` ```{code-cell} ipython3 # Grab the bulk and cut the slab we are interested in bulk = Bulk(bulk_src_id_from_db="mp-33", bulk_db_path=BULK_PKL_PATH) slab = Slab.from_bulk_get_specific_millers(bulk=bulk, specific_millers=(0, 0, 1)) ``` ```{code-cell} ipython3 # Perform site enumeration # For AdsorbML num_sites = 100, but we use 5 here for brevity. This should be increased for practical use. reactant_configs = AdsorbateSlabConfig( slab=slab[0], adsorbate=reactant, mode="random_site_heuristic_placement", num_sites=10, ).atoms_list product1_configs = AdsorbateSlabConfig( slab=slab[0], adsorbate=product1, mode="random_site_heuristic_placement", num_sites=10, ).atoms_list product2_configs = AdsorbateSlabConfig( slab=slab[0], adsorbate=product2, mode="random_site_heuristic_placement", num_sites=10, ).atoms_list ``` ```{code-cell} ipython3 # Instantiate the calculator predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") ``` ```{code-cell} ipython3 # Relax the reactant systems reactant_energies = [] for config in reactant_configs: config.calc = calc config.pbc = True opt = BFGS(config) opt.run(fmax=0.05, steps=optimization_steps) reactant_energies.append(config.get_potential_energy()) ``` ```{code-cell} ipython3 # Relax the product systems product1_energies = [] for config in product1_configs: config.calc = calc config.pbc = True opt = BFGS(config) opt.run(fmax=0.05, steps=optimization_steps) product1_energies.append(config.get_potential_energy()) ``` ```{code-cell} ipython3 product2_energies = [] for config in product2_configs: config.calc = calc config.pbc = True opt = BFGS(config) opt.run(fmax=0.05, steps=optimization_steps) product2_energies.append(config.get_potential_energy()) ``` ## Enumerate NEBs ![](dissociation_scheme.png) ```{code-cell} ipython3 af = AutoFrameDissociation( reaction=reaction, reactant_system=reactant_configs[reactant_energies.index(min(reactant_energies))], product1_systems=product1_configs, product1_energies=product1_energies, product2_systems=product2_configs, product2_energies=product2_energies, r_product1_max=2, # r1 in the above fig r_product2_max=3, # r3 in the above fig r_product2_min=1, # r2 in the above fig ) ``` ```{code-cell} ipython3 import random nframes = 10 random.seed( 42 ) # set the seed to make the random generation deterministic for the tutorial! frame_sets, mapping_idxs = af.get_neb_frames( calc, n_frames=nframes, n_pdt1_sites=4, # = 5 in the above fig (step 1) n_pdt2_sites=4, # = 5 in the above fig (step 2) ) ``` ## Run NEBs ```{code-cell} ipython3 :tags: [skip-execution] ## This will run all NEBs enumerated - to just run one, run the code cell below. # On GPU, each NEB takes an average of ~1 minute so this could take around a half hour on GPU # But much longer on CPU # Remember that not all NEBs will converge -- the k, nframes would be adjusted to achieve convergence fmax = 0.05 # [eV / ang**2] delta_fmax_climb = 0.4 converged_idxs = [] for idx, frame_set in enumerate(frame_sets): neb = DyNEB(frame_set, k=1) for image in frame_set: image.calc = FAIRChemCalculator(predictor, task_name="oc20") optimizer = BFGS( neb, trajectory=f"ch_dissoc_on_Ru_{idx}.traj", ) conv = optimizer.run(fmax=fmax + delta_fmax_climb, steps=optimization_steps) if conv: neb.climb = True conv = optimizer.run(fmax=fmax, steps=optimization_steps) if conv: converged_idxs.append(idx) print(converged_idxs) ``` This cell will run a shorter calculations for just a single one of the enumerated transition state pathways. You can adapt this code to run transition state searches via nudged elastic band (NEB) calculations for any reaction. ```{code-cell} ipython3 # If you run the above cell -- dont run this one fmax = 0.05 # [eV / ang**2] delta_fmax_climb = 0.4 images = frame_sets[0] neb = DyNEB(images, k=1) for image in images: image.calc = FAIRChemCalculator(predictor, task_name="oc20") optimizer = BFGS( neb, trajectory="ch_dissoc_on_Ru_0.traj", ) conv = optimizer.run(fmax=fmax + delta_fmax_climb, steps=optimization_steps) if conv: neb.climb = True conv = optimizer.run(fmax=fmax, steps=optimization_steps) ``` ## Visualize the results Finally, let's visualize the results! ```{code-cell} ipython3 optimized_neb = read("ch_dissoc_on_Ru_0.traj", ":")[-1 * nframes :] ``` ```{code-cell} ipython3 es = [] for frame in optimized_neb: frame.set_calculator(calc) es.append(frame.get_potential_energy()) ``` ```{code-cell} ipython3 # Plot the reaction coordinate es = [e - es[0] for e in es] plt.plot(es) plt.xlabel("frame number") plt.ylabel("relative energy [eV]") plt.title(f"CH dissociation on Ru(0001), Ea = {max(es):1.2f} eV") plt.savefig("CH_dissoc_on_Ru_0001.png") ``` To generalize an interactive visualization, use `ase gui` from the command line or the X3D package ```{code-cell} ipython3 :tags: [skip-execution] # Make an interative html file of the optimized neb trajectory x3d = X3D(optimized_neb) x3d.write("optimized_neb_ch_disoc_on_Ru0001.html") ``` --- Source: `docs/catalysts/examples_tutorials/ocpapi.md` # ocpapi :::{card} API Overview | Property | Value | |----------|-------| | **Purpose** | Programmatic access to Open Catalyst Demo | | **Interface** | Python async/await | | **License** | [MIT](https://github.com/facebookresearch/fairchem/blob/main/LICENSE.md) | ::: Python library for programmatic use of the [Open Catalyst Demo](https://open-catalyst.metademolab.com/). Users unfamiliar with the Open Catalyst Demo are encouraged to read more about it before continuing. ## Installation ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` ## Quickstart The following examples are used to search for *OH binding sites on Pt surfaces. They use the `find_adsorbate_binding_sites` function, which is a high-level workflow on top of other methods included in this library. Once familiar with this routine, users are encouraged to learn about lower-level methods and features that support more advanced use cases. ### Note about async methods This package relies heavily on [asyncio](https://docs.python.org/3/library/asyncio.html). The examples throughout this document can be copied to a python repl launched with: ```{code-cell} ipython3 :tags: ["skip-execution"] %%sh $ python -m asyncio ``` Alternatively, an async function can be run in a script by wrapping it with [asyncio.run()](https://docs.python.org/3/library/asyncio-runner.html#asyncio.run): ```{code-cell} ipython3 :tags: ["skip-execution"] import asyncio from fairchem.demo.ocpapi import find_adsorbate_binding_sites asyncio.run(find_adsorbate_binding_sites(...)) ``` Since this is being evaluated as a jupyter notebook, ipython will handle this for you automatically! ### Search over all surfaces ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import find_adsorbate_binding_sites results = await find_adsorbate_binding_sites( adsorbate="*OH", bulk="mp-126", ) ``` Users will be prompted to select one or more surfaces that should be relaxed. Input to this function includes: * The name of the adsorbate to place * A unique ID of the bulk structure from which surfaces will be generated This function will perform the following steps: 1. Enumerate surfaces of the bulk material 2. On each surface, enumerate initial guesses for adorbate binding sites 3. Run local force-based relaxations of each adsorbate placement In addition, this handles: * Retrying failed calls to the Open Catalyst Demo API * Retrying submission of relaxations when they are rate limited This should take 2-10 minutes to finish while tens to hundreds (depending on the number of surfaces that are selected) of individual adsorbate placements are relaxed on unique surfaces of Pt. Each of the objects in the returned list includes (among other details): * Information about the surface being searched, including its structure and Miller indices * The initial positions of the adsorbate before relaxation * The final structure after relaxation * The predicted energy of the final structure * The predicted force on each atom in the final structure +++ ### Supported bulks and adsorbates A finite set of bulk materials and adsorbates can be referenced by ID throughout the OCP API. The lists of supported values can be viewed in two ways. 1. Visit the UI at https://open-catalyst.metademolab.com/demo and explore the lists in Step 1 and Step 3. 2. Use the low-level client that ships with this library: ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import Client client = Client() bulks = await client.get_bulks() print({b.src_id: b.formula for b in bulks.bulks_supported}) adsorbates = await client.get_adsorbates() print(adsorbates.adsorbates_supported) ``` ### Skip relaxation approval prompts Calls to `find_adsorbate_binding_sites()` will, by default, show the user all pending relaxations and ask for approval before they are submitted. In order to run the relaxations automatically without manual approval, `adslab_filter` can be set to a function that automatically approves any or all adsorbate/slab (adslab) configurations. Run relaxations for all slabs that are generated: ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import find_adsorbate_binding_sites, keep_all_slabs results = await find_adsorbate_binding_sites( adsorbate="*OH", bulk="mp-126", adslab_filter=keep_all_slabs(), ) ``` Run relaxations only for slabs with Miller Indices in the input set: ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import find_adsorbate_binding_sites, keep_slabs_with_miller_indices results = await find_adsorbate_binding_sites( adsorbate="*OH", bulk="mp-126", adslab_filter=keep_slabs_with_miller_indices([(1, 0, 0), (1, 1, 1)]), ) print(results) ``` ### Persisting results **Results should be saved whenever possible in order to avoid expensive recomputation.** Assuming `results` was generated with the `find_adsorbate_binding_sites` method used above, it is an `AdsorbateBindingSites` object. This can be saved to file with: ```{code-cell} ipython3 :tags: ["skip-execution"] with open("results.json", "w") as f: f.write(results.to_json()) ``` Similarly, results can be read back from file to an `AdsorbateBindingSites` object with: ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import AdsorbateBindingSites with open("results.json", "r") as f: results = AdsorbateBindingSites.from_json(f.read()) ``` ### Viewing results in the web UI Relaxation results can be viewed in a web UI. For example, https://open-catalyst.metademolab.com/results/7eaa0d63-83aa-473f-ac84-423ffd0c67f5 shows the results of relaxing *OH on a Pt (1,1,1) surface; the uuid, "7eaa0d63-83aa-473f-ac84-423ffd0c67f5", is referred to as the `system_id`. Extending the examples above, the URLs to visualize the results of relaxations on each Pt surface can be obtained with: ```{code-cell} ipython3 :tags: ["skip-execution"] print([ slab.ui_url for slab in results.slabs ]) ``` ## Advanced usage ### Changing the model type The API currently supports two models: * `equiformer_v2_31M_s2ef_all_md` (default): https://arxiv.org/abs/2306.12059 * `gemnet_oc_base_s2ef_all_md`: https://arxiv.org/abs/2204.02782 A specific model type can be requested with: ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import find_adsorbate_binding_sites results = await find_adsorbate_binding_sites( adsorbate="*OH", bulk="mp-126", model="gemnet_oc_base_s2ef_all_md", adslab_filter=keep_slabs_with_miller_indices([(1, 1, 1)]), ) print([ slab.ui_url for slab in results.slabs ]) ``` ### Converting to [ase.Atoms](https://wiki.fysik.dtu.dk/ase/ase/atoms.html) objects **Important! The `to_ase_atoms()` method described below will fail with an import error if [ase](https://wiki.fysik.dtu.dk/ase) is not installed.** Two classes have support for generating [ase.Atoms](https://wiki.fysik.dtu.dk/ase/ase/atoms.html) objects: * `ocpapi.Atoms.to_ase_atoms()`: Adds unit cell, atomic positions, and other structural information to the returned `ase.Atoms` object. * `ocpapi.AdsorbateSlabRelaxationResult.to_ase_atoms()`: Adds the same structure information to the `ase.Atoms` object. Also adds the predicted forces and energy of the relaxed structure, which can be accessed with the `ase.Atoms.get_potential_energy()` and `ase.Atoms.get_forces()` methods. For example, the following would generate an `ase.Atoms` object for the first relaxed adsorbate configuration on the first slab generated for *OH binding on Pt: ```{code-cell} ipython3 :tags: ["skip-execution"] from fairchem.demo.ocpapi import find_adsorbate_binding_sites results = await find_adsorbate_binding_sites( adsorbate="*OH", bulk="mp-126", adslab_filter=keep_slabs_with_miller_indices([(1, 1, 1)]), ) ase_atoms = results.slabs[0].configs[0].to_ase_atoms() print(ase_atoms) ``` ### Converting to other structure formats From an `ase.Atoms` object (see previous section), is is possible to [write to other structure formats](https://wiki.fysik.dtu.dk/ase/ase/io/io.html#ase.io.write). Extending the example above, the `ase_atoms` object could be written to a [VASP POSCAR file](https://www.vasp.at/wiki/index.php/POSCAR) with: ```{code-cell} ipython3 :tags: ["skip-execution"] from ase.io import write write("POSCAR", ase_atoms, "vasp") ``` ## License `ocpapi` is released under the [MIT License](https://github.com/facebookresearch/fairchem/blob/main/LICENSE.md). ## Citing `ocpapi` If you use `ocpapi` in your research, please consider citing the [AdsorbML paper](https://www.nature.com/articles/s41524-023-01121-5) (in addition to the relevant datasets / models used): ```bibtex @article{lan2023adsorbml, title={{AdsorbML}: a leap in efficiency for adsorption energy calculations using generalizable machine learning potentials}, author={Lan*, Janice and Palizhati*, Aini and Shuaibi*, Muhammed and Wood*, Brandon M and Wander, Brook and Das, Abhishek and Uyttendaele, Matt and Zitnick, C Lawrence and Ulissi, Zachary W}, journal={npj Computational Materials}, year={2023}, } ``` --- Source: `docs/dac/datasets/summary.md` # Datasets :::{margin} ```{image} ../../assets/icons/mofs-dac.svg :alt: MOFs for Direct Air Capture :width: 100px ``` ::: Explore the Direct Air Capture (DAC) datasets for training machine learning models on metal-organic frameworks (MOFs). ::::{grid} 1 2 2 2 :::{card} ODAC25 :link: odac25 Nearly 70 million DFT calculations for CO2, H2O, N2, and O2 adsorption in nearly 15,000 MOFs. ::: :::{card} ODAC23 (Deprecated) :link: odac23 The original Open DAC dataset. Superseded by ODAC25. ::: :::: --- Source: `docs/dac/datasets/odac25.md` # ODAC25 :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | ~70 million DFT calculations | | **Domain** | Metal-organic frameworks (MOFs) | | **Adsorbates** | CO2, H2O, N2, O2 | | **MOFs** | ~15,000 structures | | **Labels** | Total energy (eV), forces (eV/A) | | **Level of Theory** | PBE+D3 (VASP) | | **License** | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | | **Download** | [HuggingFace](https://huggingface.co/facebook/ODAC25) | ::: The Open DAC 2025 (ODAC25) dataset contains nearly 70 million DFT single-point calculations for CO2, H2O, N2, and O2 adsorption in nearly 15,000 metal-organic frameworks (MOFs). This dataset represents a significant expansion upon ODAC23, introducing greater chemical and configurational diversity through functionalized MOFs, high-energy GCMC-derived placements, and synthetically generated frameworks. The dataset enables training of state-of-the-art machine-learned interatomic potentials for direct air capture applications. All structures are labeled with total energies (eV) and forces (eV/Å) computed using VASP with the PBE+D3 functional. All information about the dataset is available at the [ODAC25 Huggingface site](https://huggingface.co/facebook/ODAC25). For questions or issues, please open a GitHub issue in this repository. ## Dataset format The dataset is provided in ASE DB compatible lmdb files (`*.aselmdb`). ### Citing ODAC25 The ODAC25 dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any publications that uses this dataset: ```bib @misc{sriram2025odac25, title={The Open DAC 2025 Dataset for Sorbent Discovery in Direct Air Capture}, author={Anuroop Sriram and Logan M. Brabson and Xiaohan Yu and Sihoon Choi and Kareem Abdelmaqsoud and Elias Moubarak and Pim de Haan and Sindy Löwe and Johann Brehmer and John R. Kitchin and Max Welling and C. Lawrence Zitnick and Zachary Ulissi and Andrew J. Medford and David S. Sholl}, year={2025}, eprint={}, archivePrefix={arXiv}, primaryClass={}, url={}, } ``` --- Source: `docs/dac/datasets/odac23.md` # Open Direct Air Capture 2023 (ODAC23) :::{warning} **Deprecated**: ODAC23 has been superseded by [ODAC25](odac25.md), which contains nearly 70 million DFT calculations across 4 adsorbates (CO2, H2O, N2, O2) in nearly 15,000 MOFs. We recommend using ODAC25 for new projects. ::: :::{card} Dataset Overview | Property | Value | |----------|-------| | **Size** | ~1M structures (S2EF), ~10K structures (IS2RE) | | **Domain** | Metal-organic frameworks (MOFs) with CO2 | | **Labels** | Total energy (eV), forces (eV/A) | | **Level of Theory** | PBE+D3 (VASP) | | **License** | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) | | **Status** | Deprecated - use ODAC25 | ::: ## Structure to Energy and Forces (S2EF) task We provide precomputed LMDBs for train, validation, and the various test sets that can be used directly with the dataloaders provided in our code. The LMDBs contain input structures from all points in relaxation trajectories along with the energy of the structure and the atomic forces. The dataset contains an in-domain test set and 4 out-of-domain test sets (ood-large, ood-linker, ood-topology, and ood-linker & topology). All LMDbs are compressed into a single `.tar.gz` file. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train + Validation + Test (all splits) | 172G | 476G | [162f0660b2f1c9209c5b57f7b9e545a7](https://dl.fbaipublicfiles.com/large_objects/dac/datasets/odac23_s2ef.tar.gz ) | | | | | | The train and val splits are also available in `extxyz` formats. Each trajectory is in stored in a different `extxyz` file. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train | 232G | 781G | [381e72fd8b9c055065fd3afff6b0945b](https://dl.fbaipublicfiles.com/large_objects/dac/datasets/extxyz_train.tar.gz ) | |Val | 5.1G | 18G | [09913759c6e0f8d649f7ec9dff9e0e8b](https://dl.fbaipublicfiles.com/dac/datasets/extxyz_val.tar.gz ) | | | | | | ## Initial Structure to Relaxed Structure (IS2RS) / Relaxed Energy (IS2RE) tasks For IS2RE / IS2RS training, validation and test sets, we provide precomputed LMDBs that can be directly used with dataloaders provided in our code. The LMDBs contain input initial structures and the output relaxed structures and energies. The dataset contains an in-domain test set and 4 out-of-domain test sets (ood-large, ood-linker, ood-topology, and ood-linker & topology). All LMDBs are compressed into a single `.tar.gz` file. |Splits |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- |--- | |Train + Validation + Test (all splits) | 809M | 2.2G | [f7f2f58669a30abae8cb9ba1b7f2bcd2](https://dl.fbaipublicfiles.com/dac/datasets/odac23_is2r.tar.gz ) | | | | | | ## DDEC Charges We provide DDEC charges computed for all MOFs in the ODAC23 dataset. A small number of MOFs (~2%) are missing these charges because the DDEC calcuations failed for them. |Size of compressed version (in bytes) |Size of uncompressed version (in bytes) | MD5 checksum (download link) | |--- |--- |--- | | 147M | 534M | [81927b78d9e4184cc3c398e79760126a](https://dl.fbaipublicfiles.com/dac/datasets/ddec.tar.gz ) | | | | | ## Citing ODAC23 The OpenDAC 2023 (ODAC23) dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any research manuscript using the ODAC23 dataset: ```bibtex @article{odac23_dataset, author = {Anuroop Sriram and Sihoon Choi and Xiaohan Yu and Logan M. Brabson and Abhishek Das and Zachary Ulissi and Matt Uyttendaele and Andrew J. Medford and David S. Sholl}, title = {The Open DAC 2023 Dataset and Challenges for Sorbent Discovery in Direct Air Capture}, year = {2023}, journal={arXiv preprint arXiv:2311.00341}, } ``` --- Source: `docs/dac/models.md` # Pretrained models :::{margin} ```{image} ../assets/icons/mofs-dac.svg :alt: MOFs for Direct Air Capture :width: 100px ``` ::: :::{tip} Recommended Models **For CO2 and H2O adsorption**: Use [UMA](../core/uma), trained on all FAIR chemistry datasets including ODAC25. **For N2 and O2 adsorption**: Use the dedicated eSEN models trained on ODAC25, as UMA was not trained on these adsorbates. ::: ## ODAC25 Models As part of the ODAC25 release, we released two sets of models: 1. UMA models trained on a range of FAIR chemistry datasets including a subset of ODAC25, available at the [UMA HuggingFace site](https://huggingface.co/facebook/UMA) 2. eSEN models trained only on ODAC25, available at the [ODAC25 HuggingFace site](https://huggingface.co/facebook/ODAC25) The UMA models were only trained on CO₂ and H₂O adsorbates, and are competitive with the eSEN models for these adsorbates. Since the UMA models were not trained on N₂ and O₂, we strongly recommend using the eSEN models for these. If you use the ODAC25 trained models, please cite the following paper: ```bib @misc{sriram2025odac25, title={The Open DAC 2025 Dataset for Sorbent Discovery in Direct Air Capture}, author={Anuroop Sriram and Logan M. Brabson and Xiaohan Yu and Sihoon Choi and Kareem Abdelmaqsoud and Elias Moubarak and Pim de Haan and Sindy Löwe and Johann Brehmer and John R. Kitchin and Max Welling and C. Lawrence Zitnick and Zachary Ulissi and Andrew J. Medford and David S. Sholl}, year={2025}, eprint={}, archivePrefix={arXiv}, primaryClass={}, url={}, } ``` ## ODAC23 Models :::{warning} **Deprecated**: These models are superseded by ODAC25 models and UMA. They are provided here for reproducibility only. ::: * All config files for the ODAC23 models are available in the [`configs/odac`](https://github.com/facebookresearch/fairchem/tree/main/configs/odac) directory. ### S2EF models | Model Name |Model |Checkpoint | Config | |------------------------------|--- |--- |--- | | SchNet-S2EF-ODAC | SchNet | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/Schnet.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/schnet.yml) | | DimeNet++-S2EF-ODAC | DimeNet++ | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/DimenetPP.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/dpp.yml) | | PaiNN-S2EF-ODAC | PaiNN | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/PaiNN.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/painn.yml) | | GemNet-OC-S2EF-ODAC | GemNet-OC | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/Gemnet-OC.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/gemnet-oc.yml) | | eSCN-S2EF-ODAC | eSCN | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/eSCN.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/eSCN.yml) | | EquiformerV2-S2EF-ODAC | EquiformerV2 | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231116/eqv2_31M.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/eqv2_31M.yml) | | EquiformerV2-Large-S2EF-ODAC | EquiformerV2 (Large) | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231116/Equiformer_V2_Large.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/s2ef/eqv2_153M.yml) | ### IS2RE Direct models | Model Name | Model |Checkpoint | Config | |-------------------------|--------------|--- | --- | | Gemnet-OC-IS2RE-ODAC | Gemnet-OC | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/Gemnet-OC_Direct.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/is2re/gemnet-oc.yml) | | eSCN-IS2RE-ODAC | eSCN | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231018/eSCN_Direct.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/is2re/eSCN.yml) | | EquiformerV2-IS2RE-ODAC | EquiformerV2 | [checkpoint](https://dl.fbaipublicfiles.com/dac/checkpoints_20231116/Equiformer_V2_Direct.pt) | [config](https://github.com/facebookresearch/fairchem/tree/main/configs/odac/is2re/eqv2_31M.yml) | The models in the table above were trained to predict relaxed energy directly. Relaxed energies can also be predicted by running structural relaxations using the S2EF models from the previous section. ### IS2RS The IS2RS is solved by running structural relaxations using the S2EF models from the prior section. The Open DAC 2023 (ODAC23) dataset is licensed under a [Creative Commons Attribution 4.0 License](https://creativecommons.org/licenses/by/4.0/legalcode). Please consider citing the following paper in any research manuscript using the ODAC23 dataset: ```bibtex @article{odac23_dataset, author = {Anuroop Sriram and Sihoon Choi and Xiaohan Yu and Logan M. Brabson and Abhishek Das and Zachary Ulissi and Matt Uyttendaele and Andrew J. Medford and David S. Sholl}, title = {The Open DAC 2023 Dataset and Challenges for Sorbent Discovery in Direct Air Capture}, year = {2023}, journal={arXiv preprint arXiv:2311.00341}, } ``` --- Source: `docs/dac/examples_tutorials/summary.md` # Examples & Tutorials :::{margin} ```{image} ../../assets/icons/mofs-dac.svg :alt: MOFs for Direct Air Capture :width: 100px ``` ::: Learn how to use FAIRChem models for Direct Air Capture (DAC) applications with metal-organic frameworks (MOFs). ::::{grid} 1 2 2 2 :::{card} Adsorption Energy :link: adsorption-energy Calculate CO2 and H2O adsorption energies in MOFs, including flexible framework effects. ::: :::: --- Source: `docs/dac/examples_tutorials/adsorption_energy.md` # Adsorption Energies in MOFs :::{tip} What You Will Learn Calculate CO2 and H2O adsorption energies in metal-organic frameworks, including accounting for MOF flexibility and deformation. ::: Pre-trained ODAC models are versatile across various MOF-related tasks. To begin, we'll start with a fundamental application: calculating the adsorption energy for a single CO2 molecule. This serves as an excellent and simple demonstration of what you can achieve with these datasets and models. For predicting the adsorption energy of a single CO2 molecule within a MOF structure, the adsorption energy ($E_{\mathrm{ads}}$) is defined as: $$ E_{\mathrm{ads}} = E_{\mathrm{MOF+CO2}} - E_{\mathrm{MOF}} - E_{\mathrm{CO2}} \tag{1}$$ Each term on the right-hand side represents the energy of the relaxed state of the indicated chemical system. For a comprehensive understanding of our methodology for computing these adsorption energies, please refer to our [paper](https://doi.org/10.1021/acscentsci.3c01629). ## Loading Pre-trained Models ````{admonition} Need to install fairchem-core or get UMA access or getting permissions/401 errors? :class: dropdown 1. Install the necessary packages using pip, uv etc ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami ``` 2. Get access to any necessary huggingface gated models * Get and login to your Huggingface account * Request access to https://huggingface.co/facebook/UMA * Create a Huggingface token at https://huggingface.co/settings/tokens/ with the permission "Permissions: Read access to contents of all public gated repos you can access" * Add the token as an environment variable using `huggingface-cli login` or by setting the HF_TOKEN environment variable. ```{code-cell} ipython3 :tags: [skip-execution] # Login using the huggingface-cli utility ! huggingface-cli login # alternatively, import os os.environ['HF_TOKEN'] = 'MY_TOKEN' ``` ```` A pre-trained model can be loaded using `FAIRChemCalculator`. In this example, we'll employ UMA to determine the CO2 adsorption energies. ```{code-cell} from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="odac") ``` ## Adsorption in rigid MOFs: CO2 Adsorption Energy in Mg-MOF-74 Let's apply our knowledge to Mg-MOF-74, a widely studied MOF known for its excellent CO2 adsorption properties. Its structure comprises magnesium atomic complexes connected by a carboxylated and oxidized benzene ring, serving as an organic linker. Previous studies consistently report the CO2 adsorption energy for Mg-MOF-74 to be around -0.40 eV [[1]](https://doi.org/10.1039/C4SC02064B) [[2]](https://doi.org/10.1039/C3SC51319J) [[3]](https://doi.org/10.1021/acs.jpcc.8b00938). Our goal is to verify if we can achieve a similar value by performing a simple single-point calculation using UMA. In the ODAC23 dataset, all MOF structures are identified by their CSD (Cambridge Structural Database) code. For Mg-MOF-74, this code is **OPAGIX**. We've extracted a specific `OPAGIX+CO2` configuration from the dataset, which exhibits the lowest adsorption energy among its counterparts. ```{code-cell} import matplotlib.pyplot as plt from ase.io import read from ase.visualize.plot import plot_atoms mof_co2 = read("structures/OPAGIX_w_CO2.cif") mof = read("structures/OPAGIX.cif") co2 = read("structures/co2.xyz") fig, ax = plt.subplots(figsize=(5, 4.5), dpi=250) plot_atoms(mof_co2, ax) ax.set_axis_off() ``` The final step in calculating the adsorption energy involves connecting the `FAIRChemCalculator` to each relaxed structure: `OPAGIX+CO2`, `OPAGIX`, and `CO2`. The structures used here are already relaxed from ODAC23. For simplicity, we assume here that further relaxations can be neglected. We will show how to go beyond this assumption in the next section. ```{code-cell} mof_co2.calc = calc mof.calc = calc co2.calc = calc E_ads = ( mof_co2.get_potential_energy() - mof.get_potential_energy() - co2.get_potential_energy() ) print(f"Adsorption energy of CO2 in Mg-MOF-74: {E_ads:.3f} eV") ``` ## Adsorption in flexible MOFs The adsorption energy calculation method outlined above is typically performed with rigid MOFs for simplicity. Both experimental and modeling literature have shown, however, that MOF flexibility can be important in accurately capturing the underlying chemistry of adsorption [[1]](https://arxiv.org/abs/2506.09256) [[2]](https://pubs.acs.org/doi/10.1021/jacs.7b01688) [[3]](https://www.nature.com/articles/nature15732). In particular, uptake can be improved by treating MOFs as flexible. Two types of MOF flexibility can be considered: intrinsic flexibility and deformation induced by guest molecules. In the Open DAC Project, we consider the latter MOF deformation by allowing the atomic positions of the MOF to relax during geometry optimization [[4]](https://pubs.acs.org/doi/10.1021/acscentsci.3c01629). The addition of additional degrees of freedoms can complicate the computation of the adsorption energy and necessitates an extra step in the calculation procedure. The figure below shows water adsorption in the MOF with CSD code WOBHEB with added defects (`WOBHEB_0.11_0`) from a DFT simulation. A typical adsorption energy calculation would only seek to capture the effects shaded in purple, which include both chemisorption and non-bonded interactions between the host and guest molecule. When allowing the MOF to relax, however, the adsorption energy also includes the energetic effect of the MOF deformation highlighted in green. +++ ![](./WOBHEB_flexible.png) To account for this deformation, it is vital to use the most energetically favorable MOF geometry for the empty MOF term in Eqn. 1. Including MOF atomic coordinates as degrees of freedom can result in three possible outcomes: 1. The MOF does not deform, so the energies of the relaxed empty MOF and the MOF in the adsorbed state are the same 2. The MOF deforms to a less energetically favorable geometry than its ground state 3. The MOF locates a new energetically favorable geoemtry relative to the empty MOF relaxation The first outcome requires no additional computation because the MOF rigidity assumption is valid. The second outcome represents physical and reversible deformation where the MOF returns to its empty ground state upon removal of the guest molecule. The third outcome is often the result of the guest molecule breaking local symmetry. We also found cases in ODAC in which both outcomes 2 and 3 occur within the same MOF. To ensure the most energetically favorable empty MOF geometry is found, an addition empty MOF relaxation should be performed after MOF + adsorbate relaxation. The guest molecule should be removed, and the MOF should be relaxed starting from its geometry in the adsorbed state. If all deformation is reversible, the MOF will return to its original empty geometry. Otherwise, the lowest energy (most favorable) MOF geometry should be taken as the reference energy, $E_{\mathrm{MOF}}$, in Eqn. 1. ### H2O Adsorption Energy in Flexible WOBHEB with UMA The first part of this tutorial demonstrates how to perform a single point adsorption energy calculation using UMA. To treat MOFs as flexible, we perform all calculations on geometries determined by geometry optimization. The following example corresponds to the figure shown above (H2O adsorption in `WOBHEB_0.11_0`). **In this tutorial, $E_{x}(r_{y})$ corresponds to the energy of $x$ determined from geometry optimization of $y$.** First, we obtain the energy of the empty MOF from relaxation of only the MOF: $E_{\mathrm{MOF}}(r_{\mathrm{MOF}})$ ```{code-cell} import ase.io from ase.optimize import BFGS mof = ase.io.read("structures/WOBHEB_0.11.cif") mof.calc = calc relax = BFGS(mof) relax.run(fmax=0.05) E_mof_empty = mof.get_potential_energy() print(f"Energy of empty MOF: {E_mof_empty:.3f} eV") ``` Next, we add the H2O guest molecule and relax the MOF + adsorbate to obtain $E_{\mathrm{MOF+H2O}}(r_{\mathrm{MOF+H2O}})$. ```{code-cell} mof_h2o = ase.io.read("structures/WOBHEB_H2O.cif") mof_h2o.calc = calc relax = BFGS(mof_h2o) relax.run(fmax=0.05) E_combo = mof_h2o.get_potential_energy() print(f"Energy of MOF + H2O: {E_combo:.3f} eV") ``` We can now isolate the MOF atoms from the relaxed MOF + H2O geometry and see that the MOF has adopted a geometry that is less energetically favorable than the empty MOF by ~0.2 eV. The energy of the MOF in the adsorbed state corresponds to $E_{\mathrm{MOF}}(r_{\mathrm{MOF+H2O}})$. ```{code-cell} mof_adsorbed_state = mof_h2o[:-3] mof_adsorbed_state.calc = calc E_mof_adsorbed_state = mof_adsorbed_state.get_potential_energy() print(f"Energy of MOF in the adsorbed state: {E_mof_adsorbed_state:.3f} eV") ``` H2O adsorption in this MOF appears to correspond to Case #2 as outlined above. We can now perform re-relaxation of the empty MOF starting from the $r_{\mathrm{MOF+H2O}}$ geometry. ```{code-cell} relax = BFGS(mof_adsorbed_state) relax.run(fmax=0.05) E_mof_rerelax = mof_adsorbed_state.get_potential_energy() print(f"Energy of re-relaxed empty MOF: {E_mof_rerelax:.3f} eV") ``` The MOF returns to its original empty reference energy upon re-relaxation, confirming that this deformation is physically relevant and is induced by the adsorbate molecule. In Case #3, this re-relaxed energy will be more negative (more favorable) than the original empty MOF relaxation. Thus, we take the reference empty MOF energy ($E_{\mathrm{MOF}}$ in Eqn. 1) to be the minimum of the original empty MOF energy and the re-relaxed MOf energy: ```{code-cell} E_mof = min(E_mof_empty, E_mof_rerelax) # get adsorbate reference energy h2o = mof_h2o[-3:] h2o.calc = calc E_h2o = h2o.get_potential_energy() # compute adsorption energy E_ads = E_combo - E_mof - E_h2o print(f"Adsorption energy of H2O in WOBHEB_0.11_0: {E_ads:.3f} eV") ``` This adsorption energy closely matches that from DFT (–0.699 eV) [[1]](https://arxiv.org/abs/2506.09256). The strong adsorption energy is a consequence of both H2O chemisorption and MOF deformation. We can decompose the adsorption energy into contributions from these two factors. Assuming rigid H2O molecules, we define $E_{\mathrm{int}}$ and $E_{\mathrm{MOF,deform}}$, respectively, as $$ E_{\mathrm{int}} = E_{\mathrm{MOF+H2O}}(r_{\mathrm{MOF+H2O}}) - E_{\mathrm{MOF}}(r_{\mathrm{MOF+H2O}}) - E_{\mathrm{H2O}}(r_{\mathrm{MOF+H2O}}) \tag{2}$$ $$ E_{\mathrm{MOF,deform}} = E_{\mathrm{MOF}}(r_{\mathrm{MOF+H2O}}) - E_{\mathrm{MOF}}(r_{\mathrm{MOF}}) \tag{3}$$ $E_{\mathrm{int}}$ describes host host–guest interactions for the MOF in the adsorbed state only. $E_{\mathrm{MOF,deform}}$ quantifies the magnitude of deformation between the MOF in the adsorbed state and the most energetically favorable empty MOF geometry determined from the workflow presented here. It can be shown that $$ E_{\mathrm{ads}} = E_{\mathrm{int}} + E_{\mathrm{MOF,deform}} \tag{4}$$ For H2O adsorption in `WOBHEB_0.11`, we have ```{code-cell} E_int = E_combo - E_mof_adsorbed_state - E_h2o print(f"E_int: {E_int}") ``` ```{code-cell} E_mof_deform = E_mof_adsorbed_state - E_mof_empty print(f"E_mof_deform: {E_mof_deform}") ``` ```{code-cell} E_ads = E_int + E_mof_deform print(f"E_ads: {E_ads}") ``` $E_{\mathrm{int}}$ is equivalent to $E_{\mathrm{ads}}$ when the MOF is assumed to be rigid. In this case, failure to consider adsorbate-induced deformation would result in an overestimation of the adsorption energy magnitude. ## Acknowledgements & Authors Logan Brabson and Sihoon Choi (Georgia Tech) and the OpenDAC project. --- Source: `docs/uma_tutorials/workshop_tutorials.md` # Workshop Tutorials These longer, hands-on tutorials are designed for classrooms, workshops, and guided self-study. They deliberately revisit material available elsewhere in the documentation, but organize it into complete sessions that instructors and learners can follow from beginning to end. For a concise first calculation, use [Hello World](../core/quickstart.md). For task-specific reference material, browse [common workflows](../core/common_tasks/summary.md). ::::{grid} 1 2 2 2 :::{card} General UMA workshop :link: uma_tutorial In a 1–2 hour session, practice molecular energies, surface and bulk relaxations, molecular dynamics, vibrations, phonons, and transition states. ::: :::{card} Catalysis workshop :link: uma_catalysis_tutorial A comprehensive guided session covering bulk optimization, surface energies, Wulff construction, adsorption energies, and reaction barriers. ::: :::: --- Source: `docs/uma_tutorials/uma_tutorial.md` # General UMA Workshop This 1–2 hour tutorial is intended for classrooms, workshops, and guided self-study. It combines several UMA workflows in one linear lesson. The same topics also appear as shorter, task-focused pages elsewhere in the documentation; use [Hello World](../core/quickstart.md) if you only need a concise introduction. :::{note} Learning Objectives By the end of this tutorial, you will be able to: - Set up and configure UMA models with HuggingFace authentication - Use the FAIRChemCalculator with different task names (omol, oc20, oc22, oc25, omat, odac, omc) - Perform molecular energy calculations including spin states - Run adsorbate relaxations on catalyst surfaces - Execute bulk relaxations with cell optimization - Conduct molecular dynamics simulations - Calculate adsorption energies with proper reference energies - Compute molecular vibrations and phonon spectra - Set up and run transition state calculations (NEBs) ::: This tutorial walks through representative ways to use UMA. It is best suited to researchers who are new to UMA and already have some familiarity with ASE and molecular simulation. # Before you start / installation You need to get a HuggingFace account and request access to the UMA models. You need a Huggingface account, request access to https://huggingface.co/facebook/UMA, and to create a Huggingface token at https://huggingface.co/settings/tokens/ with these permission: Permissions: Read access to contents of all public gated repos you can access Then, add the token as an environment variable (using `huggingface-cli login`: ```{code-cell} :tags: [skip-execution] # Enter token via huggingface-cli ! huggingface-cli login ``` or you can set the token via HF_TOKEN variable: ```{code-cell} :tags: [skip-execution] # Set token via env variable import os os.environ['HF_TOKEN'] = 'MYTOKEN' ``` ## Installation process It may be enough to use `pip install fairchem-core`. This gets you the latest version on PyPi (https://pypi.org/project/fairchem-core/) Here we install some sub-packages. This can take 2-5 minutes to run. ```{code-cell} :tags: [skip-execution] ! pip install fairchem-core fairchem-data-oc fairchem-applications-cattsunami x3dase ``` ```{code-cell} # Check that packages are installed !pip list | grep fairchem ``` ```{code-cell} import fairchem.core fairchem.core.__version__ ``` # Illustrative examples These should just run, and are here to show some basic uses. :::{tip} Critical Points When using UMA, remember these key steps: 1. Create a calculator using `pretrained_mlip.get_predict_unit()` 2. Specify the appropriate **task_name** for your system (omol, oc20, oc22, oc25, omat, odac, omc) 3. Use the calculator like any other ASE calculator ::: ## Spin gap energy - OMOL This is the difference in energy between a triplet and single ground state for a CH2 radical. This downloads a ~1GB checkpoint the first time you run it. :::{tip} OMOL Task Requirements For molecular calculations with the `omol` task, you must set spin and charge in `atoms.info`: ```python atoms.info.update({"spin": 1, "charge": 0}) ``` Spin is the multiplicity (1 for singlet, 2 for doublet, 3 for triplet, etc.). ::: We don't set a device here, so we get a warning about using a CPU device. You can ignore that. If a CUDA environment is available, a GPU may be used to speed up the calculations. ```{code-cell} from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") ``` ```{code-cell} from ase.build import molecule # singlet CH2 singlet = molecule("CH2_s1A1d") singlet.info.update({"spin": 1, "charge": 0}) singlet.calc = FAIRChemCalculator(predictor, task_name="omol") # triplet CH2 triplet = molecule("CH2_s3B1d") triplet.info.update({"spin": 3, "charge": 0}) triplet.calc = FAIRChemCalculator(predictor, task_name="omol") print(triplet.get_potential_energy() - singlet.get_potential_energy()) ``` ## Example of adsorbate relaxation - OC20 Here we just setup a Cu(100) slab with a CO on it and relax it. :::{note} This is an OC20 task because it involves a slab with an adsorbate. The OC20 task is trained specifically for heterogeneous catalysis systems and understands the surface-adsorbate interaction. ::: We specify an explicit device in the predictor here, and avoid the warning. ```{code-cell} from ase.build import add_adsorbate, fcc100, molecule from ase.optimize import LBFGS from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") # Set up your system as an ASE atoms object slab = fcc100("Cu", (3, 3, 3), vacuum=8, periodic=True) adsorbate = molecule("CO") add_adsorbate(slab, adsorbate, 2.0, "bridge") slab.calc = calc # Set up LBFGS dynamics object opt = LBFGS(slab) opt.run(0.05, 100) print(slab.get_potential_energy()) ``` # Example bulk relaxation - OMAT :::{tip} OMAT Task for Bulk Materials The `omat` task is trained on bulk inorganic materials and supports stress tensor predictions, enabling cell optimization with filters like `FrechetCellFilter`. This is essential for finding equilibrium lattice constants. ::: ```{code-cell} from ase.build import bulk from ase.filters import FrechetCellFilter from ase.optimize import FIRE from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="omat") atoms = bulk("Fe") atoms.calc = calc opt = FIRE(FrechetCellFilter(atoms)) opt.run(0.05, 100) print(atoms.get_stress()) # !!!! We get stress now! ``` ## Molecular dynamics - OMOL ```{code-cell} import matplotlib.pyplot as plt from ase import units from ase.build import molecule from ase.io import Trajectory from ase.md.langevin import Langevin from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="omol") atoms = molecule("H2O") atoms.info.update(charge=0, spin=1) # For omol atoms.calc = calc dyn = Langevin( atoms, timestep=0.1 * units.fs, temperature_K=400, friction=0.001 / units.fs, ) trajectory = Trajectory("my_md.traj", "w", atoms) dyn.attach(trajectory.write, interval=1) dyn.run(steps=50) # See some results - not paper ready! traj = Trajectory("my_md.traj") plt.plot( [i * 0.1 * units.fs for i in range(len(traj))], [a.get_potential_energy() for a in traj], ) plt.xlabel("Time (fs)") plt.ylabel("Energy (eV)"); ``` # [Catalyst Adsorption energies](../catalysts/examples_tutorials/OCP-introduction) The basic approach in computing an adsorption energy is to compute this energy difference: dH = E_adslab - E_slab - E_ads We use UMA for two of these energies `E_adslab` and `E_slab`. For `E_ads` We have to do something a little different. The OC20 task is not trained for molecules or molecular fragments. We use atomic energy reference energies instead. These are tabulated below. The OC20 reference scheme is this reaction: x CO + (x + y/2 - z) H2 + (z-x) H2O + w/2 N2 + * -> CxHyOzNw* For this example we have -H2 + H2O + * -> O*. "O": -7.204 eV Where `"O": -7.204` is a constant. To get the desired reaction energy we want we add the formation energy of water. We use either DFT or experimental values for this reaction energy. 1/2O2 + H2 -> H2O Alternatives to this approach are using DFT to estimate the energy of 1/2 O2, just make sure to use consistent settings with your task. You should not use OMOL for this. ```{code-cell} from ase.build import add_adsorbate, fcc111 from ase.optimize import BFGS from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") ``` ```{code-cell} # reference energies from a linear combination of H2O/N2/CO/H2! atomic_reference_energies = { "H": -3.477, "N": -8.083, "O": -7.204, "C": -7.282, } re1 = -3.03 # Water formation energy from experiment slab = fcc111("Pt", size=(2, 2, 5), vacuum=20.0) slab.pbc = True adslab = slab.copy() add_adsorbate(adslab, "O", height=1.2, position="fcc") slab.calc = calc opt = BFGS(slab) print("Relaxing slab") opt.run(fmax=0.05, steps=100) slab_e = slab.get_potential_energy() adslab.calc = calc opt = BFGS(adslab) print("\nRelaxing adslab") opt.run(fmax=0.05, steps=100) adslab_e = adslab.get_potential_energy() ``` Now we compute the adsorption energy. ```{code-cell} # Energy for ((H2O-H2) + * -> *O) + (H2 + 1/2O2 -> H2O) leads to 1/2O2 + * -> *O! adslab_e - slab_e - atomic_reference_energies["O"] + re1 ``` How did we do? We need a reference point. In the paper below, there is an atomic adsorption energy for O on Pt(111) of about -4.264 eV. This is for the reaction O + * -> O*. To convert this to the dissociative adsorption energy, we have to add the reaction: 1/2 O2 -> O D = 2.58 eV (expt) to get a comparable energy of about -1.68 eV. There is about ~0.2 eV difference (we predicted -1.47 eV above, and the reference comparison is -1.68 eV) to account for. The biggest difference is likely due to the differences in exchange-correlation functional. The reference data used the PBE functional, and eSCN was trained on RPBE data. To additional places where there are differences include: 1. Difference in lattice constant 2. The reference energy used for the experiment references. These can differ by up to 0.5 eV from comparable DFT calculations. 2. How many layers are relaxed in the calculation Some of these differences tend to be systematic, and you can calibrate and correct these, especially if you can augment these with your own DFT calculations. It is always a good idea to visualize the geometries to make sure they look reasonable. ```{code-cell} import matplotlib.pyplot as plt from ase.visualize.plot import plot_atoms fig, axs = plt.subplots(1, 2) plot_atoms(slab, axs[0]) plot_atoms(slab, axs[1], rotation=("-90x")) axs[0].set_axis_off() axs[1].set_axis_off() ``` ```{code-cell} fig, axs = plt.subplots(1, 2) plot_atoms(adslab, axs[0]) plot_atoms(adslab, axs[1], rotation=("-90x")) axs[0].set_axis_off() axs[1].set_axis_off() ``` # Molecular vibrations ```{code-cell} from ase import Atoms from ase.optimize import BFGS predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="omol") from ase.vibrations import Vibrations n2 = Atoms("N2", [(0, 0, 0), (0, 0, 1.1)]) n2.info.update({"spin": 1, "charge": 0}) n2.calc = calc BFGS(n2).run(fmax=0.01) ``` ```{code-cell} vib = Vibrations(n2) vib.run() vib.summary() ``` # Bulk alloy phase behavior Adapted from https://kitchingroup.cheme.cmu.edu/dft-book/dft.html#orgheadline29 We manually compute the formation energy of pure compounds and some alloy compositions to assess stability. ```{code-cell} from ase.atoms import Atom, Atoms from ase.filters import FrechetCellFilter from ase.optimize import FIRE from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") cu = Atoms( [Atom("Cu", [0.000, 0.000, 0.000])], cell=[[1.818, 0.000, 1.818], [1.818, 1.818, 0.000], [0.000, 1.818, 1.818]], pbc=True, ) cu.calc = FAIRChemCalculator(predictor, task_name="omat") opt = FIRE(FrechetCellFilter(cu)) opt.run(0.05, 100) cu.get_potential_energy() ``` ```{code-cell} pd = Atoms( [Atom("Pd", [0.000, 0.000, 0.000])], cell=[[1.978, 0.000, 1.978], [1.978, 1.978, 0.000], [0.000, 1.978, 1.978]], pbc=True, ) pd.calc = FAIRChemCalculator(predictor, task_name="omat") opt = FIRE(FrechetCellFilter(pd)) opt.run(0.05, 100) pd.get_potential_energy() ``` ## Alloy formation energies ```{code-cell} cupd1 = Atoms( [Atom("Cu", [0.000, 0.000, 0.000]), Atom("Pd", [-1.652, 0.000, 2.039])], cell=[[0.000, -2.039, 2.039], [0.000, 2.039, 2.039], [-3.303, 0.000, 0.000]], pbc=True, ) # Note pbc=True is important, it is not the default and OMAT cupd1.calc = FAIRChemCalculator(predictor, task_name="omat") opt = FIRE(FrechetCellFilter(cupd1)) opt.run(0.05, 100) cupd1.get_potential_energy() ``` ```{code-cell} cupd2 = Atoms( [ Atom("Cu", [-0.049, 0.049, 0.049]), Atom("Cu", [-11.170, 11.170, 11.170]), Atom("Pd", [-7.415, 7.415, 7.415]), Atom("Pd", [-3.804, 3.804, 3.804]), ], cell=[[-5.629, 3.701, 5.629], [-3.701, 5.629, 5.629], [-5.629, 5.629, 3.701]], pbc=True, ) cupd2.calc = FAIRChemCalculator(predictor, task_name="omat") opt = FIRE(FrechetCellFilter(cupd2)) opt.run(0.05, 100) cupd2.get_potential_energy() ``` ```{code-cell} # Delta Hf cupd-1 = -0.11 eV/atom hf1 = ( cupd1.get_potential_energy() - cu.get_potential_energy() - pd.get_potential_energy() ) hf1 ``` ```{code-cell} # DFT: Delta Hf cupd-2 = -0.04 eV/atom hf2 = ( cupd2.get_potential_energy() - 2 * cu.get_potential_energy() - 2 * pd.get_potential_energy() ) hf2 ``` ```{code-cell} hf1 - hf2, (-0.11 - -0.04) ``` These indicate that cupd-1 and cupd-2 are both more stable than phase separated Cu and Pd, and that cupd-1 is more stable than cupd-2. The absolute formation energies differ from the DFT references, but the relative differences are quite close. The absolute differences could be due to DFT parameter choices (XC, psp, etc.). ## Phonon calculation This takes 4-10 minutes. Adapted from https://wiki.fysik.dtu.dk/ase/ase/phonons.html#example. Phonons have applications in computing the stability and free energy of solids. See: 1. https://www.sciencedirect.com/science/article/pii/S1359646215003127 2. https://iopscience.iop.org/book/mono/978-0-7503-2572-1/chapter/bk978-0-7503-2572-1ch1 ```{code-cell} from ase.build import bulk from ase.phonons import Phonons predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="omat") # Setup crystal atoms = bulk("Al", "fcc", a=4.05) # Phonon calculator N = 7 ph = Phonons(atoms, calc, supercell=(N, N, N), delta=0.05) ph.run() # Read forces and assemble the dynamical matrix ph.read(acoustic=True) ph.clean() path = atoms.cell.bandpath("GXULGK", npoints=100) bs = ph.get_band_structure(path) dos = ph.get_dos(kpts=(20, 20, 20)).sample_grid(npts=100, width=1e-3) ``` ```{code-cell} # Plot the band structure and DOS: import matplotlib.pyplot as plt # noqa fig = plt.figure(figsize=(7, 4)) ax = fig.add_axes([0.12, 0.07, 0.67, 0.85]) emax = 0.04 bs.plot(ax=ax, emin=0.0, emax=emax) dosax = fig.add_axes([0.8, 0.07, 0.17, 0.85]) dosax.fill_between( dos.get_weights(), dos.get_energies(), y2=0, color="grey", edgecolor="k", lw=1, ) dosax.set_ylim(0, emax) dosax.set_yticks([]) dosax.set_xticks([]) dosax.set_xlabel("DOS", fontsize=18); ``` # Transition States (NEBs) Nudged elastic band calculations are among the most costly calculations we do. UMA makes these quicker! :::{note} NEB Workflow The standard workflow for transition state calculations: 1. Get and relax the initial state 2. Get and relax the final state 3. Construct band and interpolate the images 4. Relax the band 5. Analyze and plot the band ::: We explore diffusion of an O adatom from an hcp to an fcc site on Pt(111). ## Initial state ```{code-cell} from ase.build import add_adsorbate, fcc111, molecule from ase.optimize import LBFGS from fairchem.core import FAIRChemCalculator, pretrained_mlip predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") calc = FAIRChemCalculator(predictor, task_name="oc20") # Set up your system as an ASE atoms object initial = fcc111("Pt", (3, 3, 3), vacuum=8, periodic=True) adsorbate = molecule("O") add_adsorbate(initial, adsorbate, 2.0, "fcc") initial.calc = calc # Set up LBFGS dynamics object opt = LBFGS(initial) opt.run(0.05, 100) print(initial.get_potential_energy()) ``` ## Final state ```{code-cell} # Set up your system as an ASE atoms object final = fcc111("Pt", (3, 3, 3), vacuum=8, periodic=True) adsorbate = molecule("O") add_adsorbate(final, adsorbate, 2.0, "hcp") final.calc = FAIRChemCalculator(predictor, task_name="oc20") # Set up LBFGS dynamics object opt = LBFGS(final) opt.run(0.05, 100) print(final.get_potential_energy()) ``` ## Setup and relax the band ```{code-cell} from ase.mep import NEB images = [initial] for i in range(3): image = initial.copy() image.calc = FAIRChemCalculator(predictor, task_name="oc20") images.append(image) images.append(final) neb = NEB(images) neb.interpolate() opt = LBFGS(neb, trajectory="neb.traj") opt.run(0.05, 100) ``` ```{code-cell} from ase.mep import NEBTools NEBTools(neb.images).plot_band(); ``` This could be a good initial guess to initialize an NEB in DFT. # Ideas for things you can do with UMA :::{seealso} Further Reading Here are some research papers demonstrating advanced UMA applications: 1. **FineTuna** - Use it for initial geometry optimizations then do DFT - [DOI: 10.1088/2632-2153/ac8fe0](http://doi.org/10.1088/2632-2153/ac8fe0) - [DOI: 10.1088/2632-2153/ad37f0](http://doi.org/10.1088/2632-2153/ad37f0) 2. **AdsorbML** - Prescreen adsorption sites to find relevant ones - [Nature Computational Science (2023)](https://www.nature.com/articles/s41524-023-01121-5) 3. **CatTsunami** - Screen NEBs more thoroughly - [ACS Catalysis (2024)](https://pubs.acs.org/doi/10.1021/acscatal.4c04272) 4. **Free energy estimations** - Compute vibrational modes for entropy - [J. Phys. Chem. C (2024)](https://pubs.acs.org/doi/10.1021/acs.jpcc.4c07477) 5. **Massive screening** - 685M relaxations of catalyst surfaces - [arXiv:2411.11783](https://arxiv.org/abs/2411.11783) ::: # Advanced applications These take a while to run. ## [AdsorbML](../catalysts/examples_tutorials/adsorbml_walkthrough.md) It is so cheap to run these calculations that we can screen a broad range of adsorbate sites and rank them in stability. The AdsorbML approach automates this. This takes quite a while to run here, and we don't do it in the workshop. ## [Expert adsorption energies](../catalysts/examples_tutorials/adsorption_energies/adsorption_energies.md) This tutorial reproduces Fig 6b from the following paper: Zhou, Jing, et al. “Enhanced Catalytic Activity of Bimetallic Ordered Catalysts for Nitrogen Reduction Reaction by Perturbation of Scaling Relations.” ACS Catalysis 134 (2023): 2190-2201 (https://doi.org/10.1021/acscatal.2c05877). This takes up to an hour with a GPU, and much longer with a CPU. ## [CatTsunami](../catalysts/examples_tutorials/cattsunami_tutorial.md) The CatTsunami tutorial is an example of enumerating initial and final states, and computing reaction paths between them with UMA. ## Acknowledgements This tutorial was originally compiled by John Kitchin (CMU) for the NAM29 catalysis tutorial session, using a variety of resources from the FAIR chemistry repository. --- Source: `docs/uma_tutorials/uma_catalysis_tutorial.md` # UMA Catalysis Tutorial Author: Zack Ulissi (Meta, CMU), with help from AI coding agents / LLMs Original paper: Bjarne Kreitz et al. JPCC (2021) ## Overview This tutorial demonstrates how to use the Universal Model for Atoms (UMA) machine learning potential to perform comprehensive catalyst surface analysis. We replicate key computational workflows from ["Microkinetic Modeling of CO₂ Desorption from Supported Multifaceted Ni Catalysts"](https://pubs.acs.org/doi/10.1021/acs.jpcc.0c09985) by Bjarne Kreitz (now faculty at Georgia Tech!), showing how ML potentials can accelerate computational catalysis research. ```{admonition} Learning Objectives :class: note By the end of this tutorial, you will be able to: - Optimize bulk crystal structures and extract lattice constants - Calculate surface energies using linear extrapolation methods - Construct Wulff shapes to predict nanoparticle morphologies - Compute adsorption energies with zero-point energy corrections - Study coverage-dependent binding phenomena - Calculate reaction barriers using the nudged elastic band (NEB) method - Apply D3 dispersion corrections to improve accuracy ``` ```{admonition} About UMA-S-1P2P1 :class: tip The **UMA-S-1P2P1** model is a state-of-the-art universal machine learning potential trained on the OMat24, OC20, OMol25, ODAC23, and OMC25 datasets, covering diverse materials and surface chemistries. It provides roughly a 1000× speedup over DFT while maintaining useful accuracy for screening studies. Here we use `uma-s-1p2p1`, the latest patch release of the small UMA 1.2 model. The checkpoint is open science under a lightweight license that users accept through Hugging Face. You can read more about the UMA models here: https://arxiv.org/abs/2506.23971 ``` ## Installation and Setup This tutorial uses a number of helpful open source packages: - `ase` - Atomic Simulation Environment - `fairchem` - FAIR Chemistry ML potentials (formerly OCP) - `pymatgen` - Materials analysis - `matplotlib` - Visualization - `numpy` - Numerical computing - `torch-dftd` - Dispersion corrections among many others! ### Huggingface setups You need to get a HuggingFace account and request access to the UMA models. You need a Huggingface account, request access to https://huggingface.co/facebook/UMA, and to create a Huggingface token at https://huggingface.co/settings/tokens/ with these permission: Permissions: Read access to contents of all public gated repos you can access Then, add the token as an environment variable using `huggingface-cli login`: ```{code-cell} ipython3 :tags: [skip-execution] # Enter token via huggingface-cli ! huggingface-cli login ``` or you can set the token via HF_TOKEN variable: ```{code-cell} ipython3 :tags: [skip-execution] # Set token via env variable import os os.environ["HF_TOKEN"] = "MYTOKEN" ``` ### FAIR Chemistry (UMA) installation It may be enough to use `pip install fairchem-core`. This gets you the latest version on PyPi (https://pypi.org/project/fairchem-core/) Here we install some sub-packages. This can take 2-5 minutes to run. ```{code-cell} ipython3 :tags: [skip-execution] ! pip install fairchem-core[docs] fairchem-data-oc fairchem-applications-cattsunami x3dase ``` ```{code-cell} ipython3 # Check that packages are installed !pip list | grep fairchem ``` ```{code-cell} ipython3 import fairchem.core fairchem.core.__version__ ``` ## Package imports First, let's import all necessary libraries and initialize the UMA-S-1P2P1 predictor: ```{code-cell} ipython3 from pathlib import Path import ase.io import matplotlib.pyplot as plt import numpy as np from ase import Atoms from ase.build import bulk from ase.constraints import FixBondLengths from ase.io import write from ase.mep import interpolate from ase.mep.dyneb import DyNEB from ase.optimize import FIRE, LBFGS from ase.vibrations import Vibrations from ase.visualize import view from fairchem.core import FAIRChemCalculator, pretrained_mlip from fairchem.data.oc.core import ( Adsorbate, AdsorbateSlabConfig, Bulk, MultipleAdsorbateSlabConfig, Slab, ) from pymatgen.analysis.wulff import WulffShape from pymatgen.core import Lattice, Structure from pymatgen.core.surface import SlabGenerator from pymatgen.io.ase import AseAtomsAdaptor from torch_dftd.torch_dftd3_calculator import TorchDFTD3Calculator # Set up output directory structure output_dir = Path("ni_tutorial_results") output_dir.mkdir(exist_ok=True) # Create subdirectories for each part part_dirs = { "part1": "part1-bulk-optimization", "part2": "part2-surface-energies", "part3": "part3-wulff-construction", "part4": "part4-h-adsorption", "part5": "part5-coverage-dependence", "part6": "part6-co-dissociation", } for key, dirname in part_dirs.items(): (output_dir / dirname).mkdir(exist_ok=True) # Create subdirectories for different facets in part2 for facet in ["111", "100", "110", "211"]: (output_dir / part_dirs["part2"] / f"ni{facet}").mkdir(exist_ok=True) # Initialize the UMA-S-1P2P1 predictor print("\nLoading UMA-S-1P2P1 model...") predictor = pretrained_mlip.get_predict_unit("uma-s-1p2p1") print("✓ Model loaded successfully!") ``` It is somewhat time consuming to run this. We're going to use a small number of bulks for the testing of this documentation, but otherwise run all of the results for the actual documentation. ```{code-cell} ipython3 import os fast_docs = os.environ.get("FAST_DOCS", "false").lower() == "true" if fast_docs: num_sites = 2 relaxation_steps = 20 else: num_sites = 5 relaxation_steps = 300 ``` --- ## Part 1: Bulk Crystal Optimization ### Introduction Before studying surfaces, we need to determine the equilibrium lattice constant of bulk Ni. This is crucial because surface energies and adsorbate binding depend strongly on the underlying lattice parameter. ### Theory For FCC metals like Ni, the lattice constant **a** defines the unit cell size. The experimental value for Ni is **a = 3.524 Å** at room temperature. We'll optimize both atomic positions and the cell volume to find the ML potential's equilibrium structure. ```{code-cell} ipython3 # Create initial FCC Ni structure a_initial = 3.52 # Å, close to experimental ni_bulk = bulk("Ni", "fcc", a=a_initial, cubic=True) print(f"Initial lattice constant: {a_initial:.2f} Å") print(f"Number of atoms: {len(ni_bulk)}") # Set up calculator for bulk optimization calc = FAIRChemCalculator(predictor, task_name="omat") ni_bulk.calc = calc # Use ExpCellFilter to allow cell relaxation from ase.filters import ExpCellFilter ecf = ExpCellFilter(ni_bulk) # Optimize with LBFGS opt = LBFGS( ecf, trajectory=str(output_dir / part_dirs["part1"] / "ni_bulk_opt.traj"), logfile=str(output_dir / part_dirs["part1"] / "ni_bulk_opt.log"), ) opt.run(fmax=0.05, steps=relaxation_steps) # Extract results cell = ni_bulk.get_cell() a_optimized = cell[0, 0] a_exp = 3.524 # Experimental value error = abs(a_optimized - a_exp) / a_exp * 100 print(f"\n{'='*50}") print(f"Experimental lattice constant: {a_exp:.2f} Å") print(f"Optimized lattice constant: {a_optimized:.2f} Å") print(f"Relative error: {error:.2f}%") print(f"{'='*50}") ase.io.write(str(output_dir / part_dirs["part1"] / "ni_bulk_relaxed.cif"), ni_bulk) # Store results for later use a_opt = a_optimized ``` ```{admonition} Missing UMA access? :class: dropdown, tip Don't have access to UMA yet? You can still explore this calculation! [Download example Ni bulk structure](example_configs/ni_bulk.xyz) and test it in the [UMA demo (no login required)](https://facebook-fairchem-uma-demo.hf.space/) to see how the model predicts properties for bulk Ni. ``` ```{admonition} Understanding the Results :class: tip uma-s-1p2p1 using the `omat` task name will predict lattice constants at the PBE level of DFT. For metals, PBE typically predicts lattice constants within 1-2% of experimental values. Small discrepancies arise from: - Training data biases (if your structure is far from OMAT24) - Temperature effects (0 K vs room temperature). You can do a quasi-harmonic analysis to include finite temperature effects if desired. - Quantum effects not captured by the underlying DFT/PBE simulations For surface calculations, using the ML-optimized lattice constant maintains internal consistency. ``` ```{admonition} Comparison with Paper :class: note **Paper (Table 1):** Ni lattice constant = 3.524 Å (experimental reference) The UMA-S-1P2P1 model with OMAT provides excellent agreement with experiment, as would be expected for the PBE functional for simple BCC Ni. The underlying calculations for OMat24 and the original results cited in the paper should be very similar (both PBE), so the fact that the results are a little closer to experiment than the original results is within the numerical noise of the ML model. ``` ```{admonition} Further exploration :class: seealso Try modifying the following parameters and observe the effects: 1. **Task name**: The original paper re-relaxed the structures at the RPBE level of theory before continuing. Try that with the uma-s-1p2p1 model (using the oc20 task name) and see if it matters here. 2. **Initial guess**: Change `a_initial` to 3.0 or 4.0 Å. Does the optimizer still converge to the same value? 3. **Convergence criterion**: Tighten `fmax` to 0.01 eV/Å. How many more steps are required? 4. **Different metals**: Replace `"Ni"` with `"Cu"`, `"Pd"`, or `"Pt"`. Compare predicted vs experimental lattice constants. 5. **Cell shape**: Remove `cubic=True` and allow the cell to distort. Does FCC remain stable? ``` --- ## Part 2: Surface Energy Calculations ### Introduction Surface energy (γ) quantifies the thermodynamic cost of creating a surface. It determines surface stability, morphology, and catalytic activity. We'll calculate γ for four low-index Ni facets: (111), (100), (110), and (211). ### Theory The surface energy is defined as: $$ \gamma = \frac{E_{\text{slab}} - N \cdot E_{\text{bulk}}}{2A} $$ where: - $E_{\text{slab}}$ = total energy of the slab - $N$ = number of atoms in the slab - $E_{\text{bulk}}$ = bulk energy per atom - $A$ = surface area - Factor of 2 accounts for two surfaces (top and bottom) **Challenge**: Direct calculation suffers from quantum size effects, and if you were doing DFT calculations small numerical errors in the simulation or from the K-point grid sampling can lead to small (but significant) errors in the bulk lattice energy. **Solution**: It is fairly common when calculating surface energies to use the bulk energy from a bulk relaxation in the above equation. However, because DFT often has some small numerical noise in the predictions from k-point convergence, this might lead to the wrong surface energy. Instead, two more careful schemes are either: 1. Calculate the energy of a bulk structure oriented to each slab to maximize cancellation of small numerical errors or 2. Calculate the energy of multiple slabs at multiple thicknesses and extrapolate to zero thickness. The intercept will be the surface energy, and the slope will be a fitted bulk energy. A benefit of this approach is that it also forces us to check that we have a sufficiently thick slab for a well defined surface energy; if the fit is non-linear we need thicker slabs. We'll use the linear extrapolation method here as it's more likely to work in future DFT studies if you use this code! ### Step 1: Setup and Bulk Energy Reference First, we'll set up the calculation parameters and get the bulk energy reference: ```{code-cell} ipython3 # Calculate surface energies for all facets facets = [(1, 1, 1), (1, 0, 0), (1, 1, 0), (2, 1, 1)] surface_energies = {} surface_energies_SI = {} all_fit_data = {} # Get bulk energy reference (only need to do this once) E_bulk_total = ni_bulk.get_potential_energy() N_bulk = len(ni_bulk) E_bulk_per_atom = E_bulk_total / N_bulk print(f"Bulk energy reference:") print(f" Total energy: {E_bulk_total:.2f} eV") print(f" Number of atoms: {N_bulk}") print(f" Energy per atom: {E_bulk_per_atom:.6f} eV/atom") ``` ### Step 2: Generate and Relax Slabs Now we'll loop through each facet, generating slabs at three different thicknesses: ```{code-cell} ipython3 # Convert bulk to pymatgen structure for slab generation adaptor = AseAtomsAdaptor() ni_structure = adaptor.get_structure(ni_bulk) for facet in facets: facet_str = "".join(map(str, facet)) print(f"\n{'='*60}") print(f"Calculating Ni({facet_str}) surface energy") print(f"{'='*60}") # Calculate for three thicknesses thicknesses = [4, 6, 8] # layers n_atoms_list = [] energies_list = [] for n_layers in thicknesses: print(f"\n Thickness: {n_layers} layers") # Generate slab slabgen = SlabGenerator( ni_structure, facet, min_slab_size=n_layers * a_opt / np.sqrt(sum([h**2 for h in facet])), min_vacuum_size=10.0, center_slab=True, ) pmg_slab = slabgen.get_slabs()[0] slab = adaptor.get_atoms(pmg_slab) slab.center(vacuum=10.0, axis=2) print(f" Atoms: {len(slab)}") # Relax slab (no constraints - both surfaces free) calc = FAIRChemCalculator(predictor, task_name="omat") slab.calc = calc opt = LBFGS(slab, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_slab = slab.get_potential_energy() n_atoms_list.append(len(slab)) energies_list.append(E_slab) print(f" Energy: {E_slab:.2f} eV") # Linear regression: E_slab = slope * N + intercept coeffs = np.polyfit(n_atoms_list, energies_list, 1) slope = coeffs[0] intercept = coeffs[1] # Extract surface energy from intercept cell = slab.get_cell() area = np.linalg.norm(np.cross(cell[0], cell[1])) gamma = intercept / (2 * area) # eV/Ų gamma_SI = gamma * 16.0218 # J/m² print(f"\n Linear fit:") print(f" Slope: {slope:.6f} eV/atom (cf. bulk {E_bulk_per_atom:.6f})") print(f" Intercept: {intercept:.2f} eV") print(f"\n Surface energy:") print(f" γ = {gamma:.6f} eV/Ų = {gamma_SI:.2f} J/m²") # Store results and fit data surface_energies[facet] = gamma surface_energies_SI[facet] = gamma_SI all_fit_data[facet] = { "n_atoms": n_atoms_list, "energies": energies_list, "slope": slope, "intercept": intercept, } ``` ### Step 3: Visualize Linear Fits Let's visualize the linear extrapolation for all four facets: ```{code-cell} ipython3 # Visualize linear fits for all facets fig, axes = plt.subplots(2, 2, figsize=(12, 10)) axes = axes.flatten() for idx, facet in enumerate(facets): ax = axes[idx] data = all_fit_data[facet] # Plot data points ax.scatter( data["n_atoms"], data["energies"], s=100, color="steelblue", marker="o", zorder=3, label="Calculated", ) # Plot fit line n_range = np.linspace(min(data["n_atoms"]) - 5, max(data["n_atoms"]) + 5, 100) E_fit = data["slope"] * n_range + data["intercept"] ax.plot( n_range, E_fit, "r--", linewidth=2, label=f'Fit: {data["slope"]:.2f}N + {data["intercept"]:.2f}', ) # Formatting facet_str = f"Ni({facet[0]}{facet[1]}{facet[2]})" ax.set_xlabel("Number of Atoms", fontsize=11) ax.set_ylabel("Slab Energy (eV)", fontsize=11) ax.set_title( f"{facet_str}: γ = {surface_energies_SI[facet]:.2f} J/m²", fontsize=12, fontweight="bold", ) ax.legend(fontsize=9) ax.grid(True, alpha=0.3) plt.tight_layout() plt.savefig( str(output_dir / part_dirs["part2"] / "surface_energy_fits.png"), dpi=300, bbox_inches="tight", ) plt.show() ``` ### Step 4: Compare with Literature Finally, let's compare our calculated surface energies with DFT literature values: ```{code-cell} ipython3 print(f"\n{'='*70}") print("Comparison with DFT Literature (Tran et al., 2016)") print(f"{'='*70}") lit_values = { (1, 1, 1): 1.92, (1, 0, 0): 2.21, (1, 1, 0): 2.29, (2, 1, 1): 2.24, } # J/m² for facet in facets: facet_str = f"Ni({facet[0]}{facet[1]}{facet[2]})" calc = surface_energies_SI[facet] lit = lit_values[facet] diff = abs(calc - lit) / lit * 100 print(f"{facet_str:<10} {calc:>8.2f} J/m² (Lit: {lit:.2f}, Δ={diff:.1f}%)") ``` ```{admonition} Missing UMA access? :class: dropdown, tip Don't have access to UMA yet? You can still explore this calculation! [Download example Ni(111) slab structure](example_configs/ni111_slab.xyz) and test it in the [UMA demo (no login required)](https://facebook-fairchem-uma-demo.hf.space/) to see how the model predicts energies for Ni surfaces. ``` ```{admonition} Comparison with Paper (Table 1) :class: note **Paper Results (PBE-DFT, Tran et al.):** - Ni(111): 1.92 J/m² - Ni(100): 2.21 J/m² - Ni(110): 2.29 J/m² - Ni(211): 2.24 J/m² **Key Observations:** 1. **Energy ordering preserved**: (111) < (100) < (110) ≈ (211), matching DFT 2. **Absolute errors**: Typically 10-20%, within expected range for ML potentials 3. **(111) most stable**: Both methods agree this is the lowest energy surface 4. **Physical trend correct**: Close-packed surfaces have lower energy **Why differences exist:** - Training data biases in ML model, which has seen mostly periodic bulk structures, not surfaces. - Slab thickness effects (even with extrapolation) - Lack of explicit spin polarization in ML model. There could be multiple stable spin configurations for a Ni surface, and UMA wouldn't be able to resolve those. **Caveat** Both methods here use PBE as the underlying functional in DFT. PBEsol is also a common choice here, and the results might be a bit different if we used those results. **Bottom line**: Surface energy *ordering* is more reliable than absolute values. Use ML for screening, validate critical cases with DFT. ``` ```{admonition} Why Linear Extrapolation? :class: note Single-thickness slabs suffer from: - **Quantum confinement**: Electronic structure depends on slab thickness - **Surface-surface interactions**: Bottom and top surfaces couple at small thicknesses - **Relaxation artifacts**: Atoms at center may not reach bulk-like coordination Linear extrapolation eliminates these by fitting $E_{\text{slab}}(N)$ and extracting the asymptotic surface energy. ``` ### Explore on Your Own 1. **Thickness convergence**: Add 10 and 12 layer calculations. Is the linear fit still valid? 2. **Constraint effects**: Fix the bottom 2 layers during relaxation. How does this affect γ? 3. **Vacuum size**: Vary `min_vacuum_size` from 8 to 15 Å. When does γ converge? 4. **High-index facets**: Try (311) or (331) surfaces. Are they more or less stable? 5. **Alternative fitting**: Use polynomial (degree 2) instead of linear fit. Does the intercept change? --- ## Part 3: Wulff Construction ### Introduction The **Wulff construction** predicts the equilibrium shape of a crystalline particle by minimizing total surface energy. This determines the morphology of supported catalyst nanoparticles. ### Theory The Wulff theorem states that at equilibrium, the distance from the particle center to a facet is proportional to its surface energy: $$ \frac{h_i}{\gamma_i} = \text{constant} $$ Facets with lower surface energy have larger areas in the equilibrium shape. ### Step 1: Prepare Surface Energies We'll use the surface energies calculated in Part 2 to construct the Wulff shape: ```{code-cell} ipython3 print("\nConstructing Wulff Shape") print("=" * 50) # Use optimized bulk structure adaptor = AseAtomsAdaptor() ni_structure = adaptor.get_structure(ni_bulk) miller_list = list(surface_energies_SI.keys()) energy_list = [surface_energies_SI[m] for m in miller_list] print(f"Using {len(miller_list)} facets:") for miller, energy in zip(miller_list, energy_list): print(f" {miller}: {energy:.2f} J/m²") ``` ### Step 2: Generate Wulff Construction Now we create the Wulff shape and analyze its properties: ```{code-cell} ipython3 # Create Wulff shape wulff = WulffShape(ni_structure.lattice, miller_list, energy_list) # Print properties print(f"\nWulff Shape Properties:") print(f" Volume: {wulff.volume:.2f} ų") print(f" Surface area: {wulff.surface_area:.2f} Ų") print(f" Effective radius: {wulff.effective_radius:.2f} Å") print(f" Weighted γ: {wulff.weighted_surface_energy:.2f} J/m²") # Area fractions print(f"\nFacet Area Fractions:") area_frac = wulff.area_fraction_dict for hkl, frac in sorted(area_frac.items(), key=lambda x: x[1], reverse=True): print(f" {hkl}: {frac*100:.1f}%") ``` ### Step 3: Visualize and Compare Let's visualize the Wulff shape and compare with literature: ```{code-cell} ipython3 # Visualize fig = wulff.get_plot() plt.title("Wulff Construction: Ni Nanoparticle", fontsize=14) plt.tight_layout() plt.savefig( str(output_dir / part_dirs["part3"] / "wulff_shape.png"), dpi=300, bbox_inches="tight", ) plt.show() # Compare with paper print(f"\nComparison with Paper (Table 2):") paper_fractions = {(1, 1, 1): 69.23, (1, 0, 0): 21.10, (1, 1, 0): 5.28, (2, 1, 1): 4.39} for hkl in miller_list: calc_frac = area_frac.get(hkl, 0) * 100 paper_frac = paper_fractions.get(hkl, 0) print(f" {hkl}: {calc_frac:>6.1f}% (Paper: {paper_frac:.1f}%)") ``` ```{admonition} Comparison with Paper (Table 2) :class: note **Paper Results (Wulff Construction):** - Ni(111): 69.23% of surface area - Ni(100): 21.10% - Ni(110): 5.28% - Ni(211): 4.39% **Key Findings:** 1. **(111) dominance**: Both ML and DFT show >65% of surface is (111) facets 2. **Shape prediction**: Truncated octahedron with primarily {111} and {100} faces 3. **Minor facets**: (110) and (211) have small contributions (<10%) 4. **Agreement**: Area fraction ordering matches perfectly with paper **Physical interpretation:** - Real Ni nanoparticles are (111)-terminated octahedra - (100) facets appear at corners/edges as truncations - This morphology is confirmed experimentally by TEM - Explains why (111) surface chemistry dominates catalysis **Impact on catalysis:** - Must study (111) surface for representative results - (100) sites may be important for minority reaction pathways - Edge/corner sites (not captured here) can be highly active ``` ```{admonition} Physical Interpretation :class: note The Wulff shape shows: - **(111) dominance**: Close-packed surface has lowest energy → largest area - **(100) presence**: Moderate energy → significant area fraction - **(110), (211) minor**: Higher energy → small or absent This predicts that Ni nanoparticles will be predominantly {111}-faceted octahedra with {100} truncations, matching experimental observations. ``` ### Explore on Your Own 1. **Particle size effects**: How would including edge/corner energies modify the shape? 2. **Anisotropic strain**: Apply 2% compressive strain to the lattice. How does the shape change? 3. **Temperature effects**: Surface energies decrease with T. Estimate γ(T) and recompute Wulff shape. 4. **Alloy nanoparticles**: Replace some Ni with Cu or Au. How would segregation affect the shape? 5. **Support effects**: Some facets interact more strongly with supports. Model this by reducing their γ. --- ## Part 4: H Adsorption Energy with ZPE Correction ### Introduction Hydrogen adsorption is a fundamental step in many catalytic reactions (hydrogenation, dehydrogenation, etc.). We'll calculate the binding energy with vibrational zero-point energy (ZPE) corrections. ### Theory The adsorption energy is: $$ E_{\text{ads}} = E(\text{slab+H}) - E(\text{slab}) - \frac{1}{2}E(\text{H}_2) $$ ZPE correction accounts for quantum vibrational effects: $$ E_{\text{ads}}^{\text{ZPE}} = E_{\text{ads}} + \text{ZPE}(\text{H}^*) - \frac{1}{2}\text{ZPE}(\text{H}_2) $$ The ZPE correction is calculated by analyzing the vibrational modes of the molecule/adsorbate. ### Step 1: Setup and Relax Clean Slab First, we create the Ni(111) surface and relax it: ```{code-cell} ipython3 # Create Ni(111) slab ni_bulk_atoms = bulk("Ni", "fcc", a=a_opt, cubic=True) ni_bulk_obj = Bulk(bulk_atoms=ni_bulk_atoms) ni_slabs = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj, specific_millers=(1, 1, 1) ) ni_slab = ni_slabs[0].atoms print(f" Created {len(ni_slab)} atom slab") # Set up calculators calc = FAIRChemCalculator(predictor, task_name="oc20") d3_calc = TorchDFTD3Calculator(device="cpu", damping="bj") print(" Calculators initialized (ML + D3)") ``` ### Step 2: Relax Clean Slab Relax the bare Ni(111) surface as our reference: ```{code-cell} ipython3 print("\n1. Relaxing clean Ni(111) slab...") clean_slab = ni_slab.copy() clean_slab.set_pbc([True, True, True]) clean_slab.calc = calc opt = LBFGS( clean_slab, trajectory=str(output_dir / part_dirs["part4"] / "ni111_clean.traj"), logfile=str(output_dir / part_dirs["part4"] / "ni111_clean.log"), ) opt.run(fmax=0.05, steps=relaxation_steps) E_clean_ml = clean_slab.get_potential_energy() clean_slab.calc = d3_calc E_clean_d3 = clean_slab.get_potential_energy() E_clean = E_clean_ml + E_clean_d3 print(f" E(clean): {E_clean:.2f} eV (ML: {E_clean_ml:.2f}, D3: {E_clean_d3:.2f})") # Save clean slab ase.io.write(str(output_dir / part_dirs["part4"] / "ni111_clean.xyz"), clean_slab) print(" ✓ Clean slab relaxed and saved") ``` ### Step 3: Generate H Adsorption Sites Use heuristic placement to generate multiple candidate H adsorption sites: ```{code-cell} ipython3 print("\n2. Generating 5 H adsorption sites...") ni_slab_for_ads = ni_slabs[0] ni_slab_for_ads.atoms = clean_slab.copy() adsorbate_h = Adsorbate(adsorbate_smiles_from_db="*H") ads_slab_config = AdsorbateSlabConfig( ni_slab_for_ads, adsorbate_h, mode="random_site_heuristic_placement", num_sites=num_sites, ) print(f" Generated {len(ads_slab_config.atoms_list)} initial configurations") print(" These include fcc, hcp, bridge, and top sites") ``` ### Step 4: Relax All H Configurations Relax each configuration and identify the most stable site: ```{code-cell} ipython3 print("\n3. Relaxing all H adsorption configurations...") h_energies = [] h_configs = [] h_d3_energies = [] for idx, config in enumerate(ads_slab_config.atoms_list): config_relaxed = config.copy() config_relaxed.set_pbc([True, True, True]) config_relaxed.calc = calc opt = LBFGS( config_relaxed, trajectory=str(output_dir / part_dirs["part4"] / f"h_site_{idx+1}.traj"), logfile=str(output_dir / part_dirs["part4"] / f"h_site_{idx+1}.log"), ) opt.run(fmax=0.05, steps=relaxation_steps) E_ml = config_relaxed.get_potential_energy() config_relaxed.calc = d3_calc E_d3 = config_relaxed.get_potential_energy() E_total = E_ml + E_d3 h_energies.append(E_total) h_configs.append(config_relaxed) h_d3_energies.append(E_d3) print(f" Config {idx+1}: {E_total:.2f} eV (ML: {E_ml:.2f}, D3: {E_d3:.2f})") # Save structure ase.io.write( str(output_dir / part_dirs["part4"] / f"h_site_{idx+1}.xyz"), config_relaxed ) # Select best configuration best_idx = np.argmin(h_energies) slab_with_h = h_configs[best_idx] E_with_h = h_energies[best_idx] E_with_h_d3 = h_d3_energies[best_idx] print(f"\n ✓ Best site: Config {best_idx+1}, E = {E_with_h:.2f} eV") print(f" Energy spread: {max(h_energies) - min(h_energies):.2f} eV") print(f" This spread indicates the importance of testing multiple sites!") ``` ### Step 5: Calculate H₂ Reference Energy We need the H₂ molecule energy as a reference: ```{code-cell} ipython3 print("\n4. Calculating H₂ reference energy...") h2 = Atoms("H2", positions=[[0, 0, 0], [0, 0, 0.74]]) h2.center(vacuum=10.0) h2.set_pbc([True, True, True]) h2.calc = calc opt = LBFGS( h2, trajectory=str(output_dir / part_dirs["part4"] / "h2.traj"), logfile=str(output_dir / part_dirs["part4"] / "h2.log"), ) opt.run(fmax=0.05, steps=relaxation_steps) E_h2_ml = h2.get_potential_energy() h2.calc = d3_calc E_h2_d3 = h2.get_potential_energy() E_h2 = E_h2_ml + E_h2_d3 print(f" E(H₂): {E_h2:.2f} eV (ML: {E_h2_ml:.2f}, D3: {E_h2_d3:.2f})") # Save H2 structure ase.io.write(str(output_dir / part_dirs["part4"] / "h2_optimized.xyz"), h2) ``` ### Step 6: Compute Adsorption Energy Calculate the adsorption energy using the formula: E_ads = E(slab+H) - E(slab) - 0.5×E(H₂) ```{code-cell} ipython3 print(f"\n4. Computing Adsorption Energy:") print(" E_ads = E(slab+H) - E(slab) - 0.5×E(H₂)") E_ads = E_with_h - E_clean - 0.5 * E_h2 E_ads_no_d3 = (E_with_h - E_with_h_d3) - (E_clean - E_clean_d3) - 0.5 * (E_h2 - E_h2_d3) print(f"\n Without D3: {E_ads_no_d3:.2f} eV") print(f" With D3: {E_ads:.2f} eV") print(f" D3 effect: {E_ads - E_ads_no_d3:.2f} eV") print(f"\n → D3 corrections are negligible for H* (small, covalent bonding)") ``` ### Step 7: Zero-Point Energy (ZPE) Corrections Calculate vibrational frequencies to get ZPE corrections: ```{code-cell} ipython3 print("\n6. Computing ZPE corrections...") print(" This accounts for quantum vibrational effects") h_index = len(slab_with_h) - 1 slab_with_h.calc = calc vib = Vibrations(slab_with_h, indices=[h_index], delta=0.02) vib.run() vib_energies = vib.get_energies() zpe_ads = np.sum(vib_energies) / 2.0 h2.calc = calc vib_h2 = Vibrations(h2, indices=[0, 1], delta=0.02) vib_h2.run() vib_energies_h2 = vib_h2.get_energies() zpe_h2 = np.sum(vib_energies_h2) / 2.0 E_ads_zpe = E_ads + zpe_ads - 0.5 * zpe_h2 print(f" ZPE(H*): {zpe_ads:.2f} eV") print(f" ZPE(H₂): {zpe_h2:.2f} eV") print(f" E_ads(ZPE): {E_ads_zpe:.2f} eV") # Visualize vibrational modes from IPython.display import Image, display print("\n Creating animations of vibrational modes...") vib.write_mode(n=0) try: ase.io.write("vib.0.gif", ase.io.read("vib.0.traj@:"), rotation=("-45x,0y,0z")) display(Image(filename="vib.0.gif")) except IndexError: print(" No animation frames were generated for this mode.") vib.clean() vib_h2.clean() ``` ### Step 8: Visualize and Compare Results Visualize the best configuration and compare with literature: ```{code-cell} ipython3 print("\n7. Visualizing best H* configuration...") view(slab_with_h, viewer='x3d') ``` ```{admonition} Missing UMA access? :class: dropdown, tip Don't have access to UMA yet? You can still explore this calculation! [Download example H on Ni(111) structure](example_configs/h_on_ni111.xyz) and test it in the [UMA demo (no login required)](https://facebook-fairchem-uma-demo.hf.space/) to see how the model predicts adsorption properties. ``` ```{code-cell} ipython3 # 6. Compare with literature print(f"\n{'='*60}") print("Comparison with Literature:") print(f"{'='*60}") print("Table 4 (DFT): -0.60 eV (Ni(111), ref H₂)") print(f"This work: {E_ads_zpe:.2f} eV") print(f"Difference: {abs(E_ads_zpe - (-0.60)):.2f} eV") ``` ```{admonition} D3 Dispersion Corrections :class: tip Dispersion (van der Waals) interactions are important for: - Large molecules (CO, CO₂) - Physisorption - Metal-support interfaces For H adsorption, D3 corrections are typically small (<0.1 eV) because H forms strong covalent bonds with the surface. However, always check the magnitude! ``` ### Explore on Your Own 1. **Site preference**: Identify which site (fcc, hcp, bridge, top) the H prefers. Visualize with `view(atoms, viewer='x3d')`. 2. **Coverage effects**: Place 2 H atoms on the slab. How does binding change with separation? 3. **Different facets**: Compare H adsorption on (100) and (110) surfaces. Which is strongest? 4. **Subsurface H**: Place H below the surface layer. Is it stable? 5. **ZPE uncertainty**: How sensitive is E_ads to the vibrational delta parameter (try 0.01, 0.03 Å)? --- ## Part 5: Coverage-Dependent H Adsorption ### Introduction At higher coverages, adsorbate-adsorbate interactions become significant. We'll study how H binding energy changes from dilute (1 atom) to saturated (full monolayer) coverage. ### Theory The differential adsorption energy at coverage θ is: $$ E_{\text{ads}}(\theta) = \frac{E(n\text{H}^*) - E(*) - n \cdot \frac{1}{2}E(\text{H}_2)}{n} $$ For many systems, this varies linearly: $$ E_{\text{ads}}(\theta) = E_{\text{ads}}(0) + \beta \theta $$ where β quantifies lateral interactions (repulsive if β > 0). ### Step 1: Setup Slab and Calculators Create a larger Ni(111) slab to accommodate multiple adsorbates: ```{code-cell} ipython3 # Create large Ni(111) slab ni_bulk_atoms = bulk("Ni", "fcc", a=a_opt, cubic=True) ni_bulk_obj = Bulk(bulk_atoms=ni_bulk_atoms) ni_slabs = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj, specific_millers=(1, 1, 1) ) slab = ni_slabs[0].atoms.copy() print(f" Created {len(slab)} atom slab") # Set up calculators base_calc = FAIRChemCalculator(predictor, task_name="oc20") d3_calc = TorchDFTD3Calculator(device="cpu", damping="bj") print(" ✓ Calculators initialized") ``` ### Step 2: Calculate Reference Energies Get reference energies for clean surface and H₂: ```{code-cell} ipython3 print("\n1. Relaxing clean slab...") clean_slab = slab.copy() clean_slab.pbc = True clean_slab.calc = base_calc opt = LBFGS( clean_slab, trajectory=str(output_dir / part_dirs["part5"] / "ni111_clean.traj"), logfile=str(output_dir / part_dirs["part5"] / "ni111_clean.log"), ) opt.run(fmax=0.05, steps=relaxation_steps) E_clean_ml = clean_slab.get_potential_energy() clean_slab.calc = d3_calc E_clean_d3 = clean_slab.get_potential_energy() E_clean = E_clean_ml + E_clean_d3 print(f" E(clean): {E_clean:.2f} eV") print("\n2. Calculating H₂ reference...") h2 = Atoms("H2", positions=[[0, 0, 0], [0, 0, 0.74]]) h2.center(vacuum=10.0) h2.set_pbc([True, True, True]) h2.calc = base_calc opt = LBFGS( h2, trajectory=str(output_dir / part_dirs["part5"] / "h2.traj"), logfile=str(output_dir / part_dirs["part5"] / "h2.log"), ) opt.run(fmax=0.05, steps=relaxation_steps) E_h2_ml = h2.get_potential_energy() h2.calc = d3_calc E_h2_d3 = h2.get_potential_energy() E_h2 = E_h2_ml + E_h2_d3 print(f" E(H₂): {E_h2:.2f} eV") ``` ### Step 3: Set Up Coverage Study Define the coverages we'll test (from dilute to nearly 1 ML): ```{code-cell} ipython3 # Count surface sites tags = slab.get_tags() n_sites = np.sum(tags == 1) print(f"\n3. Surface sites: {n_sites} (4×4 Ni(111))") # Test coverages: 1 H, 0.25 ML, 0.5 ML, 0.75 ML, 1.0 ML coverages_to_test = [1, 4, 8, 12, 16] print(f"\n Will test coverages: {[f'{n/n_sites:.2f} ML' for n in coverages_to_test]}") print(" This spans from dilute to nearly full monolayer") coverages = [] adsorption_energies = [] ``` ### Step 4: Generate and Relax Configurations at Each Coverage For each coverage, generate multiple configurations and find the lowest energy: ```{code-cell} ipython3 for n_h in coverages_to_test: print(f"\n3. Coverage: {n_h} H ({n_h/n_sites:.2f} ML)") # Generate configurations ni_bulk_obj_h = Bulk(bulk_atoms=ni_bulk_atoms) ni_slabs_h = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj_h, specific_millers=(1, 1, 1) ) slab_for_ads = ni_slabs_h[0] slab_for_ads.atoms = clean_slab.copy() adsorbates_list = [Adsorbate(adsorbate_smiles_from_db="*H") for _ in range(n_h)] try: multi_ads_config = MultipleAdsorbateSlabConfig( slab_for_ads, adsorbates_list, num_configurations=num_sites ) except ValueError as e: print(f" ⚠ Configuration generation failed: {e}") continue if len(multi_ads_config.atoms_list) == 0: print(f" ⚠ No configurations generated") continue print(f" Generated {len(multi_ads_config.atoms_list)} configurations") # Relax each and find best config_energies = [] for idx, config in enumerate(multi_ads_config.atoms_list): config_relaxed = config.copy() config_relaxed.set_pbc([True, True, True]) config_relaxed.calc = base_calc opt = LBFGS(config_relaxed, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_ml = config_relaxed.get_potential_energy() config_relaxed.calc = d3_calc E_d3 = config_relaxed.get_potential_energy() E_total = E_ml + E_d3 config_energies.append(E_total) print(f" Config {idx+1}: {E_total:.2f} eV") best_idx = np.argmin(config_energies) best_energy = config_energies[best_idx] best_config = multi_ads_config.atoms_list[best_idx] E_ads_per_h = (best_energy - E_clean - n_h * 0.5 * E_h2) / n_h coverage = n_h / n_sites coverages.append(coverage) adsorption_energies.append(E_ads_per_h) print(f" → E_ads/H: {E_ads_per_h:.2f} eV") # Visualize best configuration at this coverage print(f" Visualizing configuration with {n_h} H atoms...") view(best_config, viewer='x3d') print(f"\n✓ Completed coverage study: {len(coverages)} data points") ``` ### Step 5: Perform Linear Fit Fit E_ads vs coverage to extract the slope (lateral interaction strength): ```{code-cell} ipython3 print("\n4. Performing linear fit to coverage dependence...") # Linear fit from numpy.polynomial import Polynomial p = Polynomial.fit(coverages, adsorption_energies, 1) slope = p.coef[1] intercept = p.coef[0] print(f"\n{'='*60}") print(f"Linear Fit: E_ads = {intercept:.2f} + {slope:.2f}θ (eV)") print(f"Slope: {slope * 96.485:.1f} kJ/mol per ML") print(f"Paper: 8.7 kJ/mol per ML") print(f"{'='*60}") ``` ### Step 6: Visualize Coverage Dependence Create a plot showing how adsorption energy changes with coverage: ```{code-cell} ipython3 print("\n5. Plotting coverage dependence...") # Plot fig, ax = plt.subplots(figsize=(8, 6)) ax.scatter( coverages, adsorption_energies, s=100, marker="o", label="Calculated", zorder=3, color="steelblue", ) cov_fit = np.linspace(0, max(coverages), 100) ads_fit = p(cov_fit) ax.plot( cov_fit, ads_fit, "r--", label=f"Fit: {intercept:.2f} + {slope:.2f}θ", linewidth=2 ) ax.set_xlabel("H Coverage (ML)", fontsize=12) ax.set_ylabel("Adsorption Energy (eV/H)", fontsize=12) ax.set_title("Coverage-Dependent H Adsorption on Ni(111)", fontsize=14) ax.legend(fontsize=11) ax.grid(True, alpha=0.3) plt.tight_layout() plt.savefig(str(output_dir / part_dirs["part5"] / "coverage_dependence.png"), dpi=300) plt.show() print("\n✓ Coverage dependence analysis complete!") ``` ```{admonition} Missing UMA access? :class: dropdown, tip Don't have access to UMA yet? You can still explore this calculation! [Download example multiple H on Ni(111) structure](example_configs/4h_on_ni111.xyz) and test it in the [UMA demo (no login required)](https://facebook-fairchem-uma-demo.hf.space/) to see how the model handles coverage-dependent binding. ``` ```{admonition} Comparison with Paper :class: note **Expected Results from Paper:** - **Slope**: 8.7 kJ/mol per ML (indicating repulsive lateral H-H interactions) - **Physical interpretation**: H atoms repel weakly due to electrostatic and Pauli effects **What to Check:** - Your fitted slope should be close to 8.7 kJ/mol per ML - The relationship should be approximately linear for θ < 1 ML - Intercept (E_ads at θ → 0) should match the single-H result from Part 4 (~-0.60 eV) **Typical Variations:** - Slope can vary by ±2-3 kJ/mol depending on slab size and configuration sampling - Non-linearity may appear at very high coverage (θ > 0.75 ML) - Model differences can affect lateral interactions more than adsorption energies ``` ```{admonition} Physical Insights :class: note **Positive slope** (repulsive interactions): - Electrostatic: H atoms accumulate negative charge from Ni - Pauli repulsion: Overlapping electron clouds - Strain: Lattice distortions propagate **Magnitude**: - Weak (~10 kJ/mol/ML) → isolated adsorbates - Strong (>50 kJ/mol/ML) → clustering or phase separation likely The paper reports 8.7 kJ/mol/ML, indicating relatively weak lateral interactions for H on Ni(111). ``` ### Explore on Your Own 1. **Non-linear behavior**: Use polynomial (degree 2) fit. Is there curvature at high coverage? 2. **Temperature effects**: Estimate configurational entropy at each coverage. How does this affect free energy? 3. **Pattern formation**: Visualize the lowest-energy configuration at 0.5 ML. Are H atoms ordered? 4. **Other adsorbates**: Repeat for O or N. How do lateral interactions compare? 5. **Phase diagrams**: At what coverage do you expect phase separation (islands vs uniform)? --- ## Part 6: CO Formation/Dissociation Thermochemistry and Barrier ### Introduction CO dissociation (CO* → C* + O*) is the rate-limiting step in many catalytic processes (Fischer-Tropsch, CO oxidation, etc.). We'll calculate the reaction energy for C* + O* → CO* and the activation barriers in both directions using the nudged elastic band (NEB) method. ### Theory * **Forward Reaction**: C* + O* → CO* + * (recombination) * **Reverse Reaction**: CO* + *→ C* + O* (dissociation) * **Thermochemistry**: $\Delta E_{\text{rxn}} = E(\text{C}^* + \text{O}^*) - E(\text{CO}^*)$ * **Barrier**: NEB finds the minimum energy path (MEP) and transition state: $E_a = E^{\ddagger} - E_{\text{initial}}$ ### Step 1: Setup Slab and Calculators Initialize the Ni(111) surface and calculators: ```{code-cell} ipython3 # Create slab ni_bulk_atoms = bulk("Ni", "fcc", a=a_opt, cubic=True) ni_bulk_obj = Bulk(bulk_atoms=ni_bulk_atoms) ni_slabs = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj, specific_millers=(1, 1, 1) ) slab = ni_slabs[0].atoms print(f" Created {len(slab)} atom slab") base_calc = FAIRChemCalculator(predictor, task_name="oc20") d3_calc = TorchDFTD3Calculator(device="cpu", damping="bj") print(" \u2713 Calculators initialized") ``` ### Step 2: Generate and Relax Final State (CO*) Find the most stable CO adsorption configuration (this is the product of C+O recombination): ```{code-cell} ipython3 print("\n1. Final State: CO* on Ni(111)") print(" Generating CO adsorption configurations...") ni_bulk_obj_co = Bulk(bulk_atoms=ni_bulk_atoms) ni_slab_co = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj_co, specific_millers=(1, 1, 1) )[0] ni_slab_co.atoms = slab.copy() adsorbate_co = Adsorbate(adsorbate_smiles_from_db="*CO") multi_ads_config_co = MultipleAdsorbateSlabConfig( ni_slab_co, [adsorbate_co], num_configurations=num_sites ) print(f" Generated {len(multi_ads_config_co.atoms_list)} configurations") # Relax and find best co_energies = [] co_energies_ml = [] co_energies_d3 = [] co_configs = [] for idx, config in enumerate(multi_ads_config_co.atoms_list): config_relaxed = config.copy() config_relaxed.set_pbc([True, True, True]) config_relaxed.calc = base_calc opt = LBFGS(config_relaxed, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_ml = config_relaxed.get_potential_energy() config_relaxed.calc = d3_calc E_d3 = config_relaxed.get_potential_energy() E_total = E_ml + E_d3 co_energies.append(E_total) co_energies_ml.append(E_ml) co_energies_d3.append(E_d3) co_configs.append(config_relaxed) print( f" Config {idx+1}: E_total = {E_total:.2f} eV (RPBE: {E_ml:.2f}, D3: {E_d3:.2f})" ) best_co_idx = np.argmin(co_energies) final_co = co_configs[best_co_idx] E_final_co = co_energies[best_co_idx] E_final_co_ml = co_energies_ml[best_co_idx] E_final_co_d3 = co_energies_d3[best_co_idx] print(f"\n → Best CO* (Config {best_co_idx+1}):") print(f" RPBE: {E_final_co_ml:.2f} eV") print(f" D3: {E_final_co_d3:.2f} eV") print(f" Total: {E_final_co:.2f} eV") # Save best CO state ase.io.write(str(output_dir / part_dirs["part6"] / "co_final_best.traj"), final_co) print(" ✓ Best CO* structure saved") # Visualize best CO* structure print("\n Visualizing best CO* structure...") view(final_co, viewer='x3d') ``` ### Step 3: Generate and Relax Initial State (C* + O*) Find the most stable configuration for dissociated C and O (reactants): ```{code-cell} ipython3 print("\n2. Initial State: C* + O* on Ni(111)") print(" Generating C+O configurations...") ni_bulk_obj_c_o = Bulk(bulk_atoms=ni_bulk_atoms) ni_slab_c_o = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj_c_o, specific_millers=(1, 1, 1) )[0] adsorbate_c = Adsorbate(adsorbate_smiles_from_db="*C") adsorbate_o = Adsorbate(adsorbate_smiles_from_db="*O") multi_ads_config_c_o = MultipleAdsorbateSlabConfig( ni_slab_c_o, [adsorbate_c, adsorbate_o], num_configurations=num_sites ) print(f" Generated {len(multi_ads_config_c_o.atoms_list)} configurations") c_o_energies = [] c_o_energies_ml = [] c_o_energies_d3 = [] c_o_configs = [] for idx, config in enumerate(multi_ads_config_c_o.atoms_list): config_relaxed = config.copy() config_relaxed.set_pbc([True, True, True]) config_relaxed.calc = base_calc opt = LBFGS(config_relaxed, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) # Check C-O bond distance to ensure they haven't formed CO molecule c_o_dist = config_relaxed[config_relaxed.get_tags() == 2].get_distance( 0, 1, mic=True ) # CO bond length is ~1.15 Å, so if distance < 1.5 Å, they've formed a molecule if c_o_dist < 1.5: print( f" Config {idx+1}: ⚠ REJECTED - C and O formed CO molecule (d = {c_o_dist:.2f} Å)" ) continue E_ml = config_relaxed.get_potential_energy() config_relaxed.calc = d3_calc E_d3 = config_relaxed.get_potential_energy() E_total = E_ml + E_d3 c_o_energies.append(E_total) c_o_energies_ml.append(E_ml) c_o_energies_d3.append(E_d3) c_o_configs.append(config_relaxed) print( f" Config {idx+1}: E_total = {E_total:.2f} eV (RPBE: {E_ml:.2f}, D3: {E_d3:.2f}, C-O dist: {c_o_dist:.2f} Å)" ) best_c_o_idx = np.argmin(c_o_energies) initial_c_o = c_o_configs[best_c_o_idx] E_initial_c_o = c_o_energies[best_c_o_idx] E_initial_c_o_ml = c_o_energies_ml[best_c_o_idx] E_initial_c_o_d3 = c_o_energies_d3[best_c_o_idx] print(f"\n → Best C*+O* (Config {best_c_o_idx+1}):") print(f" RPBE: {E_initial_c_o_ml:.2f} eV") print(f" D3: {E_initial_c_o_d3:.2f} eV") print(f" Total: {E_initial_c_o:.2f} eV") # Save best C+O state ase.io.write(str(output_dir / part_dirs["part6"] / "co_initial_best.traj"), initial_c_o) print(" ✓ Best C*+O* structure saved") # Visualize best C*+O* structure print("\n Visualizing best C*+O* structure...") view(initial_c_o, viewer='x3d') ``` ### Step 3b: Calculate C* and O* Energies Separately Another strategy to calculate the initial energies for *C and *O at very low coverage (without interactions between the two reactants) is to do two separate relaxations. ```{code-cell} ipython3 # Clean slab ni_bulk_obj = Bulk(bulk_atoms=ni_bulk_atoms) clean_slab = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj_c_o, specific_millers=(1, 1, 1) )[0].atoms clean_slab.set_pbc([True, True, True]) clean_slab.calc = base_calc opt = LBFGS(clean_slab, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_clean_ml = clean_slab.get_potential_energy() clean_slab.calc = d3_calc E_clean_d3 = clean_slab.get_potential_energy() E_clean = E_clean_ml + E_clean_d3 print( f"\n Clean slab: E_total = {E_clean:.2f} eV (RPBE: {E_clean_ml:.2f}, D3: {E_clean_d3:.2f})" ) ``` ```{code-cell} ipython3 print(f"\n2b. Separate C* and O* Energies:") print(" Calculating energies in separate unit cells to avoid interactions") ni_bulk_obj_c_o = Bulk(bulk_atoms=ni_bulk_atoms) ni_slab_c_o = Slab.from_bulk_get_specific_millers( bulk=ni_bulk_obj_c_o, specific_millers=(1, 1, 1) )[0] print("\n Generating C* configurations...") multi_ads_config_c = MultipleAdsorbateSlabConfig( ni_slab_c_o, adsorbates=[Adsorbate(adsorbate_smiles_from_db="*C")], num_configurations=num_sites, ) c_energies = [] c_energies_ml = [] c_energies_d3 = [] c_configs = [] for idx, config in enumerate(multi_ads_config_c.atoms_list): config_relaxed = config.copy() config_relaxed.set_pbc([True, True, True]) config_relaxed.calc = base_calc opt = LBFGS(config_relaxed, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_ml = config_relaxed.get_potential_energy() config_relaxed.calc = d3_calc E_d3 = config_relaxed.get_potential_energy() E_total = E_ml + E_d3 c_energies.append(E_total) c_energies_ml.append(E_ml) c_energies_d3.append(E_d3) c_configs.append(config_relaxed) print( f" Config {idx+1}: E_total = {E_total:.2f} eV (RPBE: {E_ml:.2f}, D3: {E_d3:.2f})" ) best_c_idx = np.argmin(c_energies) c_ads = c_configs[best_c_idx] E_c = c_energies[best_c_idx] E_c_ml = c_energies_ml[best_c_idx] E_c_d3 = c_energies_d3[best_c_idx] print(f"\n → Best C* (Config {best_c_idx+1}):") print(f" RPBE: {E_c_ml:.2f} eV") print(f" D3: {E_c_d3:.2f} eV") print(f" Total: {E_c:.2f} eV") # Save best C state ase.io.write(str(output_dir / part_dirs["part6"] / "c_best.traj"), c_ads) # Visualize best C* structure print("\n Visualizing best C* structure...") view(c_ads, viewer='x3d') # Generate O* configuration print("\n Generating O* configurations...") multi_ads_config_o = MultipleAdsorbateSlabConfig( ni_slab_c_o, adsorbates=[Adsorbate(adsorbate_smiles_from_db="*O")], num_configurations=num_sites, ) o_energies = [] o_energies_ml = [] o_energies_d3 = [] o_configs = [] for idx, config in enumerate(multi_ads_config_o.atoms_list): config_relaxed = config.copy() config_relaxed.set_pbc([True, True, True]) config_relaxed.calc = base_calc opt = LBFGS(config_relaxed, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_ml = config_relaxed.get_potential_energy() config_relaxed.calc = d3_calc E_d3 = config_relaxed.get_potential_energy() E_total = E_ml + E_d3 o_energies.append(E_total) o_energies_ml.append(E_ml) o_energies_d3.append(E_d3) o_configs.append(config_relaxed) print( f" Config {idx+1}: E_total = {E_total:.2f} eV (RPBE: {E_ml:.2f}, D3: {E_d3:.2f})" ) best_o_idx = np.argmin(o_energies) o_ads = o_configs[best_o_idx] E_o = o_energies[best_o_idx] E_o_ml = o_energies_ml[best_o_idx] E_o_d3 = o_energies_d3[best_o_idx] print(f"\n → Best O* (Config {best_o_idx+1}):") print(f" RPBE: {E_o_ml:.2f} eV") print(f" D3: {E_o_d3:.2f} eV") print(f" Total: {E_o:.2f} eV") # Save best O state ase.io.write(str(output_dir / part_dirs["part6"] / "o_best.traj"), o_ads) # Visualize best O* structure print("\n Visualizing best O* structure...") view(o_ads, viewer='x3d') # Calculate combined energy for separate C* and O* E_initial_c_o_separate = E_c + E_o E_initial_c_o_separate_ml = E_c_ml + E_o_ml E_initial_c_o_separate_d3 = E_c_d3 + E_o_d3 ``` ```{code-cell} ipython3 print(f"\n Combined C* + O* (separate calculations):") print(f" RPBE: {E_initial_c_o_separate_ml:.2f} eV") print(f" D3: {E_initial_c_o_separate_d3:.2f} eV") print(f" Total: {E_initial_c_o_separate:.2f} eV") print(f"\n Comparison:") print(f" C*+O* (same cell): {E_initial_c_o - E_clean:.2f} eV") print(f" C* + O* (separate): {E_initial_c_o_separate - 2*E_clean:.2f} eV") print( f" Difference: {(E_initial_c_o - E_clean) - (E_initial_c_o_separate - 2*E_clean):.2f} eV" ) print(" ✓ Separate C* and O* energies calculated") ``` ### Step 4: Calculate Reaction Energy with ZPE Compute the thermochemistry for C* + O* → CO* with ZPE corrections: ```{code-cell} ipython3 print(f"\n3. Reaction Energy (C* + O* → CO*):") print(f" " + "=" * 60) # Electronic energies print(f"\n Electronic Energies:") print( f" Initial (C*+O*): RPBE = {E_initial_c_o_ml:.2f} eV, D3 = {E_initial_c_o_d3:.2f} eV, Total = {E_initial_c_o:.2f} eV" ) print( f" Final (CO*): RPBE = {E_final_co_ml:.2f} eV, D3 = {E_final_co_d3:.2f} eV, Total = {E_final_co:.2f} eV" ) # Reaction energies without ZPE delta_E_rpbe = E_final_co_ml - E_initial_c_o_ml delta_E_d3_contrib = E_final_co_d3 - E_initial_c_o_d3 delta_E_elec = E_final_co - E_initial_c_o print(f"\n Reaction Energies (without ZPE):") print(f" ΔE(RPBE only): {delta_E_rpbe:.2f} eV = {delta_E_rpbe*96.485:.1f} kJ/mol") print( f" ΔE(D3 contrib): {delta_E_d3_contrib:.2f} eV = {delta_E_d3_contrib*96.485:.1f} kJ/mol" ) print(f" ΔE(RPBE+D3): {delta_E_elec:.2f} eV = {delta_E_elec*96.485:.1f} kJ/mol") # Calculate ZPE for CO* (final state) print(f"\n Computing ZPE for CO*...") final_co.calc = base_calc co_indices = np.where(final_co.get_tags() == 2)[0] vib_co = Vibrations(final_co, indices=co_indices, delta=0.02, name="vib_co") vib_co.run() vib_energies_co = vib_co.get_energies() zpe_co = np.sum(vib_energies_co[vib_energies_co > 0]) / 2.0 vib_co.clean() print(f" ZPE(CO*): {zpe_co:.2f} eV ({zpe_co*1000:.1f} meV)") # Calculate ZPE for C* and O* (initial state) print(f"\n Computing ZPE for C* and O*...") initial_c_o.calc = base_calc c_o_indices = np.where(initial_c_o.get_tags() == 2)[0] vib_c_o = Vibrations(initial_c_o, indices=c_o_indices, delta=0.02, name="vib_c_o") vib_c_o.run() vib_energies_c_o = vib_c_o.get_energies() zpe_c_o = np.sum(vib_energies_c_o[vib_energies_c_o > 0]) / 2.0 vib_c_o.clean() print(f" ZPE(C*+O*): {zpe_c_o:.2f} eV ({zpe_c_o*1000:.1f} meV)") # Total reaction energy with ZPE delta_zpe = zpe_co - zpe_c_o delta_E_zpe = delta_E_elec + delta_zpe print(f"\n Reaction Energy (with ZPE):") print(f" ΔE(electronic): {delta_E_elec:.2f} eV = {delta_E_elec*96.485:.1f} kJ/mol") print( f" ΔZPE: {delta_zpe:.2f} eV = {delta_zpe*96.485:.1f} kJ/mol ({delta_zpe*1000:.1f} meV)" ) print(f" ΔE(total): {delta_E_zpe:.2f} eV = {delta_E_zpe*96.485:.1f} kJ/mol") print(f"\n Summary:") print( f" Without D3, without ZPE: {delta_E_rpbe:.2f} eV = {delta_E_rpbe*96.485:.1f} kJ/mol" ) print( f" With D3, without ZPE: {delta_E_elec:.2f} eV = {delta_E_elec*96.485:.1f} kJ/mol" ) print( f" With D3, with ZPE: {delta_E_zpe:.2f} eV = {delta_E_zpe*96.485:.1f} kJ/mol" ) print(f"\n " + "=" * 60) print(f"\n Comparison with Paper (Table 5):") print(f" Paper (DFT-D3): -142.7 kJ/mol = -1.48 eV") print(f" This work: {delta_E_zpe*96.485:.1f} kJ/mol = {delta_E_zpe:.2f} eV") print(f" Difference: {abs(delta_E_zpe - (-1.48)):.2f} eV") if delta_E_zpe < 0: print(f"\n ✓ Reaction is exothermic (C+O recombination favorable)") else: print(f"\n ⚠ Reaction is endothermic (dissociation favorable)") ``` ### Step 5: Calculate CO Adsorption Energy (Bonus) Calculate how strongly CO binds to the surface: ```{code-cell} ipython3 print(f"\n4. CO Adsorption Energy ( CO(g) + * → CO*):") print(" This helps us understand CO binding strength") # CO(g) co_gas = Atoms("CO", positions=[[0, 0, 0], [0, 0, 1.15]]) co_gas.center(vacuum=10.0) co_gas.set_pbc([True, True, True]) co_gas.calc = base_calc opt = LBFGS(co_gas, logfile=None) opt.run(fmax=0.05, steps=relaxation_steps) E_co_gas_ml = co_gas.get_potential_energy() co_gas.calc = d3_calc E_co_gas_d3 = co_gas.get_potential_energy() E_co_gas = E_co_gas_ml + E_co_gas_d3 print( f" CO(g): E_total = {E_co_gas:.2f} eV (RPBE: {E_co_gas_ml:.2f}, D3: {E_co_gas_d3:.2f})" ) # Calculate ZPE for CO(g) co_gas.calc = base_calc vib_co_gas = Vibrations(co_gas, indices=[0, 1], delta=0.01, nfree=2) vib_co_gas.clean() vib_co_gas.run() vib_energies_co_gas = vib_co_gas.get_energies() zpe_co_gas = 0.5 * np.sum(vib_energies_co_gas[vib_energies_co_gas > 0]) vib_co_gas.clean() print(f" ZPE(CO(g)): {zpe_co_gas:.2f} eV") print(f" ZPE(CO*): {zpe_co:.2f} eV (from Step 4 calculation)") # Electronic adsorption energy E_ads_co_elec = E_final_co - E_clean - E_co_gas # ZPE contribution to adsorption energy delta_zpe_ads = zpe_co - zpe_co_gas # Total adsorption energy with ZPE E_ads_co_total = E_ads_co_elec + delta_zpe_ads print(f"\n Electronic Energy Breakdown:") print(f" ΔE(RPBE only) = {(E_final_co_ml - E_clean_ml - E_co_gas_ml):.2f} eV") print(f" ΔE(D3 contrib) = {((E_final_co_d3 - E_clean_d3 - E_co_gas_d3)):.2f} eV") print(f" ΔE(RPBE+D3) = {E_ads_co_elec:.2f} eV") print(f"\n ZPE Contribution:") print(f" ΔZPE = {delta_zpe_ads:.2f} eV") print(f"\n Total Adsorption Energy:") print(f" ΔE(total) = {E_ads_co_total:.2f} eV = {E_ads_co_total*96.485:.1f} kJ/mol") print(f"\n Summary:") print( f" E_ads(CO) without ZPE = {-E_ads_co_elec:.2f} eV = {-E_ads_co_elec*96.485:.1f} kJ/mol" ) print( f" E_ads(CO) with ZPE = {-E_ads_co_total:.2f} eV = {-E_ads_co_total*96.485:.1f} kJ/mol" ) print( f" → CO binds {abs(E_ads_co_total):.2f} eV stronger than H ({abs(E_ads_co_total)/0.60:.1f}x)" ) ``` ```{admonition} Comparison with Paper Results :class: tip The paper reports a CO adsorption energy of **1.82 eV (175.6 kJ/mol)** in Table 4, calculated using DFT (RPBE functional). These results show: - **Without ZPE**: The electronic binding energy matches well with DFT predictions - **With ZPE**: The zero-point energy correction reduces the binding strength slightly - **D3 Dispersion**: Contributes to stronger binding due to van der Waals interactions ``` ### Step 6: Find guesses for nearby initial and final states for the reaction Now that we have an estimate on the reaction energy from the best possible initial and final states, we want to find a transition state (barrier) for this reaction. There are MANY possible ways that we could do this. In this case, we'll start with the *CO final state and then try and find a nearby local minimal of *C and *O, by fixing the C-O bond distance and finding a nearby local minima. Note that this approach required some insight into what the transition state might look like, and could be considerably more complicated for a reaction that did not involve breaking a single bond. ```{code-cell} ipython3 print(f"\nFinding Transition State Initial and Final States") print(" Creating initial guess with stretched C-O bond...") print(" Starting from CO* and stretching the C-O bond...") # Create a guess structure with stretched CO bond (start from CO*) initial_guess = final_co.copy() # Set up a constraint to fix the bond length to ~2 Angstroms, which should be far enough that we'll be closer to *C+*O than *CO co_indices = np.where(initial_guess.get_tags() == 2)[0] # Rotate the atoms a bit just to break the symmetry and prevent the O from going straight up to satisfy the constraint initial_slab = initial_guess[initial_guess.get_tags() != 2] initial_co = initial_guess[initial_guess.get_tags() == 2] initial_co.rotate(30, "x", center=initial_co.positions[0]) initial_guess = initial_slab + initial_co initial_guess.calc = FAIRChemCalculator(predictor, task_name="oc20") # Add constraints to keep the CO bond length extended initial_guess.constraints += [ FixBondLengths([co_indices], tolerance=1e-2, iterations=5000, bondlengths=[2.0]) ] try: opt = LBFGS( initial_guess, trajectory=output_dir / part_dirs["part6"] / "initial_guess_with_constraint.traj", ) opt.run(fmax=0.01) except RuntimeError: # The FixBondLength constraint is sometimes a little finicky, # but it's ok if it doesn't finish as it's just an initial guess # for the next step pass # Now that we have a guess, re-relax without the constraints initial_guess.constraints = initial_guess.constraints[:-1] opt = LBFGS( initial_guess, trajectory=output_dir / part_dirs["part6"] / "initial_guess_without_constraint.traj", ) opt.run(fmax=0.01) ``` ### Step 7: Run NEB to Find Activation Barrier Use the nudged elastic band method to find the minimum energy path: ```{code-cell} ipython3 print(f"\n7. NEB Barrier Calculation (C* + O* → CO*)") print(" Setting up 7-image NEB chain with TS guess in middle...") print(" Reaction: C* + O* (initial) → TS → CO* (final)") initial = initial_guess.copy() initial.calc = FAIRChemCalculator(predictor, task_name="oc20") images = [initial] # Start with C* + O* n_images = 10 for i in range(n_images): image = initial.copy() image.calc = FAIRChemCalculator(predictor, task_name="oc20") images.append(image) final = final_co.copy() final.calc = FAIRChemCalculator(predictor, task_name="oc20") images.append(final) # End with CO* # Interpolate with better initial guess dyneb = DyNEB(images, climb=True, fmax=0.05) # Interpolate first half (C*+O* → TS) print("\n Interpolating images...") dyneb.interpolate("idpp", mic=True) # Optimize print(" Optimizing NEB path (this may take a while)...") opt = FIRE( dyneb, trajectory=str(output_dir / part_dirs["part6"] / "neb.traj"), logfile=str(output_dir / part_dirs["part6"] / "neb.log"), ) opt.run(fmax=0.1, steps=relaxation_steps) # Extract barrier (from C*+O* to TS) energies = [img.get_potential_energy() for img in images] energies_rel = np.array(energies) - energies[0] E_barrier = np.max(energies_rel) print(f"\n ✓ NEB converged!") print( f"\n Forward barrier (C*+O* → CO*): {E_barrier:.2f} eV = {E_barrier*96.485:.1f} kJ/mol" ) print( f" Reverse barrier (CO* → C*+O*): {E_barrier - energies_rel[-1]:.2f} eV = {(E_barrier- energies_rel[-1])*96.485:.1f} kJ/mol" ) print(f"\n Paper (Table 5): 153 kJ/mol = 1.59 eV ") print(f" Difference: {abs(E_barrier - 1.59):.2f} eV") ``` ### Step 8: Visualize NEB Path and Key Structures Create plots showing the reaction pathway: ```{code-cell} ipython3 print("\n Creating NEB visualization...") # Plot NEB path fig, ax = plt.subplots(figsize=(10, 6)) ax.plot( range(len(energies_rel)), energies_rel, "o-", linewidth=2, markersize=10, color="steelblue", label="NEB Path", ) ax.axhline(0, color="green", linestyle="--", alpha=0.5, label="Initial: C*+O*") ax.axhline(delta_E_zpe, color="red", linestyle="--", alpha=0.5, label="Final: CO*") ax.axhline( E_barrier, color="orange", linestyle=":", alpha=0.7, linewidth=2, label=f"Forward Barrier = {E_barrier:.2f} eV", ) # Annotate transition state ts_idx = np.argmax(energies_rel) ax.annotate( f"TS\n{energies_rel[ts_idx]:.2f} eV", xy=(ts_idx, energies_rel[ts_idx]), xytext=(ts_idx, energies_rel[ts_idx] + 0.3), ha="center", fontsize=11, fontweight="bold", arrowprops=dict(arrowstyle="->", lw=1.5, color="red"), ) ax.set_xlabel("Image Number", fontsize=13) ax.set_ylabel("Relative Energy (eV)", fontsize=13) ax.set_title( "CO Formation on Ni(111): C* + O* → CO* - NEB Path", fontsize=15, fontweight="bold" ) ax.legend(fontsize=11, loc="upper left") ax.grid(True, alpha=0.3) plt.tight_layout() plt.savefig( str(output_dir / part_dirs["part6"] / "neb_path.png"), dpi=300, bbox_inches="tight" ) plt.show() # Create animation of NEB path print("\n Creating NEB path animation...") from ase.io import write as ase_write ase.io.write( str(output_dir / part_dirs["part6"] / "neb_path.gif"), images, format="gif" ) print(" → Saved as neb_path.gif") # Visualize key structures print("\n Visualizing initial state (C* + O*)...") view(initial_c_o, viewer='x3d') print("\n Visualizing transition state...") view(images[ts_idx], viewer='x3d') print("\n Visualizing final state (CO*)...") view(final_co, viewer='x3d') print("\n✓ NEB analysis complete!") ``` ```{admonition} Missing UMA access? :class: dropdown, tip Don't have access to UMA yet? You can still explore this calculation! [Download example CO on Ni(111) structure](example_configs/co_on_ni111.xyz) and [Download C+O on Ni(111) structure](example_configs/c_o_on_ni111.xyz) to test in the [UMA demo (no login required)](https://facebook-fairchem-uma-demo.hf.space/) and explore the reaction pathway. ``` ```{admonition} Comparison with Paper (Tables 4 & 5) :class: note **Expected Results from Paper:** - **Reaction Energy (C* + O* → CO*)**: **-142.7 kJ/mol = -1.48 eV** (exothermic, DFT-D3) - **Activation Barrier (C* + O* → CO*)**: **153 kJ/mol = 1.59 eV** (reverse/dissociation, DFT-D3) - **CO Adsorption Energy**: **1.82 eV = 175.6 kJ/mol** (DFT-D3) **Reaction Direction:** - Paper reports CO dissociation barrier (CO* → C* + O*), which is the **reverse** of the recombination we calculate - Forward (C* + O* → CO*): barrier = reverse_barrier - |ΔE| ≈ 1.59 - 1.48 ≈ 0.11 eV (very fast) - Reverse (CO* → C* + O*): barrier = 1.59 eV (very slow, kinetic bottleneck) **What to Check:** - Reaction energy (C*+O* → CO*) should be strongly exothermic (~-1.5 eV) - Reverse barrier (CO dissociation) should be substantial (~1.6 eV) - Forward barrier (recombination) should be very small (~0.1 eV) - CO binds much more strongly than H (1.82 eV vs 0.60 eV) **Typical Variations:** - Reaction energies typically accurate within 0.1-0.2 eV - Barriers more sensitive: expect ±0.2-0.3 eV variation - ZPE corrections typically add 0.05-0.15 eV to reaction energies - NEB convergence affects barrier more than reaction energy **Physical Insight:** - Large reverse barrier (1.59 eV) makes CO dissociation very slow at low T - Small forward barrier (0.11 eV) means C+O rapidly recombine to CO - This explains why Ni produces CO in Fischer-Tropsch rather than keeping C and O separate - High temperatures needed to overcome the dissociation barrier for further C-C coupling ``` ```{admonition} NEB Method Explained :class: note The **Nudged Elastic Band (NEB)** method finds the minimum energy path between reactants and products: 1. **Interpolate** between initial and final states (5-9 images typical) 2. **Add spring forces** along the chain to maintain spacing 3. **Project out** spring components perpendicular to the path 4. **Climbing image** variant: highest energy image climbs to saddle point Advantages: - No prior knowledge of transition state needed - Finds entire reaction coordinate - Robust for complex reactions Limitations: - Computationally expensive (optimize N images) - May find wrong path if initial interpolation is poor ``` ### Explore on Your Own 1. **Image convergence**: Run with 7 or 9 images. Does the barrier change? 2. **Spring constant**: Modify the NEB spring constant. How does this affect convergence? 3. **Alternative paths**: Try different initial CO/final C+O configurations. Are there multiple pathways? 4. **Reverse barrier**: Calculate E_a(reverse) = E_a(forward) - ΔE. Check Brønsted-Evans-Polanyi relationship. 5. **Diffusion barriers**: Compute NEB for C or O diffusion on the surface. How do they compare? --- ## Summary and Best Practices ### Key Takeaways 1. **ML Potentials**: uma-s-1p2p1 provides ~1000× speedup over DFT with reasonable accuracy 2. **Bulk optimization**: Always use the ML-optimized lattice constant for consistency 3. **Surface energies**: Linear extrapolation eliminates finite-size effects 4. **Adsorption**: Test multiple sites; lowest energy may not be intuitive 5. **Coverage**: Lateral interactions become significant above ~0.3 ML 6. **Barriers**: NEB requires careful setup but yields full reaction pathway ### Recommended Workflow for New Systems 1. **Optimize Bulk** - Determine equilibrium lattice constant 2. **Calculate Surface Energies** - Identify stable facets 3. **Wulff Construction** - Predict nanoparticle morphology 4. **Low-Coverage Adsorption** - Find binding sites and energies 5. **Coverage Study** (if coverage-dependent effects are important) - Determine lateral interactions 6. **Reaction Barriers** - Calculate activation energies using NEB 7. **Microkinetic Modeling** - Predict overall catalytic performance ### Accuracy Considerations | Property | Typical Error | When Critical | |----------|--------------|---------------| | Lattice constants | 1-2% | Strain effects, alloys | | Surface energies | 10-20% | Nanoparticle shapes | | Adsorption energies | 0.1-0.3 eV | Thermochemistry | | Barriers | 0.2-0.5 eV | Kinetics, selectivity | **Rule of thumb**: Use ML for screening → DFT for validation → Experiment for verification ### Further Reading - **UMA Paper**: [Wood et al. 2025](https://arxiv.org/abs/2506.23971) - **OMat24 Paper**: [Barroso-Luque et al., 2024](https://arxiv.org/abs/2410.12771) - **OC20 Dataset**: [Chanussot et al., ACS Catalysis, 2021](https://pubs.acs.org/doi/full/10.1021/acscatal.0c04525) - **ASE Tutorial**: [https://wiki.fysik.dtu.dk/ase/](https://wiki.fysik.dtu.dk/ase/) --- ## Appendix: Troubleshooting ### Common Issues **Problem**: Convergence failures - **Solution**: Reduce `fmax` to 0.1 initially, tighten later - Check if system is metastable (try different starting geometry) **Problem**: NEB fails to find transition state - **Solution**: Use more images (9-11) or better initial guess - Try fixed-end NEB first, then climbing image **Problem**: Unexpected adsorption energies - **Solution**: Visualize structures - check for distortions - Compare with multiple sites - Add D3 corrections **Problem**: Out of memory - **Solution**: Reduce system size (smaller supercells) - Use fewer NEB images - Run on HPC with more RAM ### Performance Tips 1. **Use batching**: Relax multiple configurations in parallel 2. **Start with DEBUG_MAX_STEPS=50**: Get quick results, refine later 3. **Cache bulk energies**: Don't recalculate reference systems 4. **Trajectory analysis**: Monitor optimization progress with ASE GUI --- ## Caveats and Pitfalls ```{admonition} Important Considerations :class: warning When using ML potentials for surface catalysis, be aware of these critical issues! ``` ### 1. Task Selection: OMAT vs OC20 **Critical choice**: Which task_name to use? - **`task_name="omat"`**: Optimized for bulk and clean surface calculations - Use for: Part 1 (bulk), Part 2 (surface energies), Part 3 (Wulff) - Better for structural relaxations without adsorbates - **`task_name="oc20"`**: Optimized for surface chemistry with adsorbates - Use for: Part 4-6 (all adsorbate calculations) - Trained on Open Catalyst data with adsorbate-surface interactions **Impact**: Using wrong task can lead to 0.1-0.3 eV errors in adsorption energies! ### 2. D3 Dispersion Corrections **Multiple decisions required**: 1. **Whether to use D3 at all?** - Small adsorbates (H, O, N): D3 effect ~0.01-0.05 eV (often negligible) - Large molecules (CO, CO₂, aromatics): D3 effect ~0.1-0.3 eV (important!) - Physisorption: D3 critical (can change binding from repulsive to attractive) - RPBE was originally fit for chemisorption energies without D3 corrections, so adding D3 corrections may actually cause small adsorbates to overbind. However, it probably would be important for larger molecules. It's relatively uncommon to see RPBE+D3 as a choice in the catalysis literature (compared to PBE+D3, or RPBE, or BEEF-vdW). 2. **Which DFT functional for D3?** - This tutorial uses `method="PBE"` consistently for the D3 correction. This is often implied when papers say they use a D3 correction, but the results can be different if use the RPBE parameterizations. - Original paper used PBE for bulk/surfaces, RPBE for adsorption. It's not specified what D3 parameterization they used, but it's likely PBE. 3. **When to apply D3?** - **End-point correction** (used here): Fast, run ML optimization then add D3 energy - **During optimization**: Slower but more accurate geometries - **Impact**: Usually <0.05 eV difference, but can be larger for weak interactions ### 3. Coverage Dependence Challenges **Non-linearity at high coverage**: - This tutorial assumes linear E_ads(θ) = E₀ + βθ - Reality: Often non-linear, especially near θ = 1 ML. See the plots generated - there is a linear regime for relatively high coverage, and relatively low coverage, but it's not uniformly linear everywhere. As long as you consistently in one regime or the other a linear assumption is probably ok, but you could get into problems if solving microkinetic models where the coverage of the species in question changes significantly from very low to high. - **Why**: Phase transitions, adsorbate ordering, surface reconstruction - **Solution**: Test polynomial fits, look for ordering in visualizations **Low coverage limit**: - At θ < 0.1 ML, coverage effects are tiny (<0.01 eV) - Hard to distinguish from numerical noise - **Best practice**: Focus on 0.25-1.0 ML range for fitting ### 4. Periodic Boundary Conditions **UMa requires PBC=True in all directions!** ```python atoms.set_pbc([True, True, True]) # Always required ``` - Forgetting this causes crashes or wrong energies - Even for "gas phase" molecules in vacuum ### 5. Gas Phase Reference Energies **Tricky cases**: - **H₂(g)**: UMa handles well (used in this tutorial) - **H(g)**: May not be reliable (use H₂/2 instead) - **CO(g)**, **O₂(g)**: Usually okay, but check against DFT - **Radicals**: Often problematic **Best practice**: Always use stable molecules as references (H₂, not H; H₂O, not OH) ### 6. Spin Polarization **Key limitation**: OC20/UMa does not include spin! - Paper used spin-polarized DFT - **Impact**: Usually small (0.05-0.1 eV) - **Larger** for: - Magnetic metals (Fe, Co, Ni) - Open-shell adsorbates (O*, OH*) - Reaction barriers with radicals ### 7. Constraint Philosophy **Clean slabs** (Part 2): No constraints (both surfaces relax) - Best for surface energy calculations - More physical for symmetric slabs **Adsorbate slabs** (Part 4-6): Bottom layers fixed - Faster convergence - Prevents adsorbate-induced reconstruction - Standard practice in surface chemistry **Fairchem helper functions**: Automatically apply sensible constraints - Trust their heuristics unless you have good reason not to - Check `atoms.constraints` to see what was applied ### 8. Complex Surface Structures **This tutorial uses low-index facets** (111, 100, 110, 211) - Well-defined, symmetric - Easy to generate and analyze **Real catalysts** have: - Steps, kinks, grain boundaries - Support interfaces - Defects and vacancies - **Challenge**: Harder to generate, more configurations to test ### 9. Slab Thickness and Vacuum **Convergence tests critical** but expensive: - This tutorial uses "reasonable" values (4-8 layers, 10 Å vacuum) - **Always check** convergence for new systems - **Especially important** for: - Metals with long electron screening (Au, Ag) - Charged adsorbates - Strong adsorbate-induced reconstruction ### 10. NEB Convergence **Most computationally expensive part**: - May need 7-11 images (not just 5) - Initial guess matters a lot - Can get stuck in local minima **Tricks**: 1. Use dimer method to find better TS guess (as shown in Part 6) 2. Start with coarse convergence (fmax=0.2), refine later 3. Visualize the path - does it make chemical sense? 4. Try different spring constants (0.1-1.0 eV/Å) ### 11. Lattice Constant Source **Consistency is key**: - Use ML-optimized lattice constant throughout (as done here) - **Don't mix**: ML lattice + DFT surface energies = inconsistent - Alternative: Use experimental lattice constant for everything ### 12. Adsorbate Placement **Multiple local minima**: - Surface chemistry is **not** convex! - Always test multiple adsorption sites - Fairchem helpers generate ~5 configurations in this tutorial, but you may need more to search many modes. You can already try methods like minima hopping or other global optimization methods to sample more configurations. **For complex adsorbates**: - Test different orientations - May need 10-20 configurations - Consider genetic algorithms or basin hopping --- ```{admonition} Congratulations! 🎉 :class: tip You've completed a comprehensive computational catalysis workflow using state-of-the-art ML potentials. You can now: - Characterize catalyst surfaces computationally - Predict nanoparticle shapes - Calculate reaction thermodynamics and kinetics - Apply these methods to your own research questions **Next steps**: - Apply to your catalyst system of interest - Validate key results with DFT - Develop microkinetic models - Publish your findings! ``` --- Source: `docs/core/fair_chemistry_papers.md` # FAIR Chemistry Papers Research publications from the FAIR Chemistry team at Meta, advancing machine learning for atomic simulations across molecules, materials, and catalysts. :::{tip} Click on any paper card to expand and see the full abstract and author list. ::: ## Universal Models & Architectures State-of-the-art neural network architectures for atomic property prediction. ::::{grid} 1 1 2 2 :::{card} UMA: A Family of Universal Models for Atoms :link: https://arxiv.org/abs/2506.23971 **2025** — Trained on 500M+ structures with 1.4B parameters but only ~50M active per structure using mixture of linear experts ::: :::{card} eSEN: Learning Smooth and Expressive Interatomic Potentials :link: https://arxiv.org/abs/2502.12147 **2025** — Passes energy conservation tests; SOTA on thermal conductivity, phonons, and materials stability prediction ::: :::{card} EquiformerV2: Improved Equivariant Transformer :link: https://arxiv.org/abs/2306.12059 **2023** — Up to 9% improvement on forces, 4% on energies; 2x reduction in DFT calculations for adsorption energies ::: :::{card} eSCN: Reducing SO(3) Convolutions to SO(2) :link: https://arxiv.org/abs/2302.03655 **2023** — Reduces equivariant convolution complexity from O(L⁶) to O(L³); SOTA on OC-20 and OC-22 ::: :::: ```{admonition} More Architecture Papers :class: dropdown ### GemNet-OC: Graph Neural Networks for Large Datasets **arXiv:** [2204.02782](https://arxiv.org/abs/2204.02782) (2022) **Authors:** Johannes Gasteiger, Muhammed Shuaibi, Anuroop Sriram, Stephan Günnemann, Zachary Ulissi, C. Lawrence Zitnick, Abhishek Das **Abstract:** Recent years have seen the advent of molecular simulation datasets that are orders of magnitude larger and more diverse. These new datasets differ substantially in four aspects of complexity: 1. Chemical diversity, 2. system size, 3. dataset size, and 4. domain shift. We develop GemNet-OC which outperforms the previous state-of-the-art on OC20 by 16% while reducing training time by a factor of 10. --- ### Spherical Channels for Modeling Atomic Interactions **arXiv:** [2206.14331](https://arxiv.org/abs/2206.14331) (2022) **Authors:** C. Lawrence Zitnick, Abhishek Das, Adeesh Kolluru, Janice Lan, Muhammed Shuaibi, Anuroop Sriram, Zachary Ulissi, Brandon Wood **Abstract:** We propose the Spherical Channel Network (SCN) to model atomic energies and forces. The SCN is a graph neural network where atom embeddings are spherical functions represented using spherical harmonics. We demonstrate state-of-the-art results on the large-scale Open Catalyst dataset. --- ### ForceNet: A Graph Neural Network for Large-Scale Quantum Calculations **arXiv:** [2103.01436](https://arxiv.org/abs/2103.01436) (2021) **Authors:** Weihua Hu, Muhammed Shuaibi, Abhishek Das, Siddharth Goyal, Anuroop Sriram, Jure Leskovec, Devi Parikh, C. Lawrence Zitnick **Abstract:** By not imposing explicit physical constraints, we can flexibly design expressive models while maintaining computational efficiency. Physical constraints are implicitly imposed through physics-based data augmentation. ForceNet predicts atomic forces more accurately than state-of-the-art physics-based GNNs while being faster. --- ### Rotation Invariant Graph Neural Networks using Spin Convolutions **arXiv:** [2106.09575](https://arxiv.org/abs/2106.09575) (2021) **Authors:** Muhammed Shuaibi, Adeesh Kolluru, Abhishek Das, Aditya Grover, Anuroop Sriram, Zachary Ulissi, C. Lawrence Zitnick **Abstract:** We introduce a novel approach to modeling angular information between atoms using per-edge local coordinate frames and spin convolutions. State-of-the-art results are demonstrated on the large-scale Open Catalyst 2020 dataset. --- ### Towards Training Billion Parameter Graph Neural Networks **arXiv:** [2203.09697](https://arxiv.org/abs/2203.09697) (2022) **Authors:** Anuroop Sriram, Abhishek Das, Brandon M. Wood, Siddharth Goyal, C. Lawrence Zitnick **Abstract:** We introduce Graph Parallelism, a method to distribute input graphs across multiple GPUs, enabling training of very large GNNs with billions of parameters. On OC20, graph-parallelized models achieve 15% relative improvement on force MAE. ``` --- ## Datasets Large-scale open datasets powering the next generation of ML models for chemistry. ::::{grid} 1 1 2 2 :::{card} OMol25: Open Molecules 2025 :link: https://arxiv.org/abs/2505.08762 **2025** — 100M+ DFT calculations, 83 elements, molecules up to 350 atoms ::: :::{card} OMat24: Open Materials 2024 :link: https://arxiv.org/abs/2410.12771 **2024** — 110M+ DFT calculations for inorganic materials ::: :::{card} OMC25: Open Molecular Crystals 2025 :link: https://arxiv.org/abs/2508.02651 **2025** — 27M+ molecular crystal structures with DFT labels ::: :::{card} ODAC25: Open DAC 2025 :link: https://arxiv.org/abs/2508.03162 **2025** — 60M DFT calculations for MOF sorbent discovery ::: :::: ```{admonition} All Dataset Papers :class: dropdown ### The Open Polymers 2026 (OPoly26) Dataset **arXiv:** [2512.23117](https://arxiv.org/abs/2512.23117) (2025) **Authors:** Daniel S. Levine, Nicholas Liesen, Lauren Chua, James Diffenderfer, et al. **Abstract:** We create the Open Polymers 2026 (OPoly26) dataset containing more than 6.57 million DFT calculations on polymer clusters comprising over 1.2 billion total atoms. OPoly26 captures chemical diversity including variations in monomer composition, degree of polymerization, chain architectures, and solvation environments. --- ### The Open Catalyst 2025 (OC25) Dataset for Solid-Liquid Interfaces **arXiv:** [2509.17862](https://arxiv.org/abs/2509.17862) (2025) **Authors:** Sushree Jagriti Sahoo, Mikael Maraschin, Daniel S. Levine, Zachary Ulissi, et al. **Abstract:** We introduce OC25, consisting of 7,801,261 calculations across 1,511,270 unique explicit solvent environments. OC25 is the largest solid-liquid interface dataset available, spanning 88 elements with commonly used solvents/ions. --- ### The Open Molecules 2025 (OMol25) Dataset **arXiv:** [2505.08762](https://arxiv.org/abs/2505.08762) (2025) **Authors:** Daniel S. Levine, Muhammed Shuaibi, Evan Walter Clark Spotte-Smith, et al. **Abstract:** OMol25 is composed of more than 100 million DFT calculations at the ωB97M-V/def2-TZVPD level. It uniquely blends 83 elements, diverse intra- and intermolecular interactions, explicit solvation, variable charge/spin, conformers, and reactive structures across ~83M unique molecular systems. --- ### Open Molecular Crystals 2025 (OMC25) Dataset **arXiv:** [2508.02651](https://arxiv.org/abs/2508.02651) (2025) **Authors:** Vahe Gharakhanyan, Luis Barroso-Luque, Yi Yang, Muhammed Shuaibi, et al. **Abstract:** OMC25 contains over 27 million molecular crystal structures with 12 elements and up to 300 atoms per unit cell, generated from dispersion-inclusive DFT relaxation trajectories of over 230,000 structures. --- ### The Open DAC 2025 (ODAC25) Dataset **arXiv:** [2508.03162](https://arxiv.org/abs/2508.03162) (2025) **Authors:** Anuroop Sriram, Logan M. Brabson, Xiaohan Yu, Sihoon Choi, et al. **Abstract:** ODAC25 comprises nearly 60 million DFT calculations for CO₂, H₂O, N₂, and O₂ adsorption in 15,000 MOFs. It introduces chemical diversity through functionalized MOFs, GCMC-derived placements, and synthetically generated frameworks. --- ### Open Materials 2024 (OMat24) Dataset **arXiv:** [2410.12771](https://arxiv.org/abs/2410.12771) (2024) **Authors:** Luis Barroso-Luque, Muhammed Shuaibi, Xiang Fu, Brandon M. Wood, et al. **Abstract:** OMat24 contains over 110 million DFT calculations focused on structural and compositional diversity. EquiformerV2 models achieve state-of-the-art on Matbench Discovery with F1 score above 0.9. --- ### Open Catalyst Experiments 2024 (OCx24) **arXiv:** [2411.11783](https://arxiv.org/abs/2411.11783) (2024) **Authors:** Jehad Abed, Jiheon Kim, Muhammed Shuaibi, Brook Wander, et al. **Abstract:** OCX24 bridges experiments and computational models with 572 synthesized samples and 441 gas diffusion electrodes evaluated for CO₂RR and HER. DFT-verified adsorption energies were calculated on ~20,000 materials requiring 685 million AI-accelerated relaxations. --- ### The Open Catalyst 2022 (OC22) Dataset **arXiv:** [2206.08917](https://arxiv.org/abs/2206.08917) (2022) **Authors:** Richard Tran, Janice Lan, Muhammed Shuaibi, Brandon M. Wood, et al. **Abstract:** OC22 consists of 62,331 DFT relaxations (~9,854,504 single point calculations) across oxide materials, coverages, and adsorbates, critical for OER catalyst development. --- ### The Open DAC 2023 (ODAC23) Dataset **arXiv:** [2311.00341](https://arxiv.org/abs/2311.00341) (2023) **Authors:** Anuroop Sriram, Sihoon Choi, Xiaohan Yu, Logan M. Brabson, et al. **Abstract:** ODAC23 consists of more than 38M DFT calculations on more than 8,400 MOF materials containing adsorbed CO₂ and/or H₂O for direct air capture applications. --- ### The Open Catalyst 2020 (OC20) Dataset **arXiv:** [2010.09990](https://arxiv.org/abs/2010.09990) (2020) **Authors:** Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, et al. **Abstract:** OC20 consists of 1,281,040 DFT relaxations (~264,890,000 single point evaluations) across materials, surfaces, and adsorbates. The foundational dataset for the Open Catalyst Project. ``` --- ## Generative Models Generating novel molecules, materials, and crystal structures. ::::{grid} 1 1 2 2 :::{card} ADiT: All-atom Diffusion Transformers :link: https://arxiv.org/abs/2503.03965 **2025** — Unified generative model for molecules and materials ::: :::{card} FlowMM: Riemannian Flow Matching for Materials :link: https://arxiv.org/abs/2406.04713 **2024** — 3x more efficient at finding stable materials ::: :::{card} FlowLLM: LLMs + Flow Matching for Crystals :link: https://arxiv.org/abs/2410.23405 **2024** — 3x higher generation rate of stable materials ::: :::{card} Space Group Conditional Flow Matching :link: https://arxiv.org/abs/2509.23822 **2025** — Symmetry-aware crystal generation ::: :::: ```{admonition} More Generative Model Papers :class: dropdown ### All-atom Diffusion Transformers (ADiT) **arXiv:** [2503.03965](https://arxiv.org/abs/2503.03965) (2025) **Authors:** Chaitanya K. Joshi, Xiang Fu, Yi-Lun Liao, Vahe Gharakhanyan, Benjamin Kurt Miller, Anuroop Sriram, Zachary W. Ulissi **Abstract:** We introduce ADiT, a unified latent diffusion framework for jointly generating both periodic materials and non-periodic molecular systems. An autoencoder maps molecules and materials to a shared latent space, then diffusion generates new embeddings. ADiT achieves state-of-the-art results on MP20, QM9, and GEOM-DRUGS with significant speedups over equivariant diffusion models. --- ### Space Group Conditional Flow Matching **arXiv:** [2509.23822](https://arxiv.org/abs/2509.23822) (2025) **Authors:** Omri Puny, Yaron Lipman, Benjamin Kurt Miller **Abstract:** We introduce a generative framework that samples highly-symmetric, stable crystals by conditioning on space groups and Wyckoff positions. Using efficient group averaging, we achieve state-of-the-art on crystal structure prediction and de novo generation benchmarks. --- ### FlowLLM: LLMs and Riemannian Flow Matching for Materials **arXiv:** [2410.23405](https://arxiv.org/abs/2410.23405) (2024) **Authors:** Anuroop Sriram, Benjamin Kurt Miller, Ricky T. Q. Chen, Brandon M. Wood **Abstract:** FlowLLM combines fine-tuned LLMs with Riemannian flow matching. It outperforms state-of-the-art by 3x on stable material generation rate and ~50% on stable, unique, and novel crystals. --- ### FlowMM: Riemannian Flow Matching for Materials **arXiv:** [2406.04713](https://arxiv.org/abs/2406.04713) (2024) **Authors:** Benjamin Kurt Miller, Ricky T. Q. Chen, Anuroop Sriram, Brandon M Wood **Abstract:** We generalize Riemannian Flow Matching to crystals with translation, rotation, permutation, and periodic boundary conditions. FlowMM is ~3x more efficient at finding stable materials compared to previous methods. --- ### Fine-Tuned Language Models Generate Stable Inorganic Materials **arXiv:** [2402.04379](https://arxiv.org/abs/2402.04379) (2024) **Authors:** Nate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson, C. Lawrence Zitnick, Zachary Ulissi **Abstract:** Fine-tuned LLaMA-2 70B generates materials predicted to be metastable at about twice the rate (49% vs 28%) of CDVAE. Around 90% of sampled structures obey physical constraints, demonstrating LLMs' surprising suitability for atomistic data. --- ### FastCSP: Accelerated Molecular Crystal Structure Prediction **arXiv:** [2508.02641](https://arxiv.org/abs/2508.02641) (2025) **Authors:** Vahe Gharakhanyan, Yi Yang, Luis Barroso-Luque, Muhammed Shuaibi, et al. **Abstract:** FastCSP combines random structure generation with UMA-powered relaxation and free energy calculations. Results for a single system can be obtained within hours on tens of GPUs, making high-throughput crystal structure prediction feasible. ``` --- ## Sampling & Molecular Dynamics Advanced methods for sampling molecular configurations and running simulations. ::::{grid} 1 1 2 2 :::{card} Adjoint Sampling: Highly Scalable Diffusion Samplers :link: https://arxiv.org/abs/2504.11713 **2025** — First on-policy approach that scales to large problems ::: :::{card} Adjoint Schrödinger Bridge Sampler :link: https://arxiv.org/abs/2506.22565 **2025** — Kinetic-optimal transportation for molecular sampling ::: :::{card} Enhancing Diffusion Sampling with Collective Variables :link: https://arxiv.org/abs/2510.11923 **2025** — First reactive sampling with diffusion-based samplers ::: :::: ```{admonition} Sampling Paper Details :class: dropdown ### Adjoint Sampling **arXiv:** [2504.11713](https://arxiv.org/abs/2504.11713) (2025) **Authors:** Aaron Havens, Benjamin Kurt Miller, Bing Yan, Carles Domingo-Enrich, Anuroop Sriram, Brandon Wood, Daniel Levine, et al. **Abstract:** We introduce Adjoint Sampling, a highly scalable algorithm for learning diffusion processes that sample from unnormalized densities. It's the first on-policy approach allowing significantly more gradient updates than energy evaluations. We demonstrate amortized conformer generation across many molecular systems. --- ### Adjoint Schrödinger Bridge Sampler (ASBS) **arXiv:** [2506.22565](https://arxiv.org/abs/2506.22565) (2025) **Authors:** Guan-Horng Liu, Jaemoo Choi, Yongxin Chen, Benjamin Kurt Miller, Ricky T. Q. Chen **Abstract:** ASBS employs simple matching-based objectives without needing target samples during training. Grounded in the Schrödinger Bridge for kinetic-optimal transportation, ASBS generalizes to arbitrary source distributions for sampling from molecular Boltzmann distributions. --- ### Enhancing Diffusion-Based Sampling with Molecular Collective Variables **arXiv:** [2510.11923](https://arxiv.org/abs/2510.11923) (2025) **Authors:** Juno Nam, Bálint Máté, Artur P. Toshev, Manasa Kaniselvan, Rafael Gómez-Bombarelli, et al. **Abstract:** We introduce a sequential bias along collective variables to improve mode discovery and enable free energy estimation. We are the first to demonstrate reactive sampling using a diffusion-based sampler, capturing bond breaking and formation with universal interatomic potentials. ``` --- ## Applications & Discovery Practical applications accelerating catalyst and materials discovery. ::::{grid} 1 1 2 2 :::{card} CatTSunami: Accelerating Transition State Calculations :link: https://arxiv.org/abs/2405.02078 **2024** — 1500x speedup for reaction network enumeration ::: :::{card} AdsorbML: Efficient Adsorption Energy Calculations :link: https://arxiv.org/abs/2211.16486 **2022** — 2000x speedup with 87% accuracy ::: :::{card} Catlas: Automated Catalyst Discovery Framework :link: https://arxiv.org/abs/2208.12717 **2022** — High-throughput screening for syngas conversion ::: :::{card} GA-Accelerated Liquid Crystal Polymer Discovery :link: https://arxiv.org/abs/2505.13477 **2025** — Genetic algorithms + first-principles for VR/AR materials ::: :::: ```{admonition} Application Paper Details :class: dropdown ### CatTSunami: Accelerating Transition State Energy Calculations **arXiv:** [2405.02078](https://arxiv.org/abs/2405.02078) (2024) **Authors:** Brook Wander, Muhammed Shuaibi, John R. Kitchin, Zachary W. Ulissi, C. Lawrence Zitnick **Abstract:** We show GNN potentials trained on OC20 find transition states within 0.1 eV of DFT 91% of the time with 28x speedup. We replicated a reaction network with 61 intermediates and 174 reactions at DFT resolution using just 12 GPU days (vs 52 GPU years), a 1500x speedup. --- ### AdsorbML: A Leap in Efficiency for Adsorption Energy Calculations **arXiv:** [2211.16486](https://arxiv.org/abs/2211.16486) (2022) **Authors:** Janice Lan, Aini Palizhati, Muhammed Shuaibi, Brandon M. Wood, et al. **Abstract:** We demonstrate ML potentials can identify low energy adsorbate-surface configurations with 87.36% accuracy while achieving a 2000x speedup. We introduce the Open Catalyst Dense dataset with ~1,000 surfaces and 100,000 configurations. --- ### Catlas: Automated Framework for Catalyst Discovery **arXiv:** [2208.12717](https://arxiv.org/abs/2208.12717) (2022) **Authors:** Brook Wander, Kirby Broderick, Zachary W. Ulissi **Abstract:** Catlas explores large design spaces using pre-trained ML models without upfront training. For syngas-to-oxygenates conversion, we explored 947 binary intermetallics, identifying 144 candidate materials including Pt-Ti, Pd-V, Ni-Nb, and Ti-Zn. --- ### Genetic Algorithm-Accelerated Discovery of Liquid Crystal Polymers **arXiv:** [2505.13477](https://arxiv.org/abs/2505.13477) (2025) **Authors:** Jianing Zhou, Yuge Huang, Arman Boromand, Keian Noori, et al. **Abstract:** We integrate first-principles calculations with genetic algorithms to discover liquid crystal polymers with low visible absorption and high refractive index for VR/AR/MR technologies. --- ### Adapting OC20-trained EquiformerV2 for High-Entropy Materials **arXiv:** [2403.09811](https://arxiv.org/abs/2403.09811) (2024) **Authors:** Christian M. Clausen, Jan Rossmeisl, Zachary W. Ulissi **Abstract:** We show that through energy filtering and few-shot fine-tuning, EquiformerV2 achieves state-of-the-art accuracy on high-entropy alloys (Ag-Ir-Pd-Pt-Ru), demonstrating that foundational knowledge from ordered structures extrapolates to disordered solid-solutions. ``` --- ## Training Methods & Techniques Innovations in model training, active learning, and pre-training strategies. ::::{grid} 1 1 2 2 :::{card} DeNS: Generalizing Denoising to Non-Equilibrium Structures :link: https://arxiv.org/abs/2403.09549 **2024** — New SOTA on OC20 and OC22 ::: :::{card} JMP: Joint Multi-domain Pre-training :link: https://arxiv.org/abs/2310.16802 **2023** — 59% average improvement over training from scratch ::: :::{card} FINETUNA: Fine-tuning Accelerated Simulations :link: https://arxiv.org/abs/2205.01223 **2022** — 91% reduction in DFT calculations ::: :::: ```{admonition} Training Methods Paper Details :class: dropdown ### Generalizing Denoising to Non-Equilibrium Structures (DeNS) **arXiv:** [2403.09549](https://arxiv.org/abs/2403.09549) (2024) **Authors:** Yi-Lun Liao, Tess Smidt, Muhammed Shuaibi, Abhishek Das **Abstract:** We propose denoising non-equilibrium structures as an auxiliary training task. By encoding forces to specify which structure we're denoising, DeNS achieves new state-of-the-art on OC20 and OC22 while significantly improving training efficiency on MD17. --- ### From Molecules to Materials: Joint Multi-domain Pre-training (JMP) **arXiv:** [2310.16802](https://arxiv.org/abs/2310.16802) (2023) **Authors:** Nima Shoghi, Adeesh Kolluru, John R. Kitchin, Zachary W. Ulissi, C. Lawrence Zitnick, Brandon M. Wood **Abstract:** JMP simultaneously trains on multiple datasets (~120M systems from OC20, OC22, ANI-1x, Transition-1x) within a multi-task framework. It demonstrates 59% average improvement over training from scratch, matching or setting SOTA on 34 of 40 tasks. --- ### FINETUNA: Fine-tuning Accelerated Molecular Simulations **arXiv:** [2205.01223](https://arxiv.org/abs/2205.01223) (2022) **Authors:** Joseph Musielewicz, Xiaoxiao Wang, Tian Tian, Zachary Ulissi **Abstract:** An online active learning framework that incorporates prior information from pre-trained models. It accelerates simulations by reducing DFT calculations by 91% while meeting accuracy thresholds 93% of the time. --- ### Robust and Scalable Uncertainty Estimation with Conformal Prediction **arXiv:** [2208.08337](https://arxiv.org/abs/2208.08337) (2022) **Authors:** Yuge Hu, Joseph Musielewicz, Zachary Ulissi, Andrew J. Medford **Abstract:** We combine conformal prediction with latent space distances to estimate uncertainty of neural network force fields. The method is calibrated, sharp, and scalable to training datasets of 1 million images. --- ### Generalization of Graph-Based Active Learning Relaxation Strategies **arXiv:** [2311.01987](https://arxiv.org/abs/2311.01987) (2023) **Authors:** Xiaoxiao Wang, Joseph Musielewicz, Richard Tran, et al. **Abstract:** We investigate Finetuna on out-of-domain systems: larger adsorbates, metal-oxides with spin polarization, and 3D structures like zeolites and MOFs. The framework reduces DFT calculations by 80% for alcohols/3D structures and 42% for oxides. --- ### Enabling Robust Offline Active Learning for ML Potentials **arXiv:** [2008.10773](https://arxiv.org/abs/2008.10773) (2020) **Authors:** Muhammed Shuaibi, Saurabh Sivakumar, Rui Qi Chen, Zachary W. Ulissi **Abstract:** We demonstrate a Δ-machine learning approach enabling stable convergence in offline active learning by avoiding unphysical configurations, reducing first-principles calculations by 70-90%. ``` --- ## Electronic Structure & Properties Predicting electronic structure and chemical properties beyond energies and forces. ::::{grid} 1 1 2 2 :::{card} HELM: Hamiltonian-trained Electronic-structure Learning :link: https://arxiv.org/abs/2510.00224 **2025** — Hamiltonian prediction for 100+ atom structures ::: :::{card} Chemical Properties from GNN-Predicted Electron Densities :link: https://arxiv.org/abs/2309.04811 **2023** — Atomic charges and dipole moments from electron density ::: :::{card} XRD Loss Landscape Analysis :link: https://arxiv.org/abs/2512.04036 **2025** — Insights for crystal structure from powder diffraction ::: :::: ```{admonition} Electronic Structure Paper Details :class: dropdown ### HELM: Learning from Electronic Structure Across the Periodic Table **arXiv:** [2510.00224](https://arxiv.org/abs/2510.00224) (2025) **Authors:** Manasa Kaniselvan, Benjamin Kurt Miller, Meng Gao, Juno Nam, Daniel S. Levine **Abstract:** We introduce HELM, a state-of-the-art Hamiltonian prediction model scaling to 100+ atom structures with 58 elements and large basis sets. We release the 'OMol_CSH_58k' dataset and demonstrate 'Hamiltonian pretraining' improves energy-prediction in low-data regimes. --- ### Chemical Properties from Graph Neural Network-Predicted Electron Densities **arXiv:** [2309.04811](https://arxiv.org/abs/2309.04811) (2023) **Authors:** Ethan M. Sunshine, Muhammed Shuaibi, Zachary W. Ulissi, John R. Kitchin **Abstract:** We demonstrate using established physical methods to obtain chemical properties from model-predicted electron densities. Without training to predict charges, the model predicts atomic charges with an order of magnitude lower error than sum of atomic densities. --- ### The Loss Landscape of Powder X-Ray Diffraction-Based Structure Optimization **arXiv:** [2512.04036](https://arxiv.org/abs/2512.04036) (2025) **Authors:** Nofit Segal, Akshay Subramanian, Mingda Li, Benjamin Kurt Miller, Rafael Gomez-Bombarelli **Abstract:** We study the powder XRD-to-structure mapping using gradient descent. Commonly used XRD similarity metrics result in highly non-convex landscapes. Constraining optimization to the ground-truth crystal family significantly improves structure recovery. --- ### A Physically-informed Graph-based Order Parameter **arXiv:** [2106.08215](https://arxiv.org/abs/2106.08215) (2021) **Authors:** James Chapman, Nir Goldman, Brandon Wood **Abstract:** A universal, transferable graph-based order parameter for characterizing atomistic structures. Validated on liquid lithium up to 300 GPa, carbon phases including nanotubes, and diverse aluminum configurations. Outperforms all existing crystalline order parameters. ``` --- ## Perspectives & Introductions Overview papers and perspectives on the field. ```{admonition} Perspectives & Introductions :class: dropdown ### Open Challenges in Developing Large Scale ML Models for Catalyst Discovery **arXiv:** [2206.02005](https://arxiv.org/abs/2206.02005) (2022) **Authors:** Adeesh Kolluru, Muhammed Shuaibi, Aini Palizhati, Nima Shoghi, et al. **Abstract:** We discuss challenges and findings of developments on OC20, examining performance across materials and adsorbates. We cover energy-conservation, finding local minima, and augmentation of off-equilibrium data. --- ### An Introduction to Electrocatalyst Design using Machine Learning **arXiv:** [2010.09435](https://arxiv.org/abs/2010.09435) (2020) **Authors:** C. Lawrence Zitnick, Lowik Chanussot, Abhishek Das, Siddharth Goyal, et al. **Abstract:** An introduction to challenges in finding electrocatalysts for renewable energy storage, how machine learning may be applied, and the use of the Open Catalyst Project OC20 dataset for model training. --- ### Computational Catalyst Discovery: Active Classification through Myopic Multiscale Sampling **arXiv:** [2102.01528](https://arxiv.org/abs/2102.01528) (2021) **Authors:** Kevin Tran, Willie Neiswanger, Kirby Broderick, Eric Xing, Jeff Schneider, Zachary W. Ulissi **Abstract:** We present myopic multiscale sampling, which combines multiscale modeling with automated DFT selection. This active classification strategy achieves ~7-16x speedup in catalyst classification relative to random sampling. ``` --- Source: `docs/faq.md` # Frequently Asked Questions :::{margin} ```{image} assets/icons/catalysis.svg :alt: Catalysis :width: 100px ``` ::: Answers to common questions about UMA, FAIR Chemistry, and catalysis applications are collected here. ## UMA :::{admonition} How do I choose the right task? :class: dropdown Choose the task based on your application domain and required level of theory: - **`omol`**: Molecules, organic chemistry, pharmaceuticals, and polymers. - **`omat`**: Inorganic materials, alloys, and bulk properties. - **`oc20`**: Heterogeneous catalysis without oxides or explicit solvent. - **`oc22`**: Oxide catalysts and supports. - **`oc25`**: Catalyst-electrolyte interfaces with explicit solvent. - **`odac`**: Metal-organic frameworks and direct air capture. - **`omc`**: Molecular crystals and organic electronics. See the [UMA task guide](./core/uma.md#the-uma-task) for the corresponding DFT methods and limitations. ::: :::{admonition} Why am I getting a 401 error when loading UMA? :class: dropdown UMA is gated on Hugging Face. Request access to the [UMA repository](https://huggingface.co/facebook/UMA), create a token with read access to public gated repositories, and authenticate with `huggingface-cli login` or the `HF_TOKEN` environment variable. ::: :::{admonition} Which checkpoint should I use? :class: dropdown Start with `uma-s-1p2p1`. It is the fastest current UMA model while retaining state-of-the-art accuracy on most supported benchmarks. Use `uma-m-1p1` when its higher accuracy justifies greater memory use and slower inference. ::: :::{admonition} Can I use the `omol` task for periodic systems? :class: dropdown The `omol` task was trained on aperiodic molecular data, so periodic systems should be treated with caution. Use `omat`, `omc`, `odac`, or a catalysis task when one of those training domains matches your system. ::: :::{admonition} How do I specify molecular charge and spin? :class: dropdown For `omol`, set total charge and spin multiplicity on the ASE object: ```python atoms.info.update({"charge": 0, "spin": 1}) ``` Other tasks expect the values documented in the [UMA model guide](./core/uma.md). ::: ## Open Catalyst Project ::::{grid} 1 2 2 4 :::{grid-item-card} General :link: #general Basic concepts and terminology ::: :::{grid-item-card} ML Questions :link: #ml-questions Model details and predictions ::: :::{grid-item-card} Catalysis :link: #catalysis-questions Using predictions for research ::: :::{grid-item-card} Computational :link: #computational-catalysis-questions Technical details and DFT ::: :::: --- (general)= ### General :::{admonition} What is a catalyst? What is a bulk? What is a surface? What is an adsorbate? :class: dropdown * **Catalyst:** A catalyst is a material that makes a chemical reaction happen faster without itself being consumed. The catalysts shown here are one specific type of catalyst (heterogeneous) catalyst where the reactions happen as molecules interact with a catalyst surface. * **Bulk:** An inorganic crystal structure that will form the base structure of our catalyst surface and from which we will identify possible surfaces. * **Surface:** A surface formed from a bulk crystal structure by cutting along a specific plane (specified by the Miller index) * **Adsorbate:** The small organic molecule interacting with the catalyst surface that is adsorbed on the catalyst surface. ::: :::{admonition} What are these models and simulations useful for? :class: dropdown Simulations like these are used to design materials (catalysts) that make chemical transportations more efficient. Example uses of catalysis in day-to-day life: * The catalytic converter in your car is a special catalyst that breaks down harmful automotive exhaust to more benign species * Catalysts are responsible for taking nitrogen in the air and making ammonia for fertilizer and agricultural use. We couldn't feed the world's population without catalysts! * Energy can be stored as renewable hydrogen produced by splitting water over electrocatalysts * Fuel cell vehicles rely on catalysts to combine hydrogen and oxygen into water while producing electricity * Carbon capture systems can use catalysts to turn waste or captured CO2 into useful chemicals These simulations can help us design cheaper, more efficient, and more durable catalysts. These simulations can also help design catalysts for new chemical reactions that are currently difficult or infeasible to drive at scale. These simulations could also be used to design materials that are resistant to surface oxidation or corrosion. ::: :::{admonition} Why does Meta care about catalysis? :class: dropdown * Catalysis is important for climate/renewable energy and achieving net-zero carbon emissions in the near future. Meta has a number of decarbonization goals that rely on societal solutions for renewable energy and energy storage. Catalysis is key to many of these technologies, and helping the community develop more efficient processes will help us reach those goals! * Meta Fundamental AI Research (FAIR) is interested in pushing forward the state-of-the-art in many applications of AI/ML to the world, including vision, language, robotics, etc. Applications of AI in science and chemistry is one such exciting area of research. * The core models powering this API are from a class of models known as [graph neural networks](https://distill.pub/2021/gnn-intro/) (GNNs). GNNs are an exciting area of research in AI. ::: :::{admonition} Are there any rate limits on incoming requests? :class: dropdown Rate limits may apply when using the OCP API. Check the [API documentation](./catalysts/examples_tutorials/ocpapi.md) for current limits. ::: :::{admonition} How do I know which surface/adsorbate/etc to pick when running this? :class: dropdown **If you are new to catalysis:** * Try selecting some simple (one element) or complicated (multiple elements) bulk structures. * Try selecting some small and larger adsorbates to see how interesting the structures may be * Watch some of the relaxations and see how subtle some of these relaxations can be! * Develop some intuition for how the energies depend on different surfaces and bulk structures If you have ideas on how to build better ML models to predict the final energies or relaxation pathways, you should visit opencatalystproject.org! Hopefully this website gives you an idea of the challenge here and what the data will look like. **If you are an experienced catalyst researcher:** * You might want to see how common descriptors for your favorite reactions (like *CO or *CH2) might vary across different surfaces or compositions * If you are interested in reaction pathways for a single surface, you might want to predict properties for several adsorbates on the same surface * Let us know if you have other suggestions here! We'd love to hear what you're using this for! Hopefully this gives you an idea of what the state of the art is for reactive catalyst potentials! The world is your oyster! ::: :::{admonition} What are the black lines and boxes in all of the visualizations? :class: dropdown The systems considered here are periodic, i.e. they repeat over and over infinitely in the X/Y direction to approximate a very large catalyst surface. The black lines show the edges of the repeating unit cell, or [periodic boundary conditions](https://en.wikipedia.org/wiki/Periodic_boundary_conditions). ::: :::{admonition} Why do only some of the atoms move? :class: dropdown These structures are an approximation of small molecule intermediates interacting with a much larger nanoparticle surface. Even small nanoparticle catalysts will have many layers to the surface. It is very common to freeze the bottom few layers of the catalyst surface for these simulations so that they do not move and are more similar to the actual much deeper structure. You will notice mostly the small molecule moving during the relaxation, and you may also notice small movements of the top layer of surface atoms as they compensate for the adsorbate. ::: :::{admonition} Why is the adsorbate broken up / appearing on two sides of the cell? :class: dropdown We are working to fix this! ::: --- (ml-questions)= ### ML Questions :::{admonition} What sort of models are used to generate these predictions? :class: dropdown The state-of-the-art in ML for chemistry is moving extremely quickly, and we track this exciting progress through the open leaderboards available at https://opencatalystproject.org/leaderboard.html. The models that are currently available are openly available models that are near the top of the leaderboard and have a desirable compute/accuracy trade-off. We are excited to see developments in many ML model types, but the ones that have worked best for large catalyst datasets tend to be message passing or graph neural networks. ::: :::{admonition} How big are these models? :class: dropdown The models used here have >100 million free fitted parameters that are fit to the OC20 and OC22 datasets. These datasets have >100 million structures, each with O(30) energy/forces, so the number of targets available to fit is quite massive (in the context of work in AI for science). ::: :::{admonition} Can I run these models on my own computer? How long would it take? :class: dropdown Yes! All of the models and pre-trained checkpoints are open source and freely available at [https://facebookresearch.github.io/fairchem/](https://facebookresearch.github.io/fairchem/); models/datasets have varying licenses. They can run on CPUs with ~16Gb of RAM, and CUDA-compatible GPUs with >16Gb of memory. Each energy/force call usually takes O(1s) on a cpu core, and O(50ms) on a GPU, averaged over a reasonable batch size. Of course, this depends on your precise setup and your mileage may vary. ::: :::{admonition} What about the CO2 emissions associated with training and serving ML models? :class: dropdown Very relevant question! Greenhouse gas emissions for training and using large ML models like those used on this website are an active research area. This includes how to measure them, and how to reduce them using methods like inference accelerators or more efficient training strategies. In this case, the models that we are replacing are even more resource intensive -- the density function theory (DFT) calculations that would typically be required for simulations like those shown often require 1000s of core-hours to analyze many possible configurations. The emissions associated with using the ML models here are a small fraction of what would be emitted while running traditional DFT simulations. We also hope that we will see community progress in developing more energy-efficient models with similar (or better) accuracies to the models shown here! ::: :::{admonition} I ran a prediction but the structures/energies look strange to me. What should I do? :class: dropdown Please let us know by posting as a [github issue](https://github.com/facebookresearch/fairchem/issues) with your inputs, results page link, and structure so we can look into the problem on our end. Alternatively, a few things to try: * First, try a couple of the ML models available on this website for the same surface/adsorbate and see if the results differ. This gives you an idea of whether it is specific to a model, or something about the surface/adsorbate that leads to problems. * Second, you can download the structures and try other ML models from [https://facebookresearch.github.io/fairchem/](https://facebookresearch.github.io/fairchem/) to see if the problems are consistent. * Finally, if you have access to VASP you can try running the relaxations yourself to verify the results. ::: :::{admonition} Do you have any estimates on how much I should trust these predictions? Why aren't there error bars? :class: dropdown Uncertainty quantification for large GNN models like those used here is a very active research area, with many different uncertainty calibration metrics and varying additional compute costs for the prediction. * You can get some idea of average MAE from the [OCP leaderboard](https://opencatalystproject.org/leaderboard.html). The models used here have MAEs across many different catalyst/adsorbate systems of ~0.3 eV. For metals and smaller adsorbates, the residuals tend to be smaller. * Deciding which configuration corresponds to the minimum energy tends to be more robust to residuals, so it is likely that the identified structure is identified. Recent results on validation datasets suggest that approximately 50% of the time the energy identified in this process is within 0.05 eV of the DFT-computed minimum (see the [AdsorbML paper](https://arxiv.org/abs/2211.16486)) We are investigating efficient ways to add uncertainty to these predictions; more updates coming soon! ::: --- (catalysis-questions)= ### Catalysis Questions :::{admonition} All I know is the rough composition of the material I'm interested in. How do I select a bulk structure or surface in the website? :class: dropdown Generally, materials with lower formation energies (or smaller distances to the lower hull of the phase diagram) are more likely to be stable or metastable. The Materials Project and other databases have a bunch of great tools to help predict bulk stability (including under reaction conditions!). For surface structure, you have a few options: * You could see how your properties of interest vary across similar compositions or surfaces to get a feel for how surface structure or composition might impact reactivity * You could use DFT or other ML models to predict surface stability to help you select a specific surface * You could find a computational chemistry catalyst friend/collaborator to dig into the surface structure! ::: :::{admonition} How do I use these predictions to understand or predict the activity/selectivity of a catalyst? :class: dropdown The catalyst predictions exposed on this service can be used in many ways; we've tried to highlight a few potential use cases in the [examples and tutorials section](./catalysts/examples_tutorials/summary.md). A few possible use cases: * If you or others have already identified ideal adsorption energies for a particular chemistry, you can use this service to compare the adsorption energies across several facets on the catalyst, or compare different catalyst surfaces. * The adsorption energy on different sites of a particular surface may help you identify which surface is responsible for catalytic activity. * If you know the catalyst surface you are interested in, you could use these predictions to construct reaction energies and a free energy diagram to compare different reaction pathways. You could also do this using your favorite microkinetic modeling package! * You could use these calculations as a starting point for vibration calculations (e.g. IR spectra) or reaction activation barriers (transition states), If you find these useful for your work, we'd love to highlight some additional use cases/stories, so please reach out via [GitHub discussions](https://github.com/facebookresearch/fairchem/discussions)! ::: :::{admonition} Stability/durability is important for my application; can I predict that? :class: dropdown Yes, this is something you can predict, but probably using different ML models than the ones trained here. A few great options: * The formation energy, energy above the phase diagram hull in Materials Project and similar databases are popular descriptors for whether a material is stable, metastable, or unstable. Reaction conditions can be included via techniques like Pourbaix diagrams. * There are many machine learning models available to predict stability for arbitrary materials. The stability of more complex catalyst interfaces is a very interesting research area, especially as some materials are self-passivating and stable even under conditions where they decompose. ::: :::{admonition} I made a catalyst with a different crystal structure than is present in the drop-down list; what should I do? :class: dropdown If you know the crystal structure of composition of your catalyst but it's not present in the drop-down list, it's probably because the structure is either not in the Materials Project, or it's predicted to be unstable by more than 0.1 eV/atom, or our calculations failed when we relaxed the inputs with DFT/RPBE. * You can use the new [Open Catalyst API](./catalysts/examples_tutorials/ocpapi.md) to enumerate surfaces and perform the adsorbate placement using python or your web browser. Make sure you have an RPBE-relaxed structure before starting this process! * You can also do these by hand using the [Open Catalyst Project tools](./catalysts/examples_tutorials/adsorption_energies/adsorption_energies.md) ::: :::{admonition} I think my material is more complex than the surfaces shown here (surface segregation, additional terminations, etc); what should I do? :class: dropdown The models used here may be able to predict the adsorption on more complex surfaces (for example, solid solutions, single atom alloys, segregated materials, etc), but the models have not been validated in these situations. We recommend downloading and using the pre-trained models on your own machine using the ASE calculator interface and using them to predict the adsorption energy ([see the adsorption energies tutorial](./catalysts/examples_tutorials/adsorption_energies/adsorption_energies.md)). If you find the models work well for your application, we'd love to hear from you! And if they don't, feel free to reach out via a [github issue](https://github.com/facebookresearch/fairchem/issues). We're considering allowing predictions on more diverse surfaces, but there are some computational nuances that make doing so a bit difficult. Feel free to reach out via [GitHub discussions](https://github.com/facebookresearch/fairchem/discussions) if you have a specific use case or are interested in getting updates! ::: ::::{admonition} I'm interested in electrochemistry, but I don't see any solvent effects or water layer; what should I do? :class: dropdown It's very common in the electrocatalyst modeling community to relate gas-phase adsorption energies (like those shown here) to adsorption energies in solvent by incorporating a per-adsorbate solvent correction to the adsorption energy. This correction tends to be largest for adsorbates that can hydrogen bond with a water layer (like *OH), or adsorbate that have a strong dipole moment that can be screened with a water layer. This approximation is helpful in screening millions of possible catalyst surfaces, especially because the structure of the water layers on each surface is difficult to predict. Of course, by not including a solvent the models cannot help predict the impact of the solvent on the catalyst activity/selectivity, the role of cation/anion additives in the solvent, or interesting enhancements/disturbances to proton transport, or reactions from the solvent layer to surface adsorbates. These are all exciting research areas! :::{note} Even using DFT, full fidelity detailed modeling of the solvent/catalyst interface in electrochemical conditions is an active research challenge! ::: :::: :::{admonition} I'm interested in trying this across many different surfaces or discover new catalysts but it's hard to select them all on the website; how should I try this? :class: dropdown This website is meant to acquaint you with the state-of-the-art in accelerated catalyst property prediction models, and demonstrate the capabilities of models trained on the Open Catalyst Project datasets. If you want to use these models for high-throughput simulations or programmatically across several materials, you are welcome to: * Use Open Catalyst Project API (python or REST); note that this may be rate-limited * Download and use the Open Catalyst Project models and toolkits (all open source and permissively licensed!) to predict properties on many systems If either of these are unclear or you have trouble using them, please direct questions to the github repo! ::: --- (computational-catalysis-questions)= ### Computational Catalysis Questions :::{admonition} What are the caveats of using these calculations? :class: dropdown As the statistician George Box once said, "*All models are wrong, but some are useful!*" Simulations like these are extremely common in the catalysis community to develop intuition into limitations of specific catalysts, propose new catalyst modifications or active sites, or discover new catalysts with interesting reactivity/selectivity. However, there are a number of caveats to be aware of: * These simulations are based on Density Functional Theory, a quantum mechanical simulation technique that offers reasonably predictive properties, but is nonetheless an approximation. * The DFT settings in the training dataset were chosen for a reasonable combination of speed/accuracy trade-off (kpoints, energy cutoffs, pseudopotential, etc). * All simulations neglected spin polarization (magnetic moments), which often have a small impact on adsorption energies but may be important. Of course, this only applies to spin-polarized elements like Ni, Fe, etc. * The simulations neglect long-range dispersion interactions, which are more important for large adsorbates than small ones. * These simulations approximate a complex catalyst surface with a very small and defect-free representation of an active site. Real catalysts may undergo restructuring or segregation under reaction conditions. Solid solution materials are also not included in databases like the Materials Project so are not shown here. * These simulations do not include solvent interactions, which may be important for your application. However, adsorption energies with and without solvent are often strongly correlated, so this doesn't mean the calculations can't be used for catalysis in condensed phases! * The reaction energies here are enthalpies, and do not include entropic contributions to stability (as they would be if they were the Gibbs free energy). * The ML models shown are imperfect and rapidly developing. The mean absolute error for the OC20 validation datasets is ~0.3-4 eV (less for metals and small molecules, more for complex catalysts like nitrides or larger adsorbates) (see the [OCP leaderboard](https://opencatalystproject.org/leaderboard.html)). The best adsorption site according to DFT is one of the best 5 predicted by these methods about 90% of the time. These caveats are all extremely common in computational catalysis, and are provided here not to scare you but to remind you of possible limitations. If you aren't sure whether these caveats are important for your intended use case, talk to your favorite computational catalyst researcher friend, or post your questions on the discussion board! ::: :::{admonition} I'm interested in a different reaction energy than the ones listed for a particular adsorbate; what do I do? :class: dropdown The reaction energies listed on this website are helpful because they all use a single set of reference gas phase species. This makes it easy to mix and match reaction energies to make new reaction energies, perhaps using gas phase species. For example, adsorption of *NH3 on a surface would show up on this site as: `* + 1/2 N2(g) + 3/2 H2(g) -> *NH3` If you wanted to know the adsorption energy of *NH3 relative to gas phase NH3, we could add this reaction energy with the (known) gas phase formation energy of ammonia: `* + 1/2 N2(g) + 3/2 H2(g) -> *NH3` (E1) `NH3(g) -> 1/2 N2(g) + 3/2 H2(g)` (E2) To yield: `* + NH3(g) -> *NH3` (E1 + E2) Similar approaches can be used to prepare surface reaction energies. ::: :::{admonition} Why are the reaction energies enthalpies and not free energies? How should I calculate dG? :class: dropdown Calculating the free energy of reaction requires several additional steps in a typical computational catalysis workflow: 1. Vibrational calculations to obtain vibrational modes and energies 2. Calculation of the zero point energy (ZPE) and entropic contributions (assuming harmonic or anharmonic vibrational modes) 3. (optional) identification of adsorbate-specific solvent or frequency-dependent corrections 4. Calculating the Gibbs free energy at a specific temperature We expect that the ML models used here can also help predict vibrational modes and are working to validate these approaches. As a placeholder, we have provided rough estimates for Gibbs free energy of reaction for the reactions here by calculating corrections for some representative simple metal surfaces. These free energy corrections are usually fairly constant across different catalyst surfaces, but can change significantly if the binding configuration (mono- or bi-dentate) or active site changes. ::: :::{admonition} Where did the list of bulk materials come from? How did you select them from the Materials Project? :class: dropdown The bulk materials for this demo were prepared by: 1. Selecting all materials from the Materials Project using catalyst elements seen in OC20 2. Filtering the materials to negative formation energies and energies above the phase diagram hull less than 0.1 eV/atom 3. Isotropic relaxations of the structures with DFT and the RPBE functional, which had a small failure rate. We're interested in adding the ability to upload custom bulk structures, but making sure the materials are not strained from the perspective of the RPBE functional is a bit nuanced and it's hard to know how well our models will work on very different structures. If you're interested in this feature, reach out via a [github issue](https://github.com/facebookresearch/fairchem/issues)! ::: :::{admonition} Where did the list of adsorbates come from? How did you select them? :class: dropdown These adsorbates were the original 82 adsorbates used to construct OC20, which allows us to have some intuition and statistics on how well the models may work. We're considering allowing users to upload custom adsorbate structures. If you're interested in this feature, reach out as a [github issue](https://github.com/facebookresearch/fairchem/issues)! ::: :::{admonition} My favorite catalyst composition/surface/adsorbate is missing. What should I do? :class: dropdown This website is meant to acquaint you with the state-of-the-art in accelerated catalyst property prediction models, and demonstrate the capabilities of models trained on the Open Catalyst Project datasets. If you want to use these models for high-throughput simulations or programmatically across several materials, you are welcome to: * Use Open Catalyst Project API (python or REST); note that this may be rate-limited * Download and use the Open Catalyst Project models and toolkits (all open source and permissively licensed!) to predict properties on many systems If either of these are unclear or you have trouble using them, please direct questions to the [GitHub discussions](https://github.com/facebookresearch/fairchem/discussions)! ::: :::{admonition} These are all low coverage energies, but I'm interested in adsorption energies at high coverage. Can we do that too? :class: dropdown This is a very exciting and interesting research area! We expect that OCP models like those shown here will be capable of predicting energies at higher coverages, but this is an ongoing research area. This is especially interesting for adsorbates with long-range interactions like *CO, which has a strong dipole moment. More training data or more sophisticated models may be needed to accurately resolve high-coverage adsorption energies. ::: ::::{admonition} Reaction energies are great, but I'm interested in reaction kinetics (activation energies). Can I use OCP models for kinetics? :class: dropdown This is a very exciting and interesting research area! We expect that OCP models like those shown here will be capable of predicting activation energies and reaction barriers, but this is an ongoing research area. More training data or more sophisticated models may be needed to accurately resolve transition state energies. :::{tip} Check out the [CatTsunami tutorial](./catalysts/examples_tutorials/cattsunami_tutorial.md) for transition state calculations using NEB methods. ::: :::: :::{admonition} What DFT settings should I use to verify the single-points? How would I reproduce these energies with DFT? :class: dropdown The [Open Catalyst Dataset repo](https://github.com/facebookresearch/fairchem/tree/main/src/fairchem/data/oc) can be used to create DFT inputs. Specifically, the DFT settings are provided [here](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/data/oc/utils/vasp.py) to ensure consistency with the underlying OC20 DFT-level theory. ::: :::{admonition} What does the "shift" mean in the surface information? :class: dropdown The shift here is the absolute position of the cut and the resulting termination, and is used internally in [pymatgen](https://pymatgen.org/) for producing the surfaces. For very simple surfaces (e.g. Pt(111) or Cu(211)) there is only one unique surface termination so we don't need to specify the shift. However, more complicated materials such as bimetallic systems may have several unique terminations that we can specify with the shift. It is sometimes possible to identify the best possible termination, but for bimetallic surfaces this is often not obvious so we include all of the possibilities here. ::: :::{admonition} I found a better adsorbate configuration than the one shown on this website; what should I do? :class: dropdown This is expected to happen when one of the following happens: * Our adsorbate placement / guessing strategy isn't exhaustive enough to find the best possible configuration * The trained ML models encounter errors or produce bad relaxations Either way, post on the discussion board so we can improve the models and datasets! ::: --- Source: `docs/videos.md` # Videos Build intuition with introductory videos, then go deeper with technical talks from FAIR Chemistry researchers. ## Why model atoms for climate? This introductory series by Larry Zitnick is designed for viewers without a computational chemistry background. It explains how machine learning can accelerate chemistry research for climate applications. [Watch the introductory series on YouTube](https://www.youtube.com/playlist?list=PLU7acyFOb6DXgCTAi2TwKXaFD_i3C6hSL). ## Technical presentations ::::{grid} 1 :::{card} FAIR Chemistry technical overview An overview of the Open Catalyst Project and FAIR Chemistry framework, including model architectures, training strategies, and catalysis applications. [Watch the technical overview on YouTube](https://www.youtube.com/watch?v=xpgrV0t91AE). ::: :::{card} Research applications and future directions Research applications of FAIR Chemistry models and directions for machine learning in computational chemistry and materials science. [Watch the research applications talk on YouTube](https://www.youtube.com/watch?v=6TunVLsj0YE). ::: :::: ## Read the research Browse [FAIR Chemistry papers](./core/fair_chemistry_papers.md) for the methods, datasets, and applications behind the project. --- Source: `docs/core/fairchemv1_v2.md` # `fairchem>=2.0` Fairchem 2 is the current platform for pretrained UMA models, multi-task training, and scalable atomistic inference. It replaces the version 1 trainer, calculator, and model interfaces with a Hydra-configured system built around a shared backbone and task-aware prediction heads. New projects should use fairchem 2 and the current `FAIRChemCalculator` interface. :::{warning} `fairchem>=2.0` is a major upgrade with completely rewritten trainer, fine-tuning, models, and calculators. The old `OCPCalculator` and trainer code will NOT be revived. ::: We plan to bring back the following models compatible with Fairchem V2 soon: * Gemnet-OC * EquiformerV2 * eSEN Start with the [installation guide](./install.md) and [Hello World](./quickstart.md), then use the [common workflows](./common_tasks/summary.md) for training, evaluation, and scaled inference. ## Using Fairchem V1 If you need to use models from fairchem version 1, you can still do so by installing version 1: ```bash pip install fairchem-core==1.10 ``` And using the `OCPCalculator`: ```python from fairchem.core import OCPCalculator calc = OCPCalculator( model_name="EquiformerV2-31M-S2EF-OC20-All+MD", local_cache="pretrained_models", cpu=False, ) ``` ## Projects and models built on `fairchem` version v2 ::::{grid} 1 :::{card} UMA (Universal Model for Atoms) :link: https://ai.meta.com/research/publications/uma-a-family-of-universal-models-for-atoms/ The latest universal model for atomistic simulations. [[arXiv]](https://ai.meta.com/research/publications/uma-a-family-of-universal-models-for-atoms/) | [[code]](https://github.com/facebookresearch/fairchem/tree/main/src/fairchem/core/models/uma) ::: :::: ## Projects and models built on `fairchem` version v1 :::{note} You can still find these in the v1 version of fairchem github. However, many of these implementations are no longer actively supported. ::: ::::{grid} 1 2 2 3 :::{card} GemNet-dT [[arXiv]](https://arxiv.org/abs/2106.08903) | [[code]](https://github.com/facebookresearch/fairchem/blob/main/src/fairchem/core/models/gemnet) ::: :::{card} PaiNN [[arXiv]](https://arxiv.org/abs/2102.03150) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/core/models/painn) ::: :::{card} Graph Parallelism [[arXiv]](https://arxiv.org/abs/2203.09697) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/core/models/gemnet_gp) ::: :::{card} GemNet-OC [[arXiv]](https://arxiv.org/abs/2204.02782) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/core/models/gemnet_oc) ::: :::{card} SCN [[arXiv]](https://arxiv.org/abs/2206.14331) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/core/models/scn) ::: :::{card} AdsorbML [[arXiv]](https://arxiv.org/abs/2211.16486) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/applications/AdsorbML) ::: :::{card} eSCN [[arXiv]](https://arxiv.org/abs/2302.03655) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/core/models/escn) ::: :::{card} EquiformerV2 [[arXiv]](https://arxiv.org/abs/2306.12059) | [[code]](https://github.com/facebookresearch/fairchem/tree/fairchem_core-1.10.0/src/fairchem/core/models/equiformer_v2) ::: :::{card} SchNet [[arXiv]](https://arxiv.org/abs/1706.08566) ::: :::{card} DimeNet++ [[arXiv]](https://arxiv.org/abs/2011.14115) ::: :::{card} CGCNN [[arXiv]](https://arxiv.org/abs/1710.10324) | [[code]](https://github.com/facebookresearch/fairchem/blob/e7a8745eb307e8a681a1aa9d30c36e8c41e9457e/ocpmodels/models/cgcnn.py) ::: :::{card} DimeNet [[arXiv]](https://arxiv.org/abs/2003.03123) | [[code]](https://github.com/facebookresearch/fairchem/blob/e7a8745eb307e8a681a1aa9d30c36e8c41e9457e/ocpmodels/models/dimenet.py) ::: :::{card} SpinConv [[arXiv]](https://arxiv.org/abs/2106.09575) | [[code]](https://github.com/facebookresearch/fairchem/blob/e7a8745eb307e8a681a1aa9d30c36e8c41e9457e/ocpmodels/models/spinconv.py) ::: :::{card} ForceNet [[arXiv]](https://arxiv.org/abs/2103.01436) | [[code]](https://github.com/facebookresearch/fairchem/blob/e7a8745eb307e8a681a1aa9d30c36e8c41e9457e/ocpmodels/models/forcenet.py) ::: ::::