ml-fairchem-finetune

Installation
SKILL.md

Fairchem Fine-tuning

Goal

To evaluate and improve the accuracy of a foundation Fairchem potential (e.g., UMA, ESEN) for a specific chemical system or physical property using the provided Python fine-tuning script.

Instructions

  1. Prepare Labeled Dataset: Obtain diverse structures with high-fidelity labels (energy, forces, stress). See the /benchmark-finetuning workflow for details.
  2. Custom Data Conversion: Read the source data format and write a customized conversion script if needed, formatting it for the subsequent preparation step.
  3. Benchmarking: Predict results on the new labels and benchmark the foundation model using ml-mlip-benchmark.
  4. Data Preparation: Execute scripts/prepare_fairchem_data.py to convert JSON structures to extxyz, generate native LMDB databases, compute dataset references, and configure a templated uma_sm_finetune_template.yaml.
  5. Fine-Tuning: Execute fairchem -c uma_sm_finetune_template.yaml job.run_dir=XXX natively.
  6. Validation: Run scripts/extract_fairchem_logs.py to extract curves and verify convergence against the benchmarked foundation metrics.
  7. Registration: Use the register_model tool to register the newly fine-tuned model checkpoint into the local registry so future research tasks can discover and reuse it.

Training Configuration

Fairchem fine-tuning relies heavily on the fairchem CLI, which uses Hydra for configuration. The script scripts/prepare_fairchem_data.py bridges standard data into the complex Fairchem directory structure and generates .aselmdb dataset formats automatically.

Basic Arguments (Data Prep Script)

Installs
5
GitHub Stars
172
First Seen
Jun 19, 2026