ml-fairchem-finetune
Installation
SKILL.md
Fairchem Fine-tuning
Goal
To evaluate and improve the accuracy of a foundation Fairchem potential (e.g., UMA, ESEN) for a specific chemical system or physical property using the provided Python fine-tuning script.
Instructions
- Prepare Labeled Dataset: Obtain diverse structures with high-fidelity labels (energy, forces, stress). See the
/benchmark-finetuningworkflow for details. - Custom Data Conversion: Read the source data format and write a customized conversion script if needed, formatting it for the subsequent preparation step.
- Benchmarking: Predict results on the new labels and benchmark the foundation model using ml-mlip-benchmark.
- Data Preparation: Execute
scripts/prepare_fairchem_data.pyto convert JSON structures to extxyz, generate native LMDB databases, compute dataset references, and configure a templateduma_sm_finetune_template.yaml. - Fine-Tuning: Execute
fairchem -c uma_sm_finetune_template.yaml job.run_dir=XXXnatively. - Validation: Run
scripts/extract_fairchem_logs.pyto extract curves and verify convergence against the benchmarked foundation metrics. - Registration: Use the
register_modeltool to register the newly fine-tuned model checkpoint into the local registry so future research tasks can discover and reuse it.
Training Configuration
Fairchem fine-tuning relies heavily on the fairchem CLI, which uses Hydra for configuration. The script scripts/prepare_fairchem_data.py bridges standard data into the complex Fairchem directory structure and generates .aselmdb dataset formats automatically.