ml-property-predictor
Installation
SKILL.md
MLIP Property Predictor Training
Goal
To leverage pre-trained GNN representations from MLIPs to train an independent readout head for any custom scalar target property (e.g., bulk modulus, bandgap, formation energy, or spin states) directly from crystal or molecular structures.
Overview
This skill allows you to leverage pre-trained GNN representations from MLIPs to train an independent readout head for any custom scalar target property, such as bulk modulus, bandgap, formation energy, or spin states.
To keep the core MLIP wrappers clean, property prediction in AtomisticSkills is handled by standalone training scripts located in the .agents/skills/ml-property-predictor/scripts/ directory.
Workflow
- Prepare Data: Build a
.jsonor.xyzdataset containing structures and the corresponding scalar property labels. JSON datasets should be lists of dicts containing astructurekey (Pymatgen format) and your target property key. - Determine Property Type: Determine if the property is
"intensive"(e.g. Bandgap, Bulk Modulus) or"extensive"(e.g. Total Energy).- MatGL logic: For intensive targets, node features undergo a global graph readout (like
Set2Set) before passing through an MLP. For extensive targets, the MLP outputs atomic properties which are then sum-pooled. - MACE logic: MACE natively supports extensive targets by predicting site-wise scalar outputs and sum-pooling them. When intensive properties are targeted, MACE still sum-pools site-wise outputs, forcing the model to internally learn the intensive invariant.
- MatGL logic: For intensive targets, node features undergo a global graph readout (like
- Execute Script: Run the MACE or MatGL property prediction script in their respective Conda environments.