data-engineering
Installation
SKILL.md
Data Engineering
Workflow
- Identify sources, destinations, schema contracts, freshness needs, volume, and ownership.
- Make ingestion, transformation, validation, and publishing steps explicit.
- Prefer incremental, idempotent pipelines with replay and backfill paths.
- Preserve lineage, timestamps, partitioning, and audit fields.
- Separate raw, staged, transformed, and serving layers when the project uses them.
Data Quality
- Validate row counts, uniqueness, nullability, accepted values, referential integrity, and time windows.
- Handle late-arriving data, duplicates, schema drift, and partial failures.
- Use warehouse-native operations for large data when possible.
- Keep PII handling and retention rules visible.