spark-data-engineering-pipeline
Installation
SKILL.md
Spark Data Engineering Pipeline
Skill by ara.so — Data Skills collection.
Overview
A production-grade ETL pipeline built with PySpark that extracts JSON data from AWS S3, performs schema validation and data quality checks, transforms the data, writes it as Parquet, and loads it into PostgreSQL. This project demonstrates end-to-end data engineering best practices including logging, testing, and CI/CD.