← Back to jobs
Irvine, CA, USA
No related jobs found
Containerized Python/PySpark applications that read source data from S3, apply territory and effective-dating business logic, and publish results downstream. Anaplan authentication and bulk import/export APIs, file formatting and upload, and reconciliation of loaded data against source. Author and maintain DAGs, custom operators, and environment-specific deployment configuration across environments. AWS S3, Delta/Parquet datasets, and PostgreSQL RDS with schema migrations. Docker image builds, Kubernetes workloads, and multi-environment promotion pipelines. Respond to pipeline failures — Spark memory and performance issues, schema drift, and reconciliation mismatches. Runbooks and lineage documentation." SKILLS_REQUIRED "• Python — production application code (modular, config-driven, unit tested). • Apache Spark / PySpark — building and tuning distributed transformations, including diagnosing memory and performance problems. • Apache Airflow — DAG authoring, scheduling and dependency design, debugging failed runs. • SQL and relational modeling — including schema migrations and data reconciliation. • AWS S3. • Production support — troubleshooting data pipeline failures and root-cause analysis." ESSENTIAL_SKILLS 10+ Years of Python and data engineering epx. DESIRABLE_SKILLS "• Python — production application code (modular, config-driven, unit tested). • Apache Spark / PySpark — building and tuning distributed transformations, including diagnosing memory and performance problems. • Apache Airflow — DAG authoring, scheduling and dependency design, debugging failed runs. • SQL and relational modeling — including schema migrations and data reconciliation. • AWS S3. • Production support — troubleshooting data pipeline failures and root-cause analysis
Bachelor's degree
No related jobs found
← Back to jobs