Description
You will build and manage data infrastructure supporting large-scale data operations.
Responsibilities
- Design, build, and maintain data pipelines across on-prem Hadoop and AWS.
- Develop and maintain Java applications, utilities, and data processing libraries.
- Manage and enhance internal Java libraries for ingestion, validation, and transformation.
- Develop and maintain Airflow DAGs for orchestration and scheduling.
- Build and optimize Spark / PySpark jobs for large-scale data processing.
Required Skills
- 5+ years of professional experience in data engineering.
- Expertise in Java development and build tools like Gradle.
- Strong proficiency with AWS services (S3, Glue, EMR, Athena, EKS basics).
- Experience with Hadoop/HDFS, Hive, and SQL.
- Proficiency in Apache Kafka for streaming ingestion.
- Hands-on experience with Apache Spark / PySpark for batch and streaming processing.
- Experience developing and maintaining Apache Airflow DAGs.
- Familiarity with Git and CI/CD workflows.
- Knowledge of observability tools such as Prometheus/Grafana.