Description
You will design and build scalable batch and real-time data pipelines for an internal observability and analytics platform.
This role is hybrid.
Responsibilities
- Design and build scalable data pipelines across structured and unstructured sources.
- Develop ETL/ELT workflows using AWS Glue, PySpark, or Airflow to ensure data quality, lineage, and reconciliation.
- Integrate analytics and observability services with upstream annotation tools and downstream ML validation systems.
- Implement observability pipelines and alerts for mission-critical metrics.
- Collaborate with product, platform, and analytics teams to define event models, metrics, and data contracts.
Required Skills
- 3–8 years of experience in data engineering or backend development in data-intensive environments.
- Proficiency in Python and SQL.
- Strong experience with cloud-native data tools: AWS (S3, Lambda, Glue, Kinesis, Firehose, RDS, Athena, QuickSight).
- Familiarity with distributed processing frameworks: PySpark, Apache Hadoop, Apache Spark.
- Experience with data lake and warehouse patterns (Delta Lake, Redshift, Snowflake).
- Working knowledge of messaging frameworks like Kafka and Firehose.
- Solid understanding of data modeling, schema design, and versioned datasets.
- Experience with CI/CD, Git, Jira, and Agile methodologies.
Preferred Skills
- Experience with observability/monitoring systems (Prometheus, Grafana, OpenTelemetry).
- Familiarity with data governance, RBAC, PII redaction, or compliance in analytics platforms.