Description
You will design, develop, and maintain scalable data pipelines and enterprise data integration solutions.
This role is remote.
Responsibilities
- Build and optimize batch and real-time streaming data pipelines using distributed processing frameworks.
- Design cloud-native data platforms and implement data warehouses, data lakes, and lakehouse architectures.
- Handle structured, semi-structured, and unstructured data using various file formats and NoSQL/relational databases.
- Implement CI/CD pipelines, DevOps practices, and infrastructure automation for data engineering workflows.
- Ensure data quality, governance, security, and observability across distributed systems.
Required Skills
- 4+ years of experience in data engineering, ETL/ELT workflows, and scalable data architectures.
- Proficiency in Python, SQL, PySpark, Spark SQL, and Scala.
- Experience with big data technologies including Apache Spark, Databricks, Hadoop, Airflow, and Kafka.
- Strong understanding of cloud platforms (AWS, Azure, or GCP) and services like S3, Glue, Redshift, or Azure Data Factory.
- Experience with database technologies such as PostgreSQL, MySQL, Oracle, MongoDB, Cassandra, and DynamoDB.
- Familiarity with CI/CD tools like GitLab CI/CD, Jenkins, or GitHub Actions.
- Knowledge of data modeling techniques including star schemas, snowflake schemas, and dimensional models.
Preferred Skills
- Experience with streaming technologies like Flink or Kinesis.
- Understanding of partitioning strategies and performance tuning for distributed systems.