Minimum 4+ years of experience in designing, developing, and maintaining scalable data pipelines, ETL/ELT workflows, and enterprise-grade data integration solutions.
Expertise in Python, SQL, PySpark, Spark SQL, Scala, and distributed data processing frameworks for building high-performance data platforms.
Experience with Apache Spark, Databricks, Hadoop, Airflow, Kafka, Snowflake, and modern big data technologies for large-scale data engineering and analytics.
Experience building cloud-native data solutions using AWS, Azure, or GCP, including services such as S3, Glue, Redshift, Athena, EMR, Kinesis, ADF, Synapse, ADLS, Databricks, and Event Hubs.
Understanding of data warehousing, dimensional modeling, Star Schema, Snowflake Schema, Data Lakes, and Lakehouse architectures for enterprise reporting and analytics.
Experience developing and optimizing batch and real-time streaming data pipelines using Kafka, Spark Streaming, Apache Flink, Kinesis, or similar event-driven technologies.
Experience working with structured, semi-structured, and unstructured data, including Parquet, Avro, ORC, CSV, JSON, and other enterprise data formats.
Experience with relational and NoSQL databases including PostgreSQL, MySQL, Oracle, SQL Server, MongoDB, Cassandra, DynamoDB, and database performance optimization.
Experience implementing CI/CD pipelines, DevOps practices, version control, and automation using Git, Jenkins, GitHub Actions, Azure DevOps, GitLab CI/CD, or similar tools.
Knowledge of performance tuning, query optimization, partitioning, indexing, data quality validation, monitoring, troubleshooting, and scalable distributed computing.
Experience integrating enterprise systems using REST APIs, Microservices, Event-Driven Architecture, Data Integration Patterns, and secure data exchange mechanisms