← Back to jobs

Synechron Logo
PySpark Data Engineer

Synechron

 

Bengaluru, Karnataka, India

Posted On: 30+ days ago
Experience: 7+ years
Availability: Hybrid
Openings: 1
Category: PySpark Data Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

Required Skills:

 

  • Proven expertise in Python programming, emphasizing clean, maintainable, and scalable code
  • Hands-on experience with PySpark in both batch and streaming workflows
  • Deep knowledge of data manipulation and feature engineering, including Pandas, NumPy, and visualization libraries (matplotlib, seaborn)
  • Experience with Spark components like Spark SQL, DataFrames, and Spark MLlib
  • Familiarity with data storage solutions: SQL and NoSQL databases (e.g., Hive, Cassandra)
  • Knowledge of ETL tools such as Apache Airflow, Jenkins, or GithHub Actions for scheduling and automation
  • Experience working with cloud environments, especially Azure or AWS for big data processing

     

Preferred Skills:

 

  • Hands-on with containerization and orchestration (Docker, Kubernetes)
  • Exposure to distributed storage solutions like Hadoop HDFS or Azure Data Lake

     

Overall Responsibilities

 

  • 5 years of experience in Design, develop, and optimize large-scale data pipelines using PySpark for structured, semi-structured, and unstructured data
  • 5 years of experience to Lead the building of ML pipelines for training, validation, and deployment of models in streaming/batch modes
  • Write high-quality, efficient code that supports data transformation, cleaning, and feature engineering
  • Collaborate with data scientists, analysts, and stakeholders to understand data requirements and deliver actionable insights
  • Build and maintain reusable code base and automation scripts for data processing and model validation
  • Monitor pipeline performance, troubleshoot issues, and implement improvements to ensure robustness and scalability
  • Stay up-to-date with the latest in big data processing, ML techniques, and analytics tools to improve system efficiency and analytics capabilities
  • Technical Skills (By Category)

    Programming Languages:

     
  • Required: Python (required), PySpark (required)
  • Preferred: Scala, Java

     

Databases & Data Management:

 

  • SQL (MySQL, SQL Server), NoSQL (Cassandra, MongoDB), Hive, Data Lakes

     

Cloud Technologies:

 

  • Azure Data Factory, Azure Synapse, AWS Glue, S3 (preferred)

     

Frameworks & Libraries:

 

  • Spark MLlib, Pandas, NumPy, seaborn, matplotlib, scikit-learn (preferred)

     

Development Tools & Methodologies:

 

  • Jupyter, PyCharm, VSCode, Git, CI/CD (Jenkins, GitHub Actions), Airflow

     

Security & Data Governance:

 

  • Data privacy principles, secure data ingestion and output, compliance

     

Experience Requirements

 

  • 7-12 years of experience in data engineering, analytics, or data science roles, with significant hands-on experience in big data processing and ML pipelines
  • Proven track record of building scalable data pipelines and supporting ML workflows in enterprise environments
  • Experience working with structured, semi-structured, and unstructured data across financial domains
  • Previous leadership or mentorship experience in a technical team is preferred

     

Day-to-Day Activities

 

  • Develop and optimize data pipelines for financial and index data using PySpark and related tools
  • Build ML workflows, feature engineering, and model deployment pipelines in both streaming and batch environments
  • Collaborate with business analysts and data scientists to refine data requirements and deliver insights
  • Automate data ingestion, transformation, and validation processes
  • Monitor system performance, troubleshoot issues, and implement tuning activities
  • Review code and pipeline health with peer teams, uphold best practices in software development and data security

     

Qualifications

 

  • Bachelor’s or Master’s degree in Computer Science, Data Science, Mathematics, or a related field
  • Relevant certifications in big data, cloud platforms, or analytics (preferred)
  • Strong portfolio showcasing data pipeline projects, analytics solutions, and ML workflows

Education

Bachelor's or Master's degrees

Related Jobs

No related jobs found

← Back to jobs