You will develop and maintain Java-based applications for big data processing, building and optimizing data pipelines using Apache Spark for both batch and streaming workloads. You will collaborate with data engineers and architects to integrate these solutions into enterprise systems, ensuring scalability, performance, and reliability of distributed applications. Your work includes writing clean, efficient code, troubleshooting Spark jobs, and resolving issues in data workflows.
Responsibilities
- Develop and maintain Java-based applications for big data processing.
- Build and optimize data pipelines using Apache Spark for batch and streaming.
- Integrate Spark solutions into enterprise systems alongside data engineers and architects.
- Ensure scalability, performance, and reliability of distributed applications.
- Troubleshoot and resolve issues in Spark jobs and data workflows.
Required Skills
- Core Java and Java 8+ proficiency.
- Hands-on experience with Apache Spark (RDDs, DataFrames, Spark SQL, Streaming).
- Knowledge of Hadoop ecosystem components (HDFS, Hive, Kafka).
- Experience with Microservices architecture and containerization (Docker/Kubernetes).
- Familiarity with SQL/NoSQL databases for data integration.
- 6+ years of working experience.
Preferred Skills
- Exposure to cloud platforms (AWS, Azure, GCP).
- Experience with real-time data processing and ETL pipelines.