Description
Key Skills: Kafka, NoSQL, Cassandra, Spark, Scala, Hadoop, MongoDB, ETL, SQL, Distributed Systems
Good to Have Skills: Strong expertise in Hadoop ecosystem (Hive, YARN, HDFS), Linux (on prem), CI-CD (Jenkins, Github), Observability tools (Dynatrace), scripting with Pandas, Pola.rs, PySpark, Ibis, understanding of system architecture concepts including caching, replication, data modeling, experience with cloud platforms (Azure), orchestration tools (NiFi, Griffin, Hamilton, Airflow), functional experience in Retail Banking, capital markets, securities processing.
Roles & Responsibilities:
- Design and develop streaming pipelines and data platforms to handle high volume data processing for both batch and realtime systems.
- Design and implement real-time streaming pipelines using Kafka for on-premises systems with high throughput and fault tolerance.
- Develop producers and consumers for high-throughput, fault-tolerant systems ensuring optimal performance and reliability.
- Implement Kafka-based event-driven architectures to support distributed system requirements and data processing workflows.
- Lead technical initiatives spending 70% of time in design, architecture, and coding activities for complex data solutions.
- Provide team mentoring, guidance, and coordination activities for 30% of time to support team development.
- Take responsibility for development and provide strong expertise to debug and fix post production issues.
- Work in Agile environments and collaborate effectively with cross-functional teams to deliver high-quality solutions