Work on complex, major, or highly visible tasks in support of multiple projects requiring multiple areas of expertise. - Provide subject matter expertise in managing Hadoop and Data Science Platform operations, focusing on Cloudera Hadoop, Jupyter Notebook, OpenShift, Docker-Container Cluster Management, and Administration. -
Integrate solutions with other applications and platforms outside the framework.
Manage day-to-day operations for platforms built on Hadoop, Spark, Kafka, Kubernetes/OpenShift, Docker/Podman, and Jupyter Notebook.
Support and maintain AI/ML platforms such as Cloudera, DataRobot, C3 AI, Panopticon, Talend, Trifacta, Selerity, ELK, KPMG Ignite, and others.
Automate platform tasks using tools like Ansible, shell scripting, and Python.
Must Haves:
Strong knowledge of Hadoop Architecture, HDFS, Hadoop Cluster, and Hadoop Administrator's role
Intimate knowledge of fully integrated AD/Kerberos authentication.
Experience setting up optimum cluster configurations
Expert-level knowledge of Cloudera Hadoop components such as HDFS, Sentry, HBase, Kafka, Impala, SOLR, Hue, Spark, Hive, YARN, Zookeeper, and Postgres.
Hands-on experience analyzing various Hadoop log files, compression, encoding, and file formats