Responsible for implementation and ongoing administration of Hadoop infrastructure
Align with the systems engineering team to propose and deploy new hardware and software environments required for Hadoop, and to expand existing environments
Work with data delivery teams to set up new Hadoop users — including Linux users, Kerberos principals, and testing HDFS, Hive, Pig, and MapReduce access
Perform cluster maintenance and node creation/removal using tools such as Ganglia, Nagios, Cloudera Manager Enterprise, Dell OpenManage, and similar
Performance tune Hadoop clusters and MapReduce routines, including Spark performance tuning
Screen Hadoop cluster job performance and support capacity planning
Monitor Hadoop cluster connectivity and security
Manage and review Hadoop log files
File system management and monitoring
HDFS support and maintenance
Provide cluster support across all data domains for data engineering jobs and warehouse queries, both on-premises and in the cloud
Support installation and proof-of-concept (POC) evaluation for new tools
Team closely with infrastructure, network, database, application, and business intelligence teams to ensure high data quality and availability
Collaborate with application teams to install OS and Hadoop updates, patches, and version upgrades as needed
Serve as point of contact for vendor escalation
DBA Responsibilities Performed by This Role
Data modeling, design, and implementation based on recognized standards
Software installation and configuration
Database backup and recovery
Database connectivity and security
Performance monitoring and tuning
Disk space management
Software patches and upgrades
Automation of manual tasks
Tools & Technical Skills
Linux
Toad, Starburst, Spark, Hadoop, S3, Nifi
Cloud platform experience (AWS, Azure, or equivalent)
SQL
Certification from AWS, Snowflake, Azure, or equivalent preferred