Description
You will maintain global infrastructure uptime and manage cloud health through proactive monitoring and incident response.
Responsibilities
- Monitor cloud health and provide first-response to issues impacting customer experience.
- Troubleshoot complex system, networking, and storage problems ranging from single droplets to cloud-wide disturbances.
- Automate manual processes and build tools to improve operational efficiency.
- Coordinate operational tasks across teams to improve the platform with minimal service impact.
Required Skills
- 5+ years of experience in system administration or cloud operations.
- Deep proficiency with Linux operating systems.
- Strong networking knowledge, specifically IPv4 troubleshooting (CCNA equivalent).
- Experience with troubleshooting virtual machine instances and virtualization technologies.
- Familiarity with containerization technologies and container troubleshooting.
- Experience with monitoring systems and incident management workflows.
- Proficiency in scripting with Bash, Python, or Ruby.
- Experience using configuration management systems.
- Foundational knowledge of storage concepts and technologies.
Preferred Skills
- Commitment to maintaining clear, technical documentation.