Description
You will be responsible for maintaining and automating the network infrastructure as a Site Reliability Engineer.
Responsibilities
- Manage and troubleshoot network connectivity issues using TCP/IP, routing, DNS, and firewalls.
- Administer and maintain Linux systems in a production environment.
- Develop and implement automation scripts using Python and Bash.
- Build and maintain CI/CD pipelines for infrastructure changes.
- Monitor system health and performance using observability tools like Prometheus, Grafana, and ELK.
Required Skills
- 5+ years of professional experience in an infrastructure or network role.
- Strong fundamentals in TCP/IP, Routing, DNS, and Firewalls.
- Proficiency in Linux administration and troubleshooting.
- Hands-on experience with scripting languages, specifically Python and Bash.
- Experience with CI/CD practices.
- Familiarity with observability stacks (Prometheus, Grafana, ELK).
- Exposure to AWS Cloud and container orchestration (Docker/Kubernetes).