← Back to jobs
San Jose, CA, USA
No related jobs found
Job Description:
Hands-on Linux Systems/Site Reliability Engineer responsible for maintaining the availability, reliability, performance, and operational health of production compute and API services. The role focuses on RHEL/Linux administration, production troubleshooting, incident response, automation, monitoring, and operational tooling. The engineer will diagnose infrastructure and service issues independently while improving scalability, reliability, and operational efficiency.
Experience:
· 4–8 years of relevant Linux administration, production operations, SRE, or infrastructure support experience.
· Hands-on experience supporting mission-critical Tier-1 production services.
· Experience with incident response, pager/on-call support, debugging, and root cause analysis.
Core Skills:
· Operating Systems: RHEL 7, RHEL 8, RHEL 9, Linux, Unix
· System Administration: Linux boot process, BIOS, UEFI, systemd, systemctl, rescue mode, emergency mode, root-password recovery
· Storage & Filesystems: Filesystems, disk utilization, inodes, SWAP, permanent mounts, /etc/fstab, Ext4, XFS
· Processes & Networking: Memory and process analysis, zombie processes, orphan processes, listening ports, service troubleshooting
· Programming & Scripting: Python, Bash, JavaScript
· Monitoring & Observability: Service metrics, dashboards, KPIs, alarms
· Production Operations: Incident triage, ticket management, runbooks, root cause analysis, operational toil reduction
· Engineering: Services, operational tools, CI/CD, scalability, reliability, API availability
Key Responsibilities:
· Maintain the operational health, availability, reliability, and low latency of core compute and API services.
· Administer and troubleshoot RHEL/Linux systems in production environments.
· Manage and triage incidents and tickets based on business and service impact.
· Diagnose filesystem, disk, memory, process, boot, service, and network-related issues.
· Build automation and operational tooling using Python, Bash, or JavaScript.
· Develop dashboards, service KPIs, monitoring systems, and actionable alerts.
· Create and automate frequently used runbooks to reduce incident triage time and operational toil.
· Collaborate with developers to improve system scalability, reliability, and development velocity.
· Participate in on-call support, incident response, debugging, and root cause analysis
Any Graduate
No related jobs found
← Back to jobs