← Back to jobs

Centraprise Logo
Linux Systems Engineer/ SRE (RHEL, Linux, SRE)

Centraprise

 

San Jose, CA, USA

Posted On: 4 days ago
Experience: 5+ years
Availability: Remote
Openings: 1
Category: Senior Linux Systems Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Job Description:

Hands-on Linux Systems/Site Reliability Engineer responsible for maintaining the availability, reliability, performance, and operational health of production compute and API services. The role focuses on RHEL/Linux administration, production troubleshooting, incident response, automation, monitoring, and operational tooling. The engineer will diagnose infrastructure and service issues independently while improving scalability, reliability, and operational efficiency.

Experience:
·       4–8 years of relevant Linux administration, production operations, SRE, or infrastructure support experience.
·       Hands-on experience supporting mission-critical Tier-1 production services.
·       Experience with incident response, pager/on-call support, debugging, and root cause analysis.

Core Skills:
·       Operating Systems: RHEL 7, RHEL 8, RHEL 9, Linux, Unix
·       System Administration: Linux boot process, BIOS, UEFI, systemd, systemctl, rescue mode, emergency mode, root-password recovery
·       Storage & Filesystems: Filesystems, disk utilization, inodes, SWAP, permanent mounts, /etc/fstab, Ext4, XFS
·       Processes & Networking: Memory and process analysis, zombie processes, orphan processes, listening ports, service troubleshooting
·       Programming & Scripting: Python, Bash, JavaScript
·       Monitoring & Observability: Service metrics, dashboards, KPIs, alarms
·       Production Operations: Incident triage, ticket management, runbooks, root cause analysis, operational toil reduction
·       Engineering: Services, operational tools, CI/CD, scalability, reliability, API availability

Key Responsibilities:
·       Maintain the operational health, availability, reliability, and low latency of core compute and API services.
·       Administer and troubleshoot RHEL/Linux systems in production environments.
·       Manage and triage incidents and tickets based on business and service impact.
·       Diagnose filesystem, disk, memory, process, boot, service, and network-related issues.
·       Build automation and operational tooling using Python, Bash, or JavaScript.
·       Develop dashboards, service KPIs, monitoring systems, and actionable alerts.
·       Create and automate frequently used runbooks to reduce incident triage time and operational toil.
·       Collaborate with developers to improve system scalability, reliability, and development velocity.
·       Participate in on-call support, incident response, debugging, and root cause analysis

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs