← Back to jobs

Prophecy Technologies Logo
Principal DevOps SRE Architect

Prophecy Technologies

 

Austin, TX, USA

Posted On: 30+ days ago
Experience: 8+ years
Availability: Hybrid
Openings: 1
Category: DevOPs SRE
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Role Overview:

This role is for a Principal DevOps, SRE & Application Infrastructure Architect, with a specific focus on candidates with ex-Apple experience. The position involves designing, deploying, and optimizing infrastructure, CI/CD pipelines, and ensuring application reliability and performance.

Key Responsibilities:

  • Infrastructure & GitOpsK8s & Containerization: Design, deploy, and optimize secure Docker/Kubernetes (AKS) environments using Helm and ArgoCD.
  • Networking & Edge: Manage cloud Ingress, Load Balancers, and end-to-end certificate management (SSL/mTLS).
  • CI/CD & Automation: Automate tasks with Shell/Python; build GitOps pipelines and manage schema migrations via Flyway.
  • SRE & Observability Reliability: Own end-to-end production availability and performance; define and track SLAs/SLOs/SLIs and error budgets.
  • Telemetry: Build observability stacks using OpenTelemetry, Prometheus, Grafana, and Splunk.
  • Incident Management: Lead P0/P1 incident response, deep-dive distributed system debugging, RCAs, and on-call rotations.
  • Application & Database Operations Polyglot DB Management: Design and operate high-availability cloud database infrastructure (Oracle, Postgres, Cassandra, Couchbase, Redis, CockroachDB).
  • Data Replication & DR: Manage Oracle GoldenGate replication, patching, purging, and execute robust P0 Disaster Recovery/failover strategies.
  • App Support: Perform deep-dive troubleshooting within Java application layers, gRPC, REST/HTTP/JSON, and caching/messaging systems (SNS/SQS, Elasticsearch, Solr).

Required Technical Skills

  • Orchestration & DevOps: Kubernetes (AKS), Docker, Helm, ArgoCD, Flyway.
  • Scripting & OS: Strong Linux/Unix internals, Shell scripting, and Python.
  • Observability: OpenTelemetry, Prometheus, Grafana, Splunk.
  • Database & Replication: SQL/PL-SQL (procedures, triggers, tuning), GoldenGate, NoSQL (Cassandra, Couchbase, Redis), and Transactional DBs (Oracle, Postgres).
  • App Troubleshooting: Java application debugging, gRPC, REST, and cloud-native caching/queues.

Preferred Skills:

  • Exposure to AliCloud or multi-cloud environments.
  • Security operations (automated password rotations, IAM privilege management).
  • Strong cross-functional collaboration with security, network, and application teams.

Essential Skills:

  • Devops
  • Digital : DevOps
  • Digital : Site Reliability Engineering (SRE)

Qualifications:

  • 8-10 years of experience.
  • Architect level profiles are sought

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs