← Back to jobs

Prophecy Technologies Logo
Senior AWS Spark Performance Engineer

Prophecy Technologies

 

Tampa, FL, USA

Posted On: 14 days ago
Experience: 6+ years
Availability: Hybrid
Openings: 1
Category: Performance Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Role Overview:

This role focuses on optimizing AWS EMR clusters and Spark applications to enhance performance, reduce costs, and improve resource utilization. The successful candidate will analyze cluster metrics, tune configurations, and implement solutions to ensure efficient data processing and operational excellence.

Key Responsibilities:

  • Analyze EMR cluster metrics, Spark application telemetry, and YARN resource utilization to identify over-provisioned memory allocations, underutilized executors, and suboptimal cluster configurations that contribute to excessive DRAM consumption.
  • Recommend and implement cluster-level optimizations, including instance family right-sizing (e.g., migrating from memory-optimized R-type to compute-optimized C-type instances), node count adjustments, EBS volume configurations, and spot/on-demand fleet composition changes.
  • Tune Spark runtime configurations at the cluster level, including executor memory/core ratios, YARN container sizing, dynamic resource allocation settings, memory overhead parameters, and shuffle service configurations, to achieve optimal memory utilization without impacting job SLAs.
  • Perform custom operations and iterative experiments using Amazon internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results.
  • Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders.
  • Monitor service health metrics and troubleshoot operational issues during and after optimization activities, ensuring zero degradation to job completion times, data processing throughput, and downstream SLAs.
  • Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture cluster optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale.
  • Extract scalable learnings from optimization engagements and develop programmatic frameworks that enable the initiative to scale across hundreds of EMR clusters, including training and enabling other vendor engineers to execute optimization playbooks.

Required Skills:

  • Digital: Amazon Web Service (AWS) Cloud Computing
  • Advanced Java Concepts
  • Microsoft SQL Server 2019
  • Java Performance Tools (Jprobe, Jensor, OptimizeIT)
  • AWS EMR Cluster Operations
  • Spark & YARN Tuning
  • Memory Optimization & Capacity Analysis
  • EC2 Right-Sizing & Cost Optimization
  • Monitoring & Troubleshooting (CloudWatch/Logs)

Qualifications:

  • 6+ years experience required

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs