← Back to jobs

GSK Solutions Inc Logo
AWS DevOps Engineer / AI Platform Engineer

GSK Solutions Inc

 

Reading, PA, USA

Posted On: 3 days ago
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: AWS DevOps Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Requirements

Experience: 8+ years in Platform Engineering, DevOps, or Site Reliability Engineering (SRE).

Cloud Expertise: Deep proficiency in AWS (IAM, CloudWatch, Bedrock, Lambda).

Observability Tools: Proven experience with Dynatrace, Jaeger, or Honeycomb, and distributed tracing standards.

AI/LLM Interest: Familiarity with the LLM lifecycle, including prompt execution, token usage, and frameworks like LangChain or AgentCore.

Automation: Advanced experience with Terraform and CI/CD pipeline design.

Collaboration: Experience working in an Agile environment with integrated tools like Microsoft Teams and Confluence.

 

User this when submitting candidates:

Also please check if the next candidate has some experience with at least 50% of below items:

Implementation of Agents on Agentcore runtime

Implementation of Agentic SDLC in Agentcore

Understanding of Strands or any other Agentic AI framework like Langraph, Langchain or Crew AI

Implementation of Bedrock Knowledge Base

Implementation of Knowledge Graph

Implementation of MCP servers in Agentcore

Implementation of Agentcore Gateway

Implementation of Agentcore Identity

Implementation of Agentic AI Observability

Implementation of Agentcore Evaluations

Implementation of AWS Bedrock

Implementation of AWS Bedrock Inference Profile

Implementation of AWS Sagemaker

AWS Services (Cloud) in General

Terraform

 

Deliverables

Observability

Assess CloudWatch, X-Ray, Bedrock logging, AgentCore traces vs. agentic workflow requirements; produce gap analysis, Setup observability in Dynatrace

Design post-deployment validation pipeline for agents & MCP servers (deployment health + tool registration checks)

Implement distributed tracing & structured logging: LLM decisions, tool selections, sub-agent calls, MCP interactions

Evaluate LangFuse / LiteLLM proxy vs. AWS-native; deliver target-state observability architecture recommendation

 

Cost Tracking & TCO

Extend tagging taxonomy to cover agent runtimes, MCP servers, vector DBs, Bedrock token consumption per namespace

Design cost visibility model: aggregate agent, MCP, vector DB, and Bedrock token costs per team/department

Build CloudWatch (or equivalent) dashboards for per-team spend; configure AWS Budgets with alerting thresholds

Automate cost reports delivered via email / Microsoft Teams; implement anomaly detection rules

 

Monitoring & Alerting

Define P1–P4 alerting rules: deployment failures, runtime errors, tool invocation failures, MCP connectivity issues

Integrate alert notifications to Microsoft Teams channels and email; route by resource ownership tags

Author runbooks linked to every alert; publish in Confluence for developer self-service resolution

Evaluate AWS-native vs. third-party monitoring stack; deliver recommendation aligned to observability architecture

 

Security & Access Control

Assess current IAM + tagging approach for multi-team isolation; identify scalability gaps and risks

Evaluate Cedar policy engine (AgentCore) for fine-grained tool access control; document enterprise-scale gaps

Design scalable ABAC-based identity model for multi-team isolation without IAM policy sprawl; deliver Terraform modules

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs