Description
Core Responsibilities
AI Guardrails & Enforcement Architecture
- Design and implement guardrail, validation, filtering, and enforcement components for AI-enabled applications.
- Determine where controls should be applied across service boundaries, including:
- Input validation
- Prompt-injection detection and mitigation
- Model input controls
- Tool-use authorization
- Output validation and filtering
- Policy enforcement
- Failure and fallback handling
- Establish reusable enforcement patterns that can be adopted by multiple engineering teams.
- Design systems to behave safely when encountering malformed, malicious, unexpected, or adversarial inputs.
- Define clear boundaries between application logic, AI orchestration, model interaction, and safety controls.
Backend & Platform Engineering
- Design, develop, test, and maintain production-quality backend services and APIs using Python and TypeScript.
- Architect scalable services that operate across distributed application environments.
- Develop shared libraries, services, APIs, and platform components with broad organizational use.
- Define stable interfaces and service contracts consumed by other engineering teams.
- Apply defensive programming principles to distributed and AI-enabled applications.
- Design for service resiliency, fault isolation, graceful degradation, and predictable failure behavior.
Agentic AI Engineering
- Design and build agentic, multi-step, or multi-agent workflows.
- Integrate LLMs with backend systems, APIs, tools, data sources, and business services.
- Design appropriate controls around agent actions and tool invocation.
- Address AI-specific risks including:
- Prompt injection
- Untrusted model output
- Hallucinated or malformed responses
- Unauthorized tool use
- Unexpected agent behavior
- Cross-service failure propagation
- Build systems that treat LLM outputs as potentially untrusted inputs requiring appropriate validation.
AWS & Infrastructure Engineering
- Design and deploy production systems in AWS.
- Work with services such as:
- AWS Lambda
- ECS/Fargate
- API Gateway
- IAM
- CloudWatch
- Related serverless and container-based AWS services
- Build and maintain production infrastructure using Terraform.
- Apply Infrastructure-as-Code practices that support repeatable deployments, secure configurations, and scalable environments.
- Participate in architectural decisions involving compute, networking, IAM, service boundaries, and deployment patterns.
Observability & Reliability
- Instrument distributed systems using:
- Structured logging
- Metrics
- Distributed tracing
- Error monitoring
- Service health indicators
- Improve visibility into behavior across service boundaries.
- Design systems that enable engineers to identify where and why enforcement or workflow failures occur.
- Consider downstream dependencies, timeouts, retries, partial failures, and failure propagation when designing services.
Technical Leadership
- Mentor junior and mid-level engineers on:
- Defensive programming
- Safe AI integration
- API and service design
- Error and exception handling
- Failure management
- Secure coding practices
- Define technical patterns and engineering standards that other developers can consistently follow.
- Communicate architectural decisions clearly through documentation, diagrams, code reviews, and technical discussions.
- Develop experience-backed technical opinions and constructively challenge designs when appropriate.
- Remain a hands-on engineer capable of implementing the systems and patterns being recommended.
Required Qualifications
- 5–8 years of professional software engineering experience.
- Strong production software development experience using Python.
- Strong production software development experience using TypeScript.
- Demonstrated experience designing and building backend services and APIs.
- Strong experience delivering production systems in AWS.
- Hands-on experience with AWS services such as Lambda, Fargate/ECS, and API Gateway.
- Strong production experience with Terraform and Infrastructure as Code.
- Experience designing systems that operate across multiple services or distributed system boundaries.
- Experience defining interfaces, contracts, reusable components, or engineering patterns used by other development teams.
- Experience with LLM integrations or GenAI-enabled applications.
- Familiarity with AI safety and LLM integration concepts such as:
- Prompt-injection detection
- Guardrail design
- Input validation
- Output filtering
- Policy enforcement
- Experience designing or building agentic workflows, multi-step AI workflows, or multi-agent systems.
- Understanding of distributed-system concerns including dependencies, failure propagation, retries, resiliency, and contract stability.
- Ability to mentor engineers and communicate defensive software engineering practices.
- Strong written and verbal technical communication skills.
Preferred Qualifications
- Hands-on experience with AWS Bedrock.
- Experience invoking and integrating foundation models through Bedrock.
- Experience with AWS Bedrock Guardrails.
- Exposure to Amazon Bedrock AgentCore or comparable agent runtime technologies.
- Experience building developer platforms, shared engineering services, or internal developer tooling.
- Experience developing centralized policy or enforcement services.
- Experience implementing distributed tracing and advanced production observability.
- Experience with secure software development or application security principles.
- Experience designing authorization or policy controls around AI agent tool usage.
What Makes Someone Successful in This Role
The strongest candidate will be an experienced software engineer who can move comfortably between writing production code and making system-level architectural decisions.
They will understand that safely integrating LLMs requires considerably more than connecting an application to a model API. They will think carefully about:
- What inputs can be trusted
- Where validation should occur
- Which service owns an enforcement decision
- What an AI agent should and should not be permitted to do
- How model output should be validated
- How failures propagate between services
- How enforcement decisions are logged and observed
- How patterns can be standardized so other development teams implement them consistently
Successful engineers in this role will be able to explain not only how they built a system, but also why specific controls were placed at particular architectural boundaries.
Ideal Candidate Profile
An especially strong candidate may have progressed through a career path similar to:
Backend Software Engineer → Senior Cloud/Platform Engineer → AWS & Terraform Engineer → GenAI/Agentic Systems Engineer → AI Safety / Guardrail Engineering
Strong candidates may come from backend engineering, cloud engineering, platform engineering, developer tooling, or application security engineering backgrounds if they have meaningful hands-on experience integrating LLMs or agentic technologies.
Candidates Who May Not Be a Strong Fit
This position is unlikely to be a fit for candidates whose experience is primarily:
- Data Science or statistical modeling
- Machine Learning research
- Model training or fine-tuning without production software engineering
- Prompt engineering without backend development
- RAG prototypes without production platform ownership
- Frontend development with limited backend architecture experience
- Python development with little or no TypeScript
- AWS usage without meaningful production architecture experience
- Terraform exposure without hands-on Infrastructure-as-Code development
- GenAI demonstrations or proofs of concept without production deployment
- Architecture or management roles where the candidate is no longer hands-on with code
Key Technologies
Primary
- Python
- TypeScript
- AWS
- Terraform
- REST APIs / Backend Services
- Distributed Systems
- LLM Integration
- Agentic AI
Relevant AWS Technologies
- AWS Lambda
- ECS / Fargate
- API Gateway
- IAM
- CloudWatch
- AWS Bedrock
AI Safety / Engineering Concepts
- Prompt Injection Detection
- LLM Guardrails
- Input Validation
- Output Filtering
- Policy Enforcement
- Agent Tool Controls
- Defensive Programming
- Failure Handling
- Multi-Agent / Agentic Workflows
Top Skills — Must Have
- Python — advanced production backend engineering
- TypeScript — strong production software engineering capability
- AWS — demonstrated delivery of production cloud systems
- Terraform — hands-on production Infrastructure-as-Code experience
- Backend/API Architecture — scalable services spanning multiple service boundaries
- LLM Integration & Safety — familiarity with prompt-injection defenses, guardrails, validation, or output filtering
- Agentic AI — experience developing agentic, multi-step, or multi-agent workflows
- Distributed Systems Thinking — understanding service contracts, dependencies, resiliency, and failure propagation