Productionise the existing RAG and scanner toolchain through CI/CD — connecting the pipeline end to end so scans, dispositions and remediations flow without manual intervention
Build the guardian / evaluation agent: an automated check that runs on every sub-agent deliverable, replacing the current brute-force knowledge-capture approach with a best-practice evaluation pattern
Implement the deterministic assertion layer as a programmatic gate — automatically rejecting any disposition that contradicts its own evidence, before a human ever sees it
Own AgentOps: trace capture, prompt / rule / model versioning, evaluation-in-CI, regression harnesses, and drift detection
Build the observability the team watches daily: pending burn-down, auto-disposition rate, accuracy against the gold set, human-minutes per item, assertion-rejection rate and cost per item
Own FinOps for the AI workload: model routing, delta-scoped runs, caching, and a per-cycle token budget tracked as a service-level objective
Guarantee provenance and auditability for a regulated environment — every decision reproducible from its evidence snapshot, rule/prompt/model version and human verdict
Qualifications
7+ years in platform / DevOps / MLOps engineering, including production LLM or ML workloads
Python — production-grade
CI/CD automation for application and ML/LLM workloads; release automation and test gating