5+ years hands-on data engineering in production environments, covering data modelling, storage design, query performance and the operational behaviour of the stores you choose.
3+ years working on agentic or LLM data patterns, with real depth in agent memory: session and conversation state, short-term and long-term memory, summarisation and compaction, expiry and retention, and isolation between users and threads. This is the defining requirement: conventional data engineering alone is not sufficient for this position.
3+ years with vector and semantic search, using Azure AI Search, PostgreSQL with pgvector, Elasticsearch, or a comparable vector store, including hybrid search, relevance tuning and index design.
3+ years designing NoSQL, document or key-value stores for high-write, low-latency workloads: partitioning and sharding strategy, item and document size constraints, time-to-live and retention, and the read and write patterns of conversational or session-based data.
2+ years working with caching or in-memory stores such as Redis or equivalent, for volatile and ephemeral state, including expiry strategy and the trade-offs against durable storage.
Working knowledge of retrieval-augmented generation, including chunking and embedding strategy and how model choice and chunking affect retrieval quality and cost. You will consume an existing vectorised knowledge base more often than you build one.
3+ years Python to production standard, building interfaces or libraries consumed by other engineers rather than scripts.
2+ years working on a major cloud platform, including managed data services, identity and access to data stores, and private networking to data services.
Experience evaluating retrieval and memory quality, using groundedness, relevance or comparable measures, rather than relying on subjective assessment