Principal Machine Learning Engineer
Purpose of the role
The successful candidate will set the technical direction for how agents and conventional ML models are built, deployed, observed and operated across the estate, and will be the escalation point when something in production needs deep infrastructure expertise rather than model expertise.
Key responsibilities
· Own the end-to-end architecture of the agent estate/foundry: build, deployment, monitoring, and reliability of agentic and conventional ML systems in production.
· Set standards and reusable patterns for CI/CD, environment management, and release processes for both traditional ML models and agentic services.
· Design and evolve the platform's use of MCP servers and A2A-based agent-to-agent communication so agents can be composed, discovered and orchestrated reliably.
· Establish production observability and evaluation practices for LLM-based and conventional ML systems (MLOps/LLMOps), including monitoring, alerting, drift detection and rollback strategies.
· Lead integration of new, greenfield agentic components with legacy services and existing data sources, balancing modern methods against operational constraints.
· Act as the senior technical authority on Azure (preferred) or AWS AI infrastructure decisions, including evaluation of Azure AI Foundry where relevant.
· Mentor the Senior Machine Learning Engineer and set technical direction across the wider engineering function, while pairing closely with the data science leadership on shared priorities.
· Represent engineering feasibility, cost and risk in architectural decisions taken jointly with the Principal Data Scientist / AI Engineer.
Essential experience
· 8+ years in machine learning or platform engineering with direct, hands-on ownership of production systems (not purely advisory or research-only experience).
· Proven experience building AI infrastructure on Azure (strongly preferred) or AWS at meaningful scale.
· Experience deploying and running conventional ML models alongside agentic/LLM-based systems side by side in the same estate.
· Track record of greenfield delivery using current methods, integrated with legacy services and data sources — not just building on a clean slate.
· Direct, practical experience with the realities of CI/CD in an ML/AI context (not just familiarity with the concept).
Essential skills
· Python at a senior/architect level.
· Docker and/or Podman for containerisation.
· Terraform and CI/CD pipeline design and ownership.
· SQL and database management.
· MCP server development and deployment.
· A2A (Agent2Agent protocol) — hard requirement, not negotiable.
· MLOps and LLMOps, including production observability and evaluation of ML/LLM systems.
· Experience with at least one agent framework or SDK (Claude Agent SDK experience is a bonus, but any comparable framework is acceptable).
Desirable / nice-to-have
· Azure AI Foundry experience.
· RAG systems and vector stores (pgvector, Pinecone).
· ETL and DAG orchestration tooling.
· Serverless architectures.
What a strong candidate looks like
A hands-on, still-coding principal or staff-level engineer who has genuinely carried production ML or platform systems through incidents, not just designed them on a whiteboard. Comfortable being the last line of defence when something breaks at 2am, and equally comfortable setting multi-quarter technical direction.
Has clearly worked across both “boring” conventional ML (batch pipelines, model serving, monitoring) and newer agentic patterns — candidates who have only ever done one or the other are unlikely to succeed here.
Sourcing notes for the agency
· Target titles: Principal/Staff Machine Learning Engineer, Head of MLOps, ML Platform Lead, AI Infrastructure Lead.
· Look for candidates coming from regulated or enterprise environments where legacy integration was unavoidable, rather than pure greenfield startups.
· A2A and MCP are both recent standards — do not filter too hard on exact protocol name-matches; prioritise candidates with strong agent-to-agent / tool-use orchestration experience who can pick up the specific protocol quickly, but confirm they understand the concept in screening.
· Azure AI Foundry is a newer product; few candidates will have deep hands-on time with it specifically — treat general Azure AI infrastructure depth as the more important signal.
Screening flags / watch-outs
· A GCP-only cloud background is explicitly called out as not workable given ramp-up time — screen for this early.
· Candidates who are strong on agent/LLM work but have never owned production infrastructure, CI/CD or on-call responsibility will not meet the “production ownership” bar for this level.
· Be wary of candidates whose “8+ years’ ML experience” is mostly modelling/data science — this is an engineering-owner role, not a data science role.