Apply on Xenon7’s siteOur Client's Digital Finance IT is scaling AI and agentic systems in production. We need an MLOps / Cloud Deployment Engineer to own the deployment, reliability, observability, and operational scale of these systems in a regulated enterprise environment. This is a cloud and platform engineering role with deep MLOps/LLMOps focus, not a model-building role. You will operate the runway that ML and GenAI systems run on, not build the models themselves. What You'll Do Own CI/CD pipelines for ML models, RAG applications, and agentic AI systems — from experiment to production Deploy and operate AI workloads on cloud-native ML/AI platforms — AWS Bedrock/SageMaker, Azure AI Foundry / Azure Machine Learning, or equivalent Build and maintain observability, tracing, and monitoring for LLM and agentic systems — latency, cost, hallucination rates, tool-call success, drift detection Implement model governance and guardrails — approval gates, kill-switches, escalation paths, audit trails Manage infrastructure-as-code (Terraform, Bicep, or equivalent) for reproducible AI/ML environments Design cost and performance optimization strategies — token usage tracking, caching, model routing, autoscaling, warehouse/cluster right-sizing Own security posture — RBAC, secret management (Key Vault / Secrets Manager), prompt-injection risk mitigation, auditability for regulated pharma Partner with data engineers, AI engineers, and Finance business stakeholders to move systems from prototype to reliable production Implement evaluation frameworks for AI systems in production — regression testing, adversarial testing, accuracy tracking, hallucination monitoring Requirements Must-Have Experience 5+ years in cloud/DevOps/MLOps engineering on AWS, Azure, or GCP Production deployment of ML or GenAI systems — CI/CD, containerization (Docker/Kubernetes), infrastructure-as-code (Terraform) MLOps tooling — MLflow, SageMaker Pipelines, Azure ML Pipelines, or equivalent LLM/GenAI operational experience — observability tools (LangSmith, Weights & Biases, or equivalent), cost monitoring, latency optimization, prompt/model versioning Cloud-native AI platforms — hands-on with at least one of: AWS Bedrock, SageMaker, Azure AI Foundry, Azure OpenAI, Vertex AI Python, Bash, and infrastructure scripting — strong Security and governance in regulated environments — RBAC, secrets, audit, compliance Nice to Have Pharma, life sciences, or regulated financial services domain Experience operating agentic AI systems in production — multi-agent orchestration, tool-calling, human-in-the-loop workflows LangChain, LangGraph, CrewAI, AutoGen, or Semantic Kernel operational experience Kubernetes-native ML platforms (Kubeflow, Ray) Snowflake or Databricks operational experience (compute governance, cost management) Certifications: AWS/Azure ML Engineer, Kubernetes CKA/CKAD, Terraform Associate What We're NOT Looking For Data Scientists or research engineers — this is a production platform role Application developers with light DevOps exposure — need real MLOps/cloud engineering depth Pure infra engineers with no AI/ML operational experience — need to understand what makes LLM systems different (evals, hallucinations, prompt versioning, RAG grounding)
This role is published by Xenon7 on workable. SwiftFit is not the employer and does not accept applications.