ML OPs engineer
MLOps Engineer
Role Overview
We are seeking an experienced MLOps Engineer to design, build, and operate the platforms, infrastructure, and processes required to deploy and manage machine learning and AI solutions at scale.
The MLOps Engineer will bridge the gap between data science, AI engineering, software engineering, data engineering, and cloud/platform teams, ensuring that models can be developed, tested, deployed, monitored, and maintained reliably in production.
The ideal candidate will have strong experience across cloud infrastructure, automation, CI/CD, machine learning lifecycle management, model deployment, monitoring, and DevOps practices.
Key Responsibilities
- Design and implement scalable MLOps platforms and architectures for machine learning and AI workloads.
- Build automated pipelines covering the full ML lifecycle, from data and model development through to production deployment and monitoring.
- Develop and maintain CI/CD and continuous training (CT) pipelines for machine learning models.
- Automate model testing, validation, deployment, rollback, and lifecycle management.
- Implement model versioning, experiment tracking, model registries, and reproducible ML workflows.
- Build and maintain infrastructure for model training, inference, and serving.
- Deploy machine learning models across cloud, containerised, and Kubernetes-based environments.
- Implement automated monitoring for model performance, data quality, drift, availability, and operational health.
- Establish processes for model retraining and continuous improvement.
- Work closely with data scientists and AI engineers to productionise models and AI applications.
- Collaborate with data engineers to integrate ML pipelines with enterprise data platforms.
- Implement infrastructure and environments using Infrastructure as Code (IaC).
- Develop reusable tooling, frameworks, templates, and deployment patterns for ML teams.
- Optimise compute, storage, model serving, and cloud infrastructure costs.
- Implement appropriate security, access control, secrets management, and compliance controls.
- Support deployment of both traditional machine learning models and Generative AI/LLM applications.
- Establish observability and operational support processes for production AI/ML systems.
- Troubleshoot complex infrastructure, deployment, pipeline, model-serving, and performance issues.
- Define and document MLOps standards, architecture patterns, engineering practices, and operational procedures.
- Mentor data scientists, AI engineers, and software engineers on production ML practices.
Required Skills and Experience
- Strong experience in MLOps, DevOps, machine learning engineering, cloud engineering, or related disciplines.
- Strong understanding of the end-to-end machine learning lifecycle.
- Strong programming and scripting skills, particularly Python.
- Experience building and managing CI/CD pipelines.
- Experience with containerisation technologies such as Docker.
- Experience with Kubernetes and container orchestration.
- Strong experience with at least one major cloud platform such as Azure, AWS, or Google Cloud Platform.
- Experience with Infrastructure as Code tools such as Terraform.
- Experience deploying and managing machine learning models in production.
- Experience with model versioning, experiment tracking, and model registries.
- Experience with ML platforms and tools such as MLflow, Kubeflow, Azure Machine Learning, AWS SageMaker, or equivalent.
- Strong understanding of Git, automated testing, deployment automation, and DevOps practices.
- Experience implementing monitoring, logging, observability, and alerting.
- Understanding of data pipelines, data quality, model performance, and data/model drift.
- Strong understanding of cloud security and identity/access management.
- Excellent troubleshooting and problem-solving skills.
Desirable Skills
- Experience supporting Generative AI and LLM workloads.
- Experience deploying RAG applications and vector search infrastructure.
- Experience with LLM evaluation, monitoring, and observability.
- Experience with platforms such as Databricks, Snowflake, Azure OpenAI, Amazon Bedrock, or Google Vertex AI.
- Experience with Apache Spark and distributed data processing.
- Experience with Kafka or other event-streaming technologies.
- Experience with GitHub Actions, GitLab CI/CD, Azure DevOps, Jenkins, or equivalent.
- Experience with Kubernetes tools such as Helm.
- Experience with cloud-native monitoring technologies.
- Experience implementing automated model retraining pipelines.
- Knowledge of responsible AI, AI governance, security, and regulatory requirements.
- Experience with FinOps and optimisation of cloud-based ML workloads.
MLOps Platform Responsibilities
The MLOps Engineer will typically be responsible for establishing and maintaining capabilities across:
- Source Control - Git-based development and version management.
- CI/CD - Automated build, test, validation, and deployment pipelines.
- Experiment Tracking - Tracking experiments, parameters, metrics, and artefacts.
- Model Registry - Model versioning, approval, promotion, and lifecycle management.
- Model Serving - Reliable and scalable online and batch inference.
- Infrastructure - Automated provisioning of ML environments and compute.
- Monitoring - Model, application, infrastructure, and data monitoring.
- Data & Model Drift - Detection and remediation of changes affecting model performance.
- Security - Identity, access control, secrets, network security, and compliance.
- Governance - Auditability, lineage, approvals, and model lifecycle controls.
- Automation - Reducing manual intervention across the ML lifecycle.
Key Competencies
- MLOps & ML Lifecycle Management
- Cloud Engineering
- DevOps & CI/CD
- Python & Automation
- Docker & Kubernetes
- Infrastructure as Code
- Model Deployment & Serving
- MLflow / ML Platforms
- Monitoring & Observability
- Model & Data Drift
- Cloud Security
- Generative AI & LLM Operations
- Performance & Cost Optimisation
- Technical Problem Solving
Typical Experience Level
4-9+ years of experience across MLOps, DevOps, cloud engineering, machine learning engineering, or related disciplines, with demonstrable experience deploying and operating machine learning or AI solutions in production.
A strong candidate should be capable of taking an ML/AI solution from development to production, establishing the automation, infrastructure, monitoring, governance, and operational processes required to run it reliably at scale.
GCS is acting as an Employment Business in relation to this vacancy.
ML OPs engineer
Other similar jobs
Popular job searches
Your next job
starts here.
JOB SPECIALISMS
LATEST JOBS
TOP SEARCHES
LOCATIONS
- IT Support & Infrastructure
- Project Management
- Engineering
- Data
- Network security consultant
- Software Development
- Controls & Automation
- DevOps
- BI & Data Analytics
- Manufacturing & Production
- Embedded Software
- Engineering Technology
LATEST JOBS
- ML OPs engineer
- AI engineer
- Data lead
- Senior data architect
- Technical Project Manager - Di...
- Technical Project Manager - UX...
- Senior FPGA Design Engineer
- Front-End Engineer (Contract)
- DevOps Engineer
- Senior FPGA Engineer
- FPGA Engineer
- Fiber Engineer
TOP SEARCHES
LOCATIONS
- Engineer
- Data Scientist
- Senior Data Scientist
- Head of Data Science
- Trainee Data Scientist
- Data Science Graduate
- Senior Financial Accountant
- Management Accountant
- Cost Accountant
- Civil Engineer
- Senior Civil Engineer
- Civil Design Engineer