Senior MLOps EngineerRole SummaryWe are seeking a Senior MLOps Engineer to lead the development and management of infrastructure... Read more
Senior MLOps Engineer
Role Summary
We are seeking a Senior MLOps Engineer to lead the development and management of infrastructure designed for training, deploying, and maintaining ML models. This role plays a critical function in operationalizing state-of-the-art systems to ensure high-performance delivery across research and production environments.
The successful candidate will be responsible for designing and implementing infrastructure to support efficient model deployment, inference, monitoring, and retraining. This includes close collaboration with cross-functional teams to integrate machine learning models into scalable and secure production pipelines, enabling the delivery of real-time, data-driven solutions across various domains.
Key Responsibilities
Inference Serving Across a Wide Model Range: Deploy and scale self-hosted open-weight models from ~7B up to ~376B parameters using engines such as vLLM, Triton, or TGI, choosing serving strategies (continuous batching, tensor/pipeline parallelism, quantization) appropriate to each model's size and SLA.Multi-Provider Gateway & Cost Governance: Operate and improve the routing layer that spans self-hosted models and external APIs (OpenAI, Anthropic, OpenRouter). Build token accounting, budget controls, and cost/latency/quality-aware routing for model consumption predictably and within budget.Training & Fine-Tuning Infrastructure: Build and maintain automated pipelines for fine-tuning, evaluation, versioning, and continuous delivery (MLflow, SageMaker Pipelines, or Kubeflow), including distributed training with DeepSpeed, FSDP, or Accelerate.Reliability & Observability: Own production reliability for ML services, including monitoring, logging, alerting, incident response, and safe rollback, to meet latency, throughput, and availability targets. Participate in on-call for the serving platform.Evaluation & Regression Safety: Stand up evaluation and verification harnesses that catch quality and performance regressions before they reach users, ensuring model or infrastructure changes ship with evidence rather than hope.Platform as a Product: Treat internal engineers and external end-users as customers by reducing friction through sensible defaults, self-service tooling, and clear contracts. Deliver Infrastructure-as-Code (IaC) and CI/CD for repeatable, secure deployments.Model & Cost Efficiency: Apply quantization, pruning, and multi-GPU/distributed inference to reduce latency and cost, especially at the high end of the model range.Qualifications
Professional Experience: 5+ years in MLOps, ML Infrastructure, or ML Engineering, owning end-to-end model lifecycles in production.Production Ownership: Experience supporting ML services in production and handling real incidents, including degraded inference, GPU OOMs, and cost escalations.Inference Serving Depth: Hands-on experience serving models across a range of sizes, with real decisions made around quantization and parallelism trade-offs under latency and cost constraints.Cost & Capacity Discipline: Demonstrable track record of measuring and optimizing GPU and cloud spend for ML workloads.Programming Skills: Strong Python skills; additional C/C++ experience for performance-sensitive workloads is advantageous.Cloud & Orchestration: Strong experience with cloud services (e.g., AWS SageMaker, EC2, EKS, Lambda), Docker, and Kubernetes. Experience across major hyperscalers is beneficial.Tooling & Distributed Training: Proficiency with MLOps frameworks (MLflow, Kubeflow, or SageMaker Pipelines) and distributed training frameworks (DeepSpeed, FSDP, Accelerate).Bonus: Experience building evaluation and verification harnesses, or multi-provider LLM gateways with token and cost management.Educational Background: Bachelor's or Master's degree in Computer Science, Machine Learning, Data Engineering, or a related field, or equivalent practical experience.
GCS is acting as an Employment Business in relation to this vacancy.
Read lessUnited Arab EmiratesContract
MLOps EngineerRole OverviewWe are seeking an experienced MLOps Engineer to design, build, and operate the platforms, infrastructure, and... Read more
We are seeking an experienced MLOps Engineer to design, build, and operate the platforms, infrastructure, and processes required to deploy and manage machine learning and AI solutions at scale.
The MLOps Engineer will bridge the gap between data science, AI engineering, software engineering, data engineering, and cloud/platform teams, ensuring that models can be developed, tested, deployed, monitored, and maintained reliably in production.
The ideal candidate will have strong experience across cloud infrastructure, automation, CI/CD, machine learning lifecycle management, model deployment, monitoring, and DevOps practices.
Key ResponsibilitiesDesign and implement scalable MLOps platforms and architectures for machine learning and AI workloads.Build automated pipelines covering the full ML lifecycle, from data and model development through to production deployment and monitoring.Develop and maintain CI/CD and continuous training (CT) pipelines for machine learning models.Automate model testing, validation, deployment, rollback, and lifecycle management.Implement model versioning, experiment tracking, model registries, and reproducible ML workflows.Build and maintain infrastructure for model training, inference, and serving.Deploy machine learning models across cloud, containerised, and Kubernetes-based environments.Implement automated monitoring for model performance, data quality, drift, availability, and operational health.Establish processes for model retraining and continuous improvement.Work closely with data scientists and AI engineers to productionise models and AI applications.Collaborate with data engineers to integrate ML pipelines with enterprise data platforms.Implement infrastructure and environments using Infrastructure as Code (IaC).Develop reusable tooling, frameworks, templates, and deployment patterns for ML teams.Optimise compute, storage, model serving, and cloud infrastructure costs.Implement appropriate security, access control, secrets management, and compliance controls.Support deployment of both traditional machine learning models and Generative AI/LLM applications.Establish observability and operational support processes for production AI/ML systems.Troubleshoot complex infrastructure, deployment, pipeline, model-serving, and performance issues.Define and document MLOps standards, architecture patterns, engineering practices, and operational procedures.Mentor data scientists, AI engineers, and software engineers on production ML practices.Required Skills and ExperienceStrong experience in MLOps, DevOps, machine learning engineering, cloud engineering, or related disciplines.Strong understanding of the end-to-end machine learning lifecycle.Strong programming and scripting skills, particularly Python.Experience building and managing CI/CD pipelines.Experience with containerisation technologies such as Docker.Experience with Kubernetes and container orchestration.Strong experience with at least one major cloud platform such as Azure, AWS, or Google Cloud Platform.Experience with Infrastructure as Code tools such as Terraform.Experience deploying and managing machine learning models in production.Experience with model versioning, experiment tracking, and model registries.Experience with ML platforms and tools such as MLflow, Kubeflow, Azure Machine Learning, AWS SageMaker, or equivalent.Strong understanding of Git, automated testing, deployment automation, and DevOps practices.Experience implementing monitoring, logging, observability, and alerting.Understanding of data pipelines, data quality, model performance, and data/model drift.Strong understanding of cloud security and identity/access management.Excellent troubleshooting and problem-solving skills.Desirable SkillsExperience supporting Generative AI and LLM workloads.Experience deploying RAG applications and vector search infrastructure.Experience with LLM evaluation, monitoring, and observability.Experience with platforms such as Databricks, Snowflake, Azure OpenAI, Amazon Bedrock, or Google Vertex AI.Experience with Apache Spark and distributed data processing.Experience with Kafka or other event-streaming technologies.Experience with GitHub Actions, GitLab CI/CD, Azure DevOps, Jenkins, or equivalent.Experience with Kubernetes tools such as Helm.Experience with cloud-native monitoring technologies.Experience implementing automated model retraining pipelines.Knowledge of responsible AI, AI governance, security, and regulatory requirements.Experience with FinOps and optimisation of cloud-based ML workloads.MLOps Platform ResponsibilitiesThe MLOps Engineer will typically be responsible for establishing and maintaining capabilities across:
Source Control - Git-based development and version management.CI/CD - Automated build, test, validation, and deployment pipelines.Experiment Tracking - Tracking experiments, parameters, metrics, and artefacts.Model Registry - Model versioning, approval, promotion, and lifecycle management.Model Serving - Reliable and scalable online and batch inference.Infrastructure - Automated provisioning of ML environments and compute.Monitoring - Model, application, infrastructure, and data monitoring.Data & Model Drift - Detection and remediation of changes affecting model performance.Security - Identity, access control, secrets, network security, and compliance.Governance - Auditability, lineage, approvals, and model lifecycle controls.Automation - Reducing manual intervention across the ML lifecycle.Key CompetenciesMLOps & ML Lifecycle ManagementCloud EngineeringDevOps & CI/CDPython & AutomationDocker & KubernetesInfrastructure as CodeModel Deployment & ServingMLflow / ML PlatformsMonitoring & ObservabilityModel & Data DriftCloud SecurityGenerative AI & LLM OperationsPerformance & Cost OptimisationTechnical Problem SolvingTypical Experience Level4-9+ years of experience across MLOps, DevOps, cloud engineering, machine learning engineering, or related disciplines, with demonstrable experience deploying and operating machine learning or AI solutions in production.
A strong candidate should be capable of taking an ML/AI solution from development to production, establishing the automation, infrastructure, monitoring, governance, and operational processes required to run it reliably at scale.
GCS is acting as an Employment Business in relation to this vacancy.
Read lessBrussels, Brussels Hoofdstedelijk Gewest, BelgiumContract
QA Automation Engineer - AI & Agentic SystemsLevel: Senior IC Location: RemoteRole SummaryWe are building agentic AI systems... Read more
QA Automation Engineer - AI & Agentic Systems
Level: Senior IC
Location: Remote
Role Summary
We are building agentic AI systems that can interpret complex data sources, documentation, and structured information to perform analysis, validation, and decision-support tasks. This role exists to make those agents trustworthy enough to act on.
This is a senior, hands-on quality role weighted toward testing non-deterministic AI and agentic systems. That is where most of your time sits and where the hiring bar is highest. You will also own the broader quality surface: integration, API, and performance testing are part of the remit, not out of scope.
The differentiator for this role is the ability to define what "good" looks like when a system reasons, calls tools, and can be wrong in subtle ways. You will scope, build, and run the frameworks yourself, with the independence of a senior engineer, and decide what to build first.
Key Responsibilities
Agentic System Testing (Primary Focus)Tool-Use & Trajectory Evaluation: Test whether agents select the right tools for the right reasons and follow sound multi-step trajectories, not just whether the final answer looks plausible. Evaluate planning, intermediate steps, and recovery when a tool fails or returns nothing.Grounding & Citation Verification: Verify that agent claims are backed by cited evidence in source data, documents, or knowledge bases, and that referenced locations actually support the answer, catching confident but unsupported outputs.Honest-Failure & Refusal Calibration: Assert that agents ask for clarification or respond appropriately when information cannot be confidently determined, rather than inventing answers.Guardrail & Adversarial Testing: Probe prompt injection, jailbreaks, and instructions hidden within ingested content, ensuring the agent treats source content as untrusted data.Multi-Turn & State: Validate follow-ups, references to prior turns, and that conversational state carries correctly across a session.Technical Requirements
AI Evaluation (Core): Hands-on experience with LLM/agent evaluation frameworks (e.g., DeepEval, TruLens, RAGAS, or custom Python evaluators) and LLM-as-judge techniques.Agent Observability: Experience tracing and debugging agent runs, including tool calls, intermediate steps, token usage, and latency, using tools such as LangSmith, Langfuse, or OpenTelemetry-based tracing.Core Automation: Expert Python for custom test harnesses and evaluation tooling (Pytest), plus standard automation libraries (Selenium/Playwright for UI, Requests for API).Performance Testing: Proven ability to design and implement performance test plans (e.g., Locust, JMeter, k6).Data Validation: Proficiency with SQL and data-validation tools, and familiarity with vector databases and retrieval corpora.CI/CD Integration: Integrating automated tests and evaluations into GitLab CI/CD pipelines and enforcing quality gates.Test Management & Reporting: Managing test reports and artifacts (e.g., TestRail, Allure) and communicating results clearly.Version Control & QE Practices: Maintaining code-based frameworks in Git/GitLab and applying modern quality-engineering practices.Traceability Tools: Familiarity with requirements-management tools (e.g., Jira, Linear, Jama, Polarion) and linking results to requirement IDs.Professional Qualifications
Experience: 5+ years in QA automation or quality engineering, with at least 2 years focused on testing ML models, LLM applications, or AI agents.Probabilistic Systems Judgement: Able to define pass/fail criteria for systems whose outputs are not identical every run and communicate confidence levels clearly to engineering leadership.Independent Operator: A senior individual contributor who scopes and builds testing and evaluation frameworks with minimal direction and prioritises what matters most.GCS is acting as an Employment Business in relation to this vacancy.
Read lessUnited KingdomContractRemote
Job Description- Implements, refines, and validates machine learning algorithms for products and applications. - Implements data pipelines consisting... Read more
Job Description
- Implements, refines, and validates machine learning algorithms for products and applications. - Implements data pipelines consisting of data ingest, data validation, data cleaning, and data monitoring.
- Trains machine learning models, validates the accuracy of the machine learning models once trained, and deploys validated machine learning models into production.
- Assists in development of proof of concept solutions and contributes to studies to support future product or application development.
- Researches, writes, and edits documentation and technical requirements, including evaluation plans, confluence pages, white papers, presentations, test results, technical manuals, formal recommendations, and reports.
- Tests and evaluates solutions. Completes case studies, testing, and reporting.
Skills
- Bachelor's degree in computer science, computer engineering, mathematics, related technical discipline, or related industry experience
- Experience with machine learning, deep learning, data mining, and/or statistical analysis tools and how to deploy and monitor machine learning models.
- Strong programming and software development skills and familiarity with Python, Java or Scala.
- Knowledge of data pipeline and cloud technologies such as Kafka, Spark, and Docker.
- 1-3 years related experience after Bachelors.
GCS is acting as an Employment Agency in relation to this vacancy.
Read lessUnited States of AmericaContractRemote
All your saved jobs are no longer available or you've already applied.
for the following search criteria