MLOps Engineer
Senior MLOps Engineer
Role Summary
We are seeking a Senior MLOps Engineer to lead the development and management of infrastructure designed for training, deploying, and maintaining ML models. This role plays a critical function in operationalizing state-of-the-art systems to ensure high-performance delivery across research and production environments.
The successful candidate will be responsible for designing and implementing infrastructure to support efficient model deployment, inference, monitoring, and retraining. This includes close collaboration with cross-functional teams to integrate machine learning models into scalable and secure production pipelines, enabling the delivery of real-time, data-driven solutions across various domains.
Key Responsibilities
- Inference Serving Across a Wide Model Range: Deploy and scale self-hosted open-weight models from ~7B up to ~376B parameters using engines such as vLLM, Triton, or TGI, choosing serving strategies (continuous batching, tensor/pipeline parallelism, quantization) appropriate to each model's size and SLA.
- Multi-Provider Gateway & Cost Governance: Operate and improve the routing layer that spans self-hosted models and external APIs (OpenAI, Anthropic, OpenRouter). Build token accounting, budget controls, and cost/latency/quality-aware routing for model consumption predictably and within budget.
- Training & Fine-Tuning Infrastructure: Build and maintain automated pipelines for fine-tuning, evaluation, versioning, and continuous delivery (MLflow, SageMaker Pipelines, or Kubeflow), including distributed training with DeepSpeed, FSDP, or Accelerate.
- Reliability & Observability: Own production reliability for ML services, including monitoring, logging, alerting, incident response, and safe rollback, to meet latency, throughput, and availability targets. Participate in on-call for the serving platform.
- Evaluation & Regression Safety: Stand up evaluation and verification harnesses that catch quality and performance regressions before they reach users, ensuring model or infrastructure changes ship with evidence rather than hope.
- Platform as a Product: Treat internal engineers and external end-users as customers by reducing friction through sensible defaults, self-service tooling, and clear contracts. Deliver Infrastructure-as-Code (IaC) and CI/CD for repeatable, secure deployments.
- Model & Cost Efficiency: Apply quantization, pruning, and multi-GPU/distributed inference to reduce latency and cost, especially at the high end of the model range.
Qualifications
- Professional Experience: 5+ years in MLOps, ML Infrastructure, or ML Engineering, owning end-to-end model lifecycles in production.
- Production Ownership: Experience supporting ML services in production and handling real incidents, including degraded inference, GPU OOMs, and cost escalations.
- Inference Serving Depth: Hands-on experience serving models across a range of sizes, with real decisions made around quantization and parallelism trade-offs under latency and cost constraints.
- Cost & Capacity Discipline: Demonstrable track record of measuring and optimizing GPU and cloud spend for ML workloads.
- Programming Skills: Strong Python skills; additional C/C++ experience for performance-sensitive workloads is advantageous.
- Cloud & Orchestration: Strong experience with cloud services (e.g., AWS SageMaker, EC2, EKS, Lambda), Docker, and Kubernetes. Experience across major hyperscalers is beneficial.
- Tooling & Distributed Training: Proficiency with MLOps frameworks (MLflow, Kubeflow, or SageMaker Pipelines) and distributed training frameworks (DeepSpeed, FSDP, Accelerate).
- Bonus: Experience building evaluation and verification harnesses, or multi-provider LLM gateways with token and cost management.
- Educational Background: Bachelor's or Master's degree in Computer Science, Machine Learning, Data Engineering, or a related field, or equivalent practical experience.
GCS is acting as an Employment Business in relation to this vacancy.