Senior Site Reliability Engineer

Reference: KK-SRE-IRE_1785928896

About the Role

Client is seeking a highly skilled Senior / Lead Site Reliability Engineer (SRE) to drive reliability, observability, and production excellence across critical business platforms. This role combines hands-on engineering expertise with leadership and governance responsibilities, ensuring services remain scalable, resilient, and aligned with business objectives.

As a key member of the technology team, you will champion SRE best practices, improve operational efficiency through automation, enhance observability, and lead the organization's approach to service reliability and incident management.

Key Responsibilities

Reliability Engineering & Production Excellence

  • Define, implement, and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets.
  • Drive data-led decisions balancing feature delivery with reliability goals.
  • Establish and enforce production readiness and operational excellence standards.

Observability & Monitoring

  • Design and implement end-to-end observability across applications and infrastructure.
  • Leverage tools such as Datadog, Azure Monitor, OpenTelemetry, and related platforms.
  • Translate monitoring and telemetry data into actionable insights that improve system performance and reliability.

Incident Management & Continuous Improvement

  • Lead the response to major incidents and critical outages.
  • Facilitate post-incident reviews and ensure remediation actions are completed.
  • Implement structured processes to reduce MTTR (Mean Time to Resolution) and improve service recovery.

Automation & Toil Reduction

  • Develop automation solutions, operational tooling, and runbooks.
  • Build self-healing and auto-remediation capabilities.
  • Integrate automation into monitoring and alerting workflows.

Governance & Standards

  • Define and promote reliability engineering standards across internal and vendor teams.
  • Assess production risks and make release readiness recommendations.
  • Ensure consistency in operational practices across multiple teams and stakeholders.

FinOps & Cloud Optimization

  • Collaborate with engineering and FinOps teams to balance reliability, performance, and cost.
  • Monitor and optimize cloud resource utilization while maintaining service quality.

Leadership & Stakeholder Engagement

  • Act as an SRE champion across the organization.
  • Mentor engineers and influence teams to adopt reliability best practices.
  • Present reliability metrics, risks, and strategic recommendations to senior stakeholders.

Required Skills & Experience

  • Extensive experience in Site Reliability Engineering (SRE) or a similar reliability-focused role.
  • Strong understanding of SLOs, SLIs, Error Budgets, and Production Readiness practices.
  • Hands-on experience with Azure, AWS, or other cloud platforms.
  • Proven expertise in:
    • Observability and monitoring solutions
    • Incident management and service restoration
    • Automation, scripting, and operational tooling
    • Reliability engineering practices
  • Experience with tools such as Datadog, Azure Monitor, OpenTelemetry, or comparable platforms.
  • Strong understanding of DevOps, SRE, and traditional Operations models.
  • Excellent stakeholder management and communication skills.

GCS is acting as an Employment Agency in relation to this vacancy.

€80,000.00 - €100,000.00
Per annum
EUR80000 - EUR100000 per annum

Republic of Ireland

Full Time

Added 05/08/2026
Reference: KK-SRE-IRE_1785928896

Senior Site Reliability Engineer

Republic of Ireland
Full Time

Other similar jobs

Senior Site Reliability Engineer (Infrastructure Focus)

Added 27/07/2026

Hybrid Dublin | Enterprise Cloud PlatformsWe're hiring experienced Site Reliability Engineers to support the evolution of large-scale cloud platforms used globally across critical production environments.Unlike traditional SRE opportunities focused heavily on Kubernetes, this role places greater emphasis on cloud infrastructure, systems engineering and platform reliability.ResponsibilitiesEngineer highly available cloud infrastructure platformsImprove reliability, resilience and operational excellenceDrive automation across provisioning, deployment and operational workflowsDevelop tools and services using Python and modern scripting technologiesSupport multi-region production environmentsParticipate in incident analysis and root cause investigationsReduce operational overhead through automation and engineering improvementsRequired ExperienceStrong cloud infrastructure experience across AWS, Azure or GCPExperience operating customer-facing or...

Learn more

Azure Site Reliability Engineer

Added 04/08/2026

Azure Site Reliability Engineer (SRE)Location: Glasgow / Knutsford (Hybrid- 2 days a week in office)Team: 6 UK / 5 IndiaEnvironment: Part of a wider multi‑cloud engineering organisation (Azure, AWS, GCP)Growth: Significant technical development opportunities across cloud engineering, automation, and platform build Role OverviewWe are looking for a hands‑on Azure SRE who can design, build, and automate enterprise‑grade Azure Landing Zones and cloud governance frameworks. This is not an application development role - it is a platform engineering role focused on controls, policies, guardrails, IaC, and DevOps automation.You will work as part of a global SRE function, collaborating with engineers in...

Learn more

Site Reliability Engineer

Added 17/07/2026

Site Reliability Engineer (SRE) - Level 3Job SummaryWe are seeking a Site Reliability Engineer (SRE) to build, maintain, and support reliable, scalable, and secure cloud platforms. The ideal candidate will have experience with cloud infrastructure, Kubernetes, automation, monitoring, and production support while working closely with engineering teams to improve platform reliability and operational excellence.Key ResponsibilitiesManage and support cloud infrastructure and Kubernetes environments.Automate infrastructure using Infrastructure as Code (IaC).Monitor, troubleshoot, and optimize production systems.Support CI/CD pipelines and deployment automation.Participate in incident response, root cause analysis, and reliability improvements.Collaborate with engineering teams on platform enhancements.Develop automation using scripting and DevOps best practices.Participate...

Learn more

Site Reliability Engineer

Added 16/07/2026

Site Reliability Engineer III (AI Platform)Location: Mount Laurel, NJ (Onsite)Duration: ContractExperience: 4+ yearsAbout the RoleWe are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This is an exciting opportunity to work on large-scale distributed systems, Kubernetes environments, and cloud-native platforms that power next-generation AI solutions.The ideal candidate has a strong background in cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation. Experience supporting production environments at scale is essential.ResponsibilitiesDesign, implement, and support highly available, scalable, and secure cloud infrastructure.Maintain...

Learn more

Site Reliability Engineer

Added 16/07/2026

Site Reliability Engineer III (AI Platform)Location: Mount Laurel, NJ (Onsite)Duration: ContractExperience: 4+ yearsAbout the RoleWe are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This is an exciting opportunity to work on large-scale distributed systems, Kubernetes environments, and cloud-native platforms that power next-generation AI solutions.The ideal candidate has a strong background in cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation. Experience supporting production environments at scale is essential.ResponsibilitiesDesign, implement, and support highly available, scalable, and secure cloud infrastructure.Maintain...

Learn more

MongoDB Site Reliability Engineer

Added 02/06/2026

MongoDB SRE (AVP) - Knutsford (Hybrid)Are you a MongoDB expert ready to step into a true engineering role? Join a global team modernising a large‑scale database estate and move beyond repetitive DBA work.What You'll DoOwn MongoDB operations end‑to‑end (clusters, sharding, replica sets, backups).Troubleshoot and resolve complex production issues across L1-L3.Build automation using Python, Ansible, TDD, Agile.Improve observability with better monitoring, alerting, and performance insights.Reduce toil by engineering tools and automation that transform the platform.Required SkillsDeep MongoDB administration expertise.Strong experience with Ops Manager and backup tooling.Solid troubleshooting and production support capability.SRE fundamentals and an automation‑first mindset.Hands‑on Python and Ansible experience.Observability experience...

Learn more

Microsoft SQL Database Site Reliability Engineer

Added 02/06/2026

Step into a high‑impact engineering role where you'll shape the future of Microsoft SQL operations at enterprise scale. As a Database SRE, you'll combine deep SQL Server expertise with modern SRE practices to build a more reliable, automated, and observable database platform for one of the world's largest financial institutions. What You'll DoLead SQL Engineering - Solve complex SQL Server 2016-2022 challenges across availability, tuning, performance, and architecture.Shape the MSSQL SRE practice - Influence standards, patterns, SLIs/SLOs, and operational models for the SQL estate.Act as the top technical escalation - Provide expert‑level guidance on incidents, root cause, and long‑term fixes.Drive...

Learn more

Azure Site Reliability Engineer

Added 29/05/2026

Azure Site Reliability Engineer (SRE)Location: Glasgow / Knutsford (Hybrid- 2 days a week in office)Team: 6 UK / 5 IndiaEnvironment: Part of a wider multi‑cloud engineering organisation (Azure, AWS, GCP)Growth: Significant technical development opportunities across cloud engineering, automation, and platform build Role OverviewWe are looking for a hands‑on Azure SRE who can design, build, and automate enterprise‑grade Azure Landing Zones and cloud governance frameworks. This is not an application development role - it is a platform engineering role focused on controls, policies, guardrails, IaC, and DevOps automation.You will work as part of a global SRE function, collaborating with engineers in...

Learn more

Site Reliability Specialist

Added 17/07/2026

Site Reliability Engineer (SRE) - Level 3Job SummaryWe are seeking a Site Reliability Engineer (SRE) to build, maintain, and support reliable, scalable, and secure cloud platforms. The ideal candidate will have experience with cloud infrastructure, Kubernetes, automation, monitoring, and production support while working closely with engineering teams to improve platform reliability and operational excellence.Key ResponsibilitiesManage and support cloud infrastructure and Kubernetes environments.Automate infrastructure using Infrastructure as Code (IaC).Monitor, troubleshoot, and optimize production systems.Support CI/CD pipelines and deployment automation.Participate in incident response, root cause analysis, and reliability improvements.Collaborate with engineering teams on platform enhancements.Develop automation using scripting and DevOps best practices.Participate...

Learn more

Site Acquisition Specialist

Added 08/07/2026

Site Acquisition Specialist is responsible for managing wireless site acquisition activities from initial site identification through project completion. This includes obtaining properly leases, coordinating zoning and permitting, working with municipalities and landlords, and processing all project milestones through AT&T's LMPS system. The role supports wireless network deployment projects for AT&T and T-Mobile while partnering with internal teams, contractors, and customers to remain on schedule.Required experience:Site leasingZoning applicationsPermittingMunicipal approvalsLandlord negotiationsUnderstanding wireless cell site development processExperience coordinating with engineering, construction, and general contractorsDay To Day Duties:Manage the complete site acquisition process for assigned projectsSubmit and track projects through AT&T LMPS systemObtain and...

Learn more

Senior Software Engineer/Data Platform Engineer (Databricks, Graph, APIs)

Added 30/04/2026

Senior Software Engineer / Data Platform Engineer (Databricks, Graph, APIs)Location: Philadelphia, PA The team sits within the network technology organisation and is responsible for building advanced data platforms that support digital twin capabilities across the access network. The group combines network design data, telemetry, mapping technologies, and graph intelligence to improve troubleshooting, planning, operational efficiency, and market competitiveness.The team works on highly scalable engineering products including large data pipelines, graph databases, APIs, and mapping platforms. Their work enables smarter network decisions, faster fault resolution, and better use of operational resources.This is a technically strong team focused on solving complex real-world...

Learn more

Senior Fiber Engineer

Added 05/08/2026

Fiber EngineerLocation: RemoteContract: 12+ MonthsOverviewWe are seeking an experienced Fiber Engineer to support fiber network testing, troubleshooting, and remediation activities across ISP and OSP environments. This role will focus on OTDR analysis, fiber characterization, bidirectional testing, and documenting findings to ensure network performance and reliability.Required Qualifications5+ years of experience in Fiber Optics or TelecommunicationsExperience with ISP and OSP fiber networksStrong expertise with OTDR testing and troubleshootingExperience performing CD/PMD, bidirectional, and power meter/laser testingKnowledge of fusion splicing, fiber repair, and remediationExperience with fiber end-face inspection and cleaningProficiency with Microsoft Word and ExcelStrong communication, documentation, and problem-solving skillsAbility to travel as required...

Learn more

Senior Software Engineer (Distributed System - C++)

Added 05/08/2026

SENIOR C++ SOFTWARE ENGINEER (DISTRIBUTED SYSTEMS - NEW DEVELOPMENT)Location: Munster (2 days a week Hybrid)Salary: €80,000 to €110,000(depends on experience) + Bonus/Benefits OverviewWe are hiring multiple C++ engineers to build new large-scale distributed software systems. This is a rare greenfield opportunity within a global engineering organisation, working on complex, high‑performance systems in Linux with a mix of new development and integration with long-standing platforms.What You'll Work OnDesigning and developing new distributed system componentsBuilding high-performance, scalable backend services in C++ (or C / Go / Rust backgrounds considered)Architecting and implementing new features within multi-node distributed environmentsIntegrating new systems with legacy platforms...

Learn more

Senior Software Engineer

Added 04/08/2026

Senior Golang Backend EngineerGolang | AWS | Kubernetes | Linux | Distributed SystemsAre you passionate about building high-performance backend systems that operate at enterprise scale? We're looking for a Senior Golang Backend Engineer to join a team developing secure, cloud-native applications and distributed services that power mission-critical platforms.This is an opportunity to work on complex engineering challenges involving microservices, authentication, cloud infrastructure, and high-throughput systems in a collaborative, fast-paced environment.What You'll DoDesign, develop, and maintain scalable backend applications using GolangBuild and enhance cloud-native microservices deployed in Kubernetes environmentsDevelop RESTful and gRPC APIs for enterprise applicationsDesign secure authentication and authorization solutionsCollaborate...

Learn more

Senior Software Engineer

Added 04/08/2026

Senior Software Engineer - Industrial Controls (CODESYS/TwinCAT)Location: Seattle/Bellevue, WA | HybridClient is looking for a hands-on Controls Software Engineer to develop the software platform behind its next-generation warehouse automation systems.In this role, you'll design and deploy reusable control software using CODESYS V3, develop robust state-machine based control logic, integrate real-world automation hardware, and help scale a modular platform across Amazon's global fulfillment network.What We're Looking For:Strong CODESYS V3 or Beckhoff TwinCAT experienceExpert-level Structured Text programmingState machine architecture and machine sequencing expertiseObject-oriented PLC development experienceEtherCAT hardware integration experience5+ years building production machine control systemsBonus Experience:Conveyor, robotics, warehouse automation, or material handling...

Learn more
At least 8 characters, 1 uppercase, 1 lowercase and 1 special character or number
Your file must be a doc, docx or pdf. No larger than 5MB.