Senior Site Reliability Engineer
(Hybrid, Cluj, Romania)
We’re Kingfisher, a team made up of over 74,000 passionate people who bring Kingfisher - and all our other brands: Castorama, B&Q, Screwfix, Brico Dépôt and Koçtaş - to life. That’s right, we’re big, but we have ambitions to become even bigger and even better. We want to become the leading home improvement company and grow the largest community of home improvers in the world. And that’s where you come in.
At Kingfisher, our customers come from all walks of life, and so do we. We want to ensure that all colleagues, future colleagues, and applicants to Kingfisher are treated equally regardless of age, gender, marital or civil partnership status, colour, ethnic or national origin, culture, religious belief, philosophical belief, political opinion, disability, gender identity, gender expression or sexual orientation.
We are open to flexible and agile working. Therefore, we offer colleagues a blend of working from home and our office, located in Cluj. Talk to us about how we can best support you!
At Kingfisher, we value the perspectives that any new team members bring, and we want to hear from you. We encourage you to apply for one of our roles, even if you do not feel you meet 100% of the requirements.
In return, we offer an inclusive environment, where what you can achieve is limited only by your imagination! We encourage new ideas, actively support experimentation, and strive to build an environment where everyone can be their best self.
We offer a competitive benefit package and plenty of opportunities to stretch and grow your career:
As a Senior Site Reliability Engineer, you will help improve the reliability, observability and resilience of Kingfisher’s digital platforms and services.
You will work closely with product engineering squads, incident management, platform teams, security, networks and observability specialists to reduce operational risk, improve service health and make reliability part of everyday engineering practice.
This is a hands-on senior engineering role. You will be expected to influence technical direction, drive improvements across teams, support better operational practices, and help product squads build and run services that are observable, scalable, secure and reliable.
Key Accountabilities / Responsibilities:
Act as a senior voice for reliability, observability and operational excellence
Define and implement SLIs, SLOs and error budgets
Improve observability with a focus on customer impact and actionable alerting
Partner with teams to reduce toil, automate operations and improve production readiness
Support incident management and post-mortems, ensuring actions are followed through
Help design resilient, scalable systems in cloud-native environments
Contribute to shared standards, tooling and SRE practices
Participate in on-call for major incidents
Required Skills & Experience:
Strong experience with SRE principles (SLOs, observability, incident response, automation)
Experience operating systems at scale in cloud environments
Solid understanding of distributed systems and reliability trade-offs
Hands-on experience with:
-Cloud: AWS / GCP / Azure
-Containers: Docker, Kubernetes
-IaC: Terraform
-CI/CD: GitLab (preferred), GitHub Actions, Jenkins
-Observability: Datadog
Scripting or programming skills (e.g. Python, Go, Bash, JavaScript)
Strong problem-solving and communication skills
DevOps mindset with focus on ownership and continuous improvement
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.