Improve deployment reliability through progressive delivery approaches, automated testing, and deployment guardrails.
Champion GitOps and Infrastructure-as-Code practices across the engineering landscape.
Cloud Architecture & Operations
Provide technical leadership and guidance on cloud architecture, best practices, and operational excellence.
Support cloud-native services including:
- GKE (Google Kubernetes Engine)
- Cloud Run
- Cloud Functions
- Pub/Sub
- Cloud SQL
- Apigee
- VPC Networking
- IAM
- Cloud Monitoring
- Design resilient and observable systems aligned to enterprise security and reliability standards.
- API Platform & Marketplace Enablement
- Build and maintain infrastructure supporting enterprise API platforms and external marketplace services.
- Enable secure onboarding of internal and external API consumers.
- Establish standardized deployment and operational processes for API products and services.
- Support API gateway technologies such as Apigee and enterprise integration platforms.
- Drive incident management, root cause analysis, and post-incident reviews.
- Improve platform availability, scalability, and performance through automation and engineering improvements.
- Monitoring & Observability
- Implement enterprise monitoring, alerting, logging, and observability solutions.
Utilize tools including:
- Prometheus
- Grafana
- Google Cloud Monitoring
- ELK/Elastic Stack
- OpenTelemetry
- Proactively identify issues before they impact customers and engineering teams.
- Develop dashboards and operational reporting for platform health and reliability.
- DevSecOps & Security
- Embed security best practices throughout the software delivery lifecycle.
- Integrate automated vulnerability scanning, code analysis, policy enforcement, and compliance controls into CI/CD pipelines.
- Partner with Cyber Security teams to ensure compliance with internal and regulatory requirements.
- Support implementation of secure identity and access management controls across cloud platforms.
Operational Excellence:
- Lead resolution of complex production incidents and service disruptions.
- Continuously improve operational procedures, runbooks, and platform support processes.
- Drive automation initiatives that reduce manual effort and improve service reliability.
- Identify opportunities for cloud cost optimization while maintaining performance and resilience.
- Mentoring & Leadership
- Mentor junior engineers and promote DevOps best practices across the engineering community.
- Collaborate with architects, developers, testers, platform teams, and product owners.
- Drive engineering excellence through knowledge sharing, technical leadership, and continuous learning.
What You'll Need
- Essential Skills & Experience
- DevOps & Cloud Engineering
- Proven experience (10+ years) in DevOps, Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure roles.
- Significant experience operating highly available production systems at enterprise scale.
- GCP Expertise (Essential)
Strong hands-on experience with:
- GKE (Google Kubernetes Engine)
- Cloud Run
- Cloud Functions
- Pub/Sub
- Cloud SQL
- VPC Networking
- IAM
- Cloud Monitoring
- Secret Manager
- Cloud Storage
- Infrastructure as Code (IaC)
- Expert-level knowledge of Terraform.
- Experience building reusable infrastructure modules and automation frameworks.
- Experience with policy-as-code and infrastructure governance.
Kubernetes & Containerization
- Strong experience with:
- Kubernetes
- Docker
- Container security
- Service mesh technologies
- Experience operating production Kubernetes environments.
- CI/CD & Release Management
Experience with:
- GitHub Actions
- Jenkins
- GitLab CI/CD
- Harness
- ArgoCD (desirable)
Strong understanding of:
- GitOps
- Deployment strategies
- Release automation
- Configuration management
- API & Integration Technologies
Experience working with API Gateway technologies such as:
- Apigee
- Azure API Management
- Kong (desirable)
- Strong understanding of:
- REST APIs
- OpenAPI Specifications
- OAuth2
- JWT
- API Security Standards
- Observability & Monitoring
Experience with:
- Prometheus
- Grafana
- ELK Stack
- OpenTelemetry
- Google Cloud Monitoring
- Scripting & Automation
Proficiency in one or more of:
- Python [optional]
- Bash
- Go
- PowerShell
- Networking & Security
Pracyva is one of THE FASTEST GROWING specialized RecruitmentConsulting firm in UK and Europe… Pracyva Limited has local presence across UK , Europe ( Ireland, Netherlands , Poland ,Germany ), USA, Middle East and India , serving Top IT clients for large volumes ..We are currently hiring for our Reputed client Only candidates based in UK and eligible to work in UK are allowed
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in the UK.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in the UK, connecting you to thousands of jobs fast!
Find the best jobs in the UK, apply in 1 click and get a job today!