We're seeking a Senior Cloud Infrastructure Engineer for a 3-month contract engagement to join our Infrastructure team and take ownership of operational excellence and SRE toil work. This is a remote, hands-on, high-velocity role where you'll keep the lights on and reduce operational burden from day one.
This contract position exists to free up our existing team to focus on roadmap initiatives. By taking over day-to-day operational work and SRE toil, you'll enable one of our current engineers to tackle strategic projects. Your success means the platform runs smoothly while the team makes forward progress on critical initiatives.
You'll bring deep operational expertise to manage production systems, respond to operational needs, and—critically—build systems and automation that reduce toil over time. This role is ideal for an experienced SRE or infrastructure engineer who thrives on operational work, can quickly understand production systems, and naturally improves everything they touch.
Our platform powers genomics and laboratory workflows for customers in highly regulated environments. You'll work with modern infrastructure tooling (HashiCorp stack, AWS, Kubernetes patterns) while ensuring we meet the reliability, security, and compliance requirements our customers depend on.
What You'll Do
Operational Excellence & SRE Work (60%)
Keep the lights on: Monitor, respond to, and resolve production incidents and operational issues
Handle toil work: Manage routine operational tasks that currently consume team capacity (deployments, configuration changes, access management, maintenance windows)
Participate in on-call rotation: Share responsibility for after-hours production support
Respond to support escalations: Work with support and development teams to troubleshoot and resolve platform issues
Manage production changes: Execute and validate infrastructure changes in production environments
Maintain operational runbooks: Update and improve documentation for operational procedures
Perform system maintenance: Handle patches, upgrades, certificate renewals, and other recurring operational tasks
Ensure service reliability: Monitor system health, respond to alerts, and maintain SLAs
Toil Reduction & Automation (30%)
Identify automation opportunities: Spot repetitive manual work and build automation to eliminate it
Improve operational tooling: Create scripts, utilities, and self-service tools to reduce operational burden
Enhance monitoring and alerting: Improve observability to catch issues before they become incidents
Streamline deployment processes: Reduce friction and manual steps in release and deployment workflows
Build self-service capabilities: Enable developers to handle routine tasks without infrastructure team involvement
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in Canada.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in Canada, connecting you to thousands of jobs fast!
Find the best jobs in Canada, apply in 1 click and get a job today!