Responsibilities:
- Implement, monitor and maintain software systems to ensure high availability and performance
- Automate repetitive tasks and processes to improve efficiency and reduce manual errors
- Develop and maintain systems and tools for monitoring and analyzing system performance and resource utilization
- Collaborate with software development teams to ensure that services are deployed and operated in a reliable and scalable manner
- Respond to and resolve incidents and outages in a timely and effective manner
- Post-incident analysis and related documentation
- Participate in on-call rotation to provide 24x7 support for critical systems
- Write business processes, business rules, or SRE best practices
- Support usage/cost allocation audits
- Support release roadmap planning
- Conduct chaos engineering activities
- Provide training on third-party platform competencies
- Load testing and other capacity management tasks
Qualifications:
- Bachelor’s degree in Computer Science, Information Systems or equivalent technical field
- At least 3 years of experience in a SRE or DevOps role
- Strong knowledge of Linux/Unix systems administration
- Experience with automation tools (e.g. Puppet, Chef, Ansible, etc)
- Experience with containerization and orchestration technologies such as Docker and Kubernetes
- Strong scripting skills in languages such as Python, Bash, or Perl
- Familiarity with cloud platforms, especially AWS and GCP
- Experience with Agile development environments
- Have very good interpersonal and communication skills and love to be part of a team
*We are an equal opportunity employer and value diversity. All employment decisions are made without regard to age, gender, disability, race, ethnicity, religion, sexual orientation, or any other protected characteristic. We encourage applicants from all backgrounds to apply.*