We’re a close-knit, ambitious team of 45, driven by a mission to make the world a better place by providing affordable, environmentally sustainable AI compute for training and deploying machine learning models at scale.
Build, maintain, and scale our observability platform to serve infrastructure and engineering teams.
Design and implement monitoring, logging, and alerting solutions to help identify and troubleshoot issues faster.
Support internal teams with dashboards, alerts, and data visualizations to improve reliability and performance.
Ensure systems are scalable and resilient, with automated lifecycle management.
Continuously improve observability tooling and practices to keep up with modern trends and best practices.
NetFlow, syslog, and related networking observability tools
Automation tools like Ansible or similar
Nice to Have:
Kubernetes monitoring & observability practices
Experience with OpenStack environments
Company equity - you’re in this with us!
Competitive salary and benefits, including health insurance, lunch benefit, annual budget to spend as you wish (i.e. sport, transport, wellness, culture)
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in the US.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast!
Find the best jobs in the US, apply in 1 click and get a job today!