Job Description - Principal Architect/Lead- Observability
Description
We are seeking an experienced and visionary Principal Architect/Lead to take ownership of our observability practices and instrumentation standards. This role is crucial in ensuring we have a robust and efficient observability stack, enabling our engineering teams to monitor and optimize our distributed AI infrastructure effectively. The successful candidate will work closely with various teams, including Platform, SRE, UI, and UX, to establish best practices and drive the adoption of industry-leading observability tools and methodologies.
Responsibilities
Own and manage the observability stack, including tooling and instrumentation standards, across the entire codebase.
Work closely with SRE, engineering, and platform teams to embed observability practices from day one of a project's lifecycle.
Ensure proper instrumentation of services and push for high-quality observability data collection.
Stay up-to-date with the latest observability tooling and technologies, such as Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
Collaborate with UI/UX designers to create intuitive and informative observability dashboards and visualizations.
Develop and maintain OLAP databases, Kafka, ETL pipelines, and streaming platforms to support observability data processing and analysis.
Explore and implement graph-databases, ontologies, and semantic layers to enhance observability and knowledge-base capabilities.
Provide expertise and guidance to engineering teams on best practices for distributed AI infrastructure observability.
Actively participate in code reviews and provide feedback to ensure proper instrumentation and observability practices are followed.
Qualifications
Expert knowledge of observability tooling, including Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
Experience with OLAP databases, Kafka, ETL pipelines, and streaming platforms (Spark, Flink) is essential.
Must be hands-on in designing and implementing code; Experience with Kubernetes; Experience with cloud technologies
Familiarity with graph-databases, ontologies, semantic layers, and knowledgebases is highly advantageous.
Understanding of distributed systems and AI infrastructure at scale.
Ability to work closely with engineering teams and drive adoption of observability practices.
Excellent communication and collaboration skills to work effectively with cross-functional teams.
Experience in leading and mentoring a team of observability experts is desirable.
A proven track record of implementing and optimizing observability stacks in large-scale projects.
Strong problem-solving and analytical skills, with the ability to identify and resolve complex issues.
A passion for staying updated with the latest advancements in observability and monitoring technologies.
Master's degree in Computer Science, Engineering, or a related field; PhD preferred.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in the US.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast!
Find the best jobs in the US, apply in 1 click and get a job today!