We are looking for an experienced software engineer to lead the development of our distributed edge compute platform. This role demands a deep understanding of hyperscalers like AWS, Azure, or GCP, and the ability to design and operate large-scale systems. The successful candidate will contribute to the platform's architecture, ensuring fault tolerance, consistency, and efficient traffic routing.
Responsibilities
Design and implement a massively distributed edge compute platform, ensuring high availability and fault tolerance.
Develop external-facing APIs, considering design, versioning, and interactions with customers and partners.
Optimize high-throughput L7 traffic routing for efficient data processing.
Utilize Kubernetes internals for resource management and container deployment.
Collaborate with a team of engineers to build and operate the platform, sharing knowledge and best practices.
Translate distributed systems complexity into clear architectural tradeoffs
Communicate infrastructure risks and mitigation strategies to executive stakeholders
Manage technical requirements from customers’ needs.
Collaborate effectively with AI/ML systems engineers, Hardware architects, Network engineering teams, Security and compliance teams
Present scalability models and reliability metrics in measurable, defensible terms
Mentor engineering teams on distributed systems best practices
Anticipate internal and external business challenges and / or regulatory issues and drive process, product or service improvements that create competitive advantage.
Stay updated with the latest trends and technologies in distributed systems and edge computing.
Establish design reviews and architectural governance standards
Ensure the platform's security and data privacy, adhering to industry standards and regulations.
Conduct code reviews and provide constructive feedback to maintain code quality.
Qualifications
Extensive experience (5+ years) with hyperscaler technologies (AWS, Azure, GCP) and managed services.
Have worked at or intimately with AWS, Azure, GCP, or similar-scale platforms - built internal cloud, compute, or infrastructure services
Proven track record in designing and operating large-scale distributed systems.
Strong fundamentals in distributed systems, including consistency models and fault tolerance.
Proficiency in API design and development, with an understanding of versioning and SDK foundations.
Experience with Kubernetes and containerization technologies for resource management.
Solid understanding of high-throughput traffic routing and network protocols.
Ability to work independently and manage complex technical tasks.
Excellent communication and collaboration skills, with a willingness to share knowledge.
Familiarity with security best practices and data privacy regulations.
Master's degree in Computer Science, Engineering, or a related field. PhD preferred.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in the US.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast!
Find the best jobs in the US, apply in 1 click and get a job today!