*This position is based onsite in Nashville, TN
We are seeking a Principal Core Infrastructure Engineer to help design, build, and operate the foundational systems behind a highly available, durable, and globally distributed object storage service. In this senior individual-contributor role, you will solve complex distributed-systems problems at massive scale and set technical direction across critical storage-platform capabilities.
You will work with engineers across storage, networking, security, control plane, and observability to deliver resilient systems that customers can trust with their most important data.
What You'll Do
* Architect and deliver core infrastructure for object storage, including metadata, data placement, replication, lifecycle management, and durability workflows.
* Lead the design of distributed systems that provide strong availability, consistency, performance, and operational simplicity at scale.
* Improve service reliability through fault isolation, automated remediation, disaster recovery design, capacity planning, and rigorous operational practices.
* Drive technical strategy for complex, cross-team initiatives; influence architecture and execution beyond your immediate team.
* Investigate and resolve challenging production issues, using deep systems expertise to prevent recurrence.
* Build tooling, automation, and observability that make the service easier to operate, diagnose, and evolve.
* Partner with security and compliance teams to embed secure-by-design practices into storage infrastructure.
* Establish engineering standards through design reviews, technical mentorship, and clear written communication.
* Strong software engineering experience in Java or a similar programming language, including designing, developing, testing, debugging, and maintaining production-quality services.
* Deep understanding of object-oriented programming principles, including abstraction, encapsulation, inheritance, polymorphism, composition, and appropriate application of common design patterns.
* Experience designing clean, extensible APIs and service interfaces, with an emphasis on maintainability, testability, backward compatibility, and operational safety.
* Demonstrated experience designing and building distributed systems, including services that operate across multiple hosts, availability domains, or regions.
* Strong knowledge of distributed-systems concepts such as replication, consistency, consensus, partition tolerance, leader election, idempotency, retries, failure handling, and eventual consistency.
* Experience designing systems for high availability, scalability, fault tolerance, and data durability.
* Experience with concurrent and multithreaded programming, performance analysis, memory management, and diagnosing latency or throughput bottlenecks in Java services.
* Proficiency with unit, integration, and end-to-end testing practices, including designing tests for failure scenarios and distributed-system edge cases.
* Experience using observability tools and practices, including metrics, logging, tracing, alerting, and production debugging.
* Ability to write clear technical design documents, evaluate architectural tradeoffs, and lead design reviews for complex services.
* Strong collaboration skills and experience partnering with engineering, security, networking, and operations teams.
* Proven ability to mentor engineers, raise engineering standards, and provide technical leadership without direct people-management responsibility.
Preferred Qualifications
* Experience with object storage systems, distributed databases, file systems, or other large-scale data platforms.
* Knowledge of storage-system concepts such as metadata management, object lifecycle operations, data placement, replication or erasure coding, and recovery from hardware or infrastructure failures.
* Experience designing systems that span multiple regions or availability domains.
* Experience operating services with demanding availability, durability, latency, or throughput requirements.
* Familiarity with RESTful APIs, asynchronous processing, message queues, and event-driven architectures.
* Experience with containerized environments, cloud infrastructure, and automated deployment pipelines.
* Knowledge of security practices for cloud services, including authentication, authorization, encryption, and secure data handling.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.