Designs, implements, and operates resilient L3/L4 network services across data centers and cloud environments. Owns well-defined projects end-to-end with autonomy, contributes to standards, and improves reliability through automation and observability. Collaborates primarily within the immediate team and with adjacent SRE, platform, and security teams to achieve team goals with measurable impact on availability, performance, and operational efficiency. Acts as a peer mentor, apply established techniques, and propose innovative improvements that enhance existing systems and processes.
Key Responsibilities
L3/L4 Network Topologies and Overlays:
Design IP addressing and routing policies (BGP/OSPF/IS-IS), EVPN/VXLAN fabrics, and SDN overlays across data centers and cloud providers.
Implement network segmentation, ECMP, high availability, and traffic engineering aligned to SLOs and capacity plans.
Execute changes via approved change processes; validate via pre/post checks and rollback plans.
Cloud Networking Implementation and Operations:
Configure and manage VPC/VNet constructs, subnets, routing tables, NAT, private/secure connectivity (TGW/VCN peering, PrivateLink/Service Endpoints), and cross-cloud interconnects.
Deploy and tune L4/L7 load balancing and service insertion (e.g., proxies, firewalls) to meet performance and resilience goals.
Ensure consistent policy and connectivity across on-prem and cloud, leveraging standard blueprints.
Network Security Controls and Compliance Guardrails:
Implement segmentation (VRFs, security zones), ACLs/policies, encryption (IPsec/TLS), and DDoS/WAF integrations following zero-trust principles.
Apply and maintain compliance controls and evidence (e.g., change records, rule reviews, least-privilege access) in partnership with Security/GRC.
Continuously review and optimize firewall and routing policies to reduce risk and improve performance.
Automation, CI/CD, and Infrastructure as Code:
Build and maintain Terraform/Ansible (or equivalent) modules/playbooks for repeatable network provisioning and configuration drift control.
Integrate network changes into CI/CD pipelines with linting, policy-as-code, unit/integration tests, and automated rollbacks.
Develop scripts and APIs for inventory, validation, and orchestration to improve change velocity and reduce error rates.
Monitoring, Telemetry, and SLOs:
Implement and operationalize SNMP/streaming telemetry, NetFlow/IPFIX, and flow logs for health, capacity, and security visibility.
Define and track SLIs/SLOs (availability, latency, packet loss) and build alerts, runbooks, and dashboards for proactive incident detection.
Perform trend and capacity analyses; recommend scaling actions and optimization.
Cross-Team Collaboration and Operations Excellence:
Partner with SRE, platform, and security teams on architecture reviews, incident response, post-incident improvements, and capacity/performance planning.
Participate in on-call rotations, execute change windows, and contribute to root-cause analysis and problem management.
Document and socialize operational models, guardrails, and standards; align designs with platform roadmaps.
Documentation, Reviews, and Standards:
Produce high-quality design docs, diagrams, and runbooks; maintain up-to-date inventories and topology maps.
Conduct and participate in design and change reviews; contribute to reference architectures, patterns, and reusable modules.
Share knowledge through peer mentoring, brown-bags, and code/document reviews to uplift team proficiency.
Core Responsibilities
Planning & Execution:
Independently manage work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements; proactively prioritize work and adapt to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
Collaboration & Partnership:
Collaborate across teams to align on expectations and achieve shared objectives; build and maintain a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships; actively listen to diverse perspectives and ask questions to ensure understanding of others.
Problem Solving:
Independently identify and address standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate; analyze data and/or information from multiple sources to troubleshoot standard and non-standard errors; contribute to knowledge sharing and best practices.
Continuous Learning:
Embrace continuous learning by actively seeking to build knowledge and new skills and/or tools, and staying current with industry trends and best practices; seek out and leverage feedback and training to improve skills; contribute to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement:
Develop ideas and recommend updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team; seek input from team members on alternative approaches and methods for improving work
Career Level - IC3
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.