Designs, implements, and operates resilient L3/L4 network services across data centers and cloud environments with broader influence beyond the immediate team. Leads moderately complex projects, drives automation and observability at scale, and contributes to standards and architectures. Mentors junior engineers, provides input on team decisions, and collaborates with SRE, platform, and security partners to deliver secure, compliant, and highly available network services with measurable impact on both team and group outcomes.
Key Responsibilities
L3/L4 Network Topologies and Overlays:
• Lead designs for IP addressing, routing policies (BGP/OSPF/IS-IS), EVPN/VXLAN fabrics, and SDN overlays spanning on-prem and cloud.
• Define HA, ECMP, and traffic-engineering patterns; validate through lab/sandbox testing, failure modeling, and staged rollouts.
• Author design docs and drive reviews involving related teams; de-risk changes with method-of-procedure (MOP), rollback, and verification plans.
Cloud Networking Implementation and Operations:
• Implement and operate VPC/VNet topologies, subnets, routing tables, NAT, Transit/Hub-and-Spoke (e.g., TGW/VCN), PrivateLink/service endpoints, and inter-cloud connectivity.
• Design and tune L4/L7 load balancing and service insertion (proxies, firewalls) to meet defined SLOs and cost targets.
• Ensure consistent network policy and connectivity across environments using standard blueprints and golden configs.
Network Security Controls and Compliance Guardrails:
• Engineer segmentation (VRFs/security zones), ACLs/policies, encryption (IPsec/TLS), and DDoS/WAF integrations aligned to zero-trust.
• Embed compliance guardrails (change control, access least privilege, evidence collection) and partner with security/GRC to pass audits.
• Periodically review rulebases and routing policies; eliminate shadow/duplicate rules and reduce attack surface.
Automation, CI/CD, and Infrastructure as Code:
• Develop reusable Terraform modules and Ansible playbooks; enforce code reviews, linting, policy-as-code, and test gates in pipelines.
• Build validation and drift-detection tooling using APIs and scripting (e.g., Python); integrate automated rollbacks and canary changes.
• Champion GitOps practices, configuration data modeling, and inventory source of truth.
Monitoring, Telemetry, and SLOs:
• Implement SNMP/streaming telemetry, NetFlow/IPFIX, and flow logs for health, capacity, and security analytics.
• Define SLIs/SLOs (availability, latency, loss, jitter) and build dashboards/alerts/runbooks; lead capacity/perf reviews and right-sizing actions.
• Correlate telemetry with incidents; drive problem management and reduction of repeat issues.
Cross-Team Collaboration and Operations Excellence:
• Partner with SRE, platform, and security teams on architecture reviews, incident response, post-incident remediation, and capacity planning.
• Lead portions of on-call response and complex change events; coordinate with related teams to minimize risk and downtime.
• Contribute to cross-LOB standards, reference architectures, and service onboarding patterns.
Documentation, Reviews, and Standards:
• Produce and maintain high-quality design docs, network diagrams, runbooks, and inventories; ensure traceability to requirements and controls.
• Facilitate and participate in design/change reviews; capture decisions and lessons learned; evolve patterns and modules accordingly.
• Mentor junior engineers via code/document reviews, pairing, and knowledge shares.
Core Responsibilities
Planning & Execution:
• Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately-sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
Collaboration & Partnership:
• Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
Problem Solving:
• Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions. Reviews, contributes to, and documents problem solving strategies.
Continuous Learning:
• Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices. Proactively seeks and leverages ongoing feedback and training to improve skills. Coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.
Continuous Improvement:
• Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluates the impact on key stakeholders. Solicits feedback from others on ideas for alternative approaches and methods for continued improvement.
Performance and Development:
• Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Career Level - IC4
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.