Our client is a fast-growing AI infrastructure company building next-generation GPU-powered AI platforms. They operate high-performance AI data centers that support large-scale machine learning, AI model training and inference workloads. The environment is highly technical, focusing on low-latency networking, scalability, automation and operational excellence.
Primary Responsibilities
Design, deploy and support high-performance data center network infrastructure for AI and GPU compute environments.
Build, configure, and optimize Spine-Leaf network architectures for large-scale GPU clusters.
Deploy and manage high-speed Ethernet and/or InfiniBand fabrics supporting AI/HPC workloads.
Configure and maintain Layer 2 and Layer 3 networking technologies, including BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC.
Monitor, troubleshoot, and resolve complex network issues across compute, storage, and AI infrastructure.
Optimize network performance, latency, and throughput to support distributed AI training and inference.
Perform firmware upgrades, network maintenance, capacity planning, and lifecycle management.
Collaborate closely with Platform, Infrastructure, Linux, DevOps, and AI Engineering teams to deliver scalable AI infrastructure.
Develop and maintain network documentation, operational procedures, and technical runbooks.
Implement network monitoring, alerting, and automation to improve operational efficiency.
Participate in incident response and on-call support for production environments.
Evaluate and recommend new networking technologies to improve scalability, reliability, and performance.
What We're Looking For
Bachelor's Degree in Computer Science, Computer Engineering, Information Technology, or a related discipline.
5+ years of experience designing or supporting enterprise or data center network infrastructure.
Strong understanding of modern data center networking principles
Hands-on experience with routing and switching protocols such as BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC.
Experience managing high-speed Ethernet networks (25G/40G/100G/200G/400G).
Exposure to AI, HPC, GPU clusters, or large-scale compute environments would be highly advantageous.
Experience working with networking platforms such as Cisco Nexus, Arista, Juniper, NVIDIA Spectrum, or Mellanox.
Good understanding of network security concepts including segmentation, ACLs, and firewall policies.
Familiarity with Linux networking fundamentals and troubleshooting.
Experience with network automation using Python, Ansible, REST APIs, or similar tools is an advantage.
Knowledge of technologies such as InfiniBand, RoCEv2, RDMA, GPUDirect, SONiC, or Cumulus Linux is a plus.
Strong analytical, troubleshooting, and problem-solving skills.
Excellent communication skills with the ability to collaborate effectively across cross-functional engineering teams.
Comfortable working in a fast-paced, high-availability production environment supporting mission-critical AI infrastructure.
Click on Apply now to find out more about this opportunity and other available positions.
EA License: 22C1396
EA Personnel: R1551466
Vouch is a specialist recruitment firm that provides extensive workforce solutions across various domains in Asia. Established in 2023, Vouch is a rising agency that aims to bridge the evident talent gap in the market. Equipped with our expertise & knowledge, we assist clients to identify exceptional talent & create career opportunities for candidates. Our approach focuses on building meaningful relationships & collaboration to forge partnerships.