jobs in Trulyyy

Trulyyy Hiring! Full Time Senior Network Engineer (AI Cloud Ops) in - Ricebowl

Senior Network Engineer (AI Cloud Ops)

Trulyyy

Singapore

Share
Save

Working Location

  • Singapore Singapore

Job Description

Responsibilities

Our client is a global technology company operating large-scale AI/HPC and data center infrastructure across multiple international markets. As its AI cloud capabilities continue to scale, the company is looking for an experienced AI Cloud Network Operations Engineer to operate and optimize high-performance network infrastructure supporting large-scale GPU computing environments.


Job Responsibilities

  • AI Network Operations — Monitor and operate large-scale data center and AI/HPC networks across switches, routers, optical links, bandwidth utilization and network health.
  • Incident & Performance Troubleshooting — Own network incidents and troubleshoot complex performance issues including latency/jitter, packet loss, GPU-to-GPU communication and NCCL throughput degradation.
  • Network Configuration & Changes — Execute and troubleshoot BGP, OSPF, VXLAN, EVPN and ECMP environments, including VLAN, routing, traffic isolation and bandwidth changes.
  • AI/HPC Fabric Operations — Support high-performance GPU cluster networking using InfiniBand and/or RoCEv2, including lossless Ethernet technologies such as PFC and ECN.
  • Monitoring & Automation — Improve network visibility, operational efficiency and reliability using monitoring platforms and Python/Go-based network automation.


Job Requirements

  • 5+ years of Network Operations / Network Engineering experience within large-scale Cloud, Internet, Data Center, Carrier or HPC environments.
  • Strong hands-on knowledge of BGP, OSPF, VXLAN, EVPN and ECMP, with independent troubleshooting capability on enterprise networking equipment.
  • Practical exposure to InfiniBand and/or RoCEv2, with understanding of PFC, ECN and lossless networking for high-performance workloads.
  • Experience operating large-scale infrastructure using vendors such as NVIDIA Spectrum/Quantum, Arista or Cisco, together with monitoring tools such as Prometheus, Grafana, Zabbix or equivalent.
  • Exposure to large-scale GPU clusters / AI infrastructure is highly advantageous, particularly NVIDIA H100/GB200 environments, NCCL/MPI, NVIDIA UFM, network automation or optical/DCI networks.


Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More