jobs in Trulyyy

全职 Senior Network Engineer (AI Cloud Ops) 工作, 薪水, Trulyyy 公司招聘中 - Ricebowl

Senior Network Engineer (AI Cloud Ops)

Trulyyy

Singapore

分享
保存

工作地点

  • Singapore Singapore

职位描述

岗位职责

Our client is a global technology company operating large-scale AI/HPC and data center infrastructure across multiple international markets. As its AI cloud capabilities continue to scale, the company is looking for an experienced AI Cloud Network Operations Engineer to operate and optimize high-performance network infrastructure supporting large-scale GPU computing environments.


Job Responsibilities

  • AI Network Operations — Monitor and operate large-scale data center and AI/HPC networks across switches, routers, optical links, bandwidth utilization and network health.
  • Incident & Performance Troubleshooting — Own network incidents and troubleshoot complex performance issues including latency/jitter, packet loss, GPU-to-GPU communication and NCCL throughput degradation.
  • Network Configuration & Changes — Execute and troubleshoot BGP, OSPF, VXLAN, EVPN and ECMP environments, including VLAN, routing, traffic isolation and bandwidth changes.
  • AI/HPC Fabric Operations — Support high-performance GPU cluster networking using InfiniBand and/or RoCEv2, including lossless Ethernet technologies such as PFC and ECN.
  • Monitoring & Automation — Improve network visibility, operational efficiency and reliability using monitoring platforms and Python/Go-based network automation.


Job Requirements

  • 5+ years of Network Operations / Network Engineering experience within large-scale Cloud, Internet, Data Center, Carrier or HPC environments.
  • Strong hands-on knowledge of BGP, OSPF, VXLAN, EVPN and ECMP, with independent troubleshooting capability on enterprise networking equipment.
  • Practical exposure to InfiniBand and/or RoCEv2, with understanding of PFC, ECN and lossless networking for high-performance workloads.
  • Experience operating large-scale infrastructure using vendors such as NVIDIA Spectrum/Quantum, Arista or Cisco, together with monitoring tools such as Prometheus, Grafana, Zabbix or equivalent.
  • Exposure to large-scale GPU clusters / AI infrastructure is highly advantageous, particularly NVIDIA H100/GB200 environments, NCCL/MPI, NVIDIA UFM, network automation or optical/DCI networks.


重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多