jobs in Raydian Cloud

全职 Senior Network Engineer 工作, 薪水, Raydian Cloud Federal Territory 公司招聘中 - Ricebowl

Senior Network Engineer

Raydian Cloud

Undisclosed

KL City, Federal Territory

分享
保存

工作地点

  • Kuala Lumpur Federal Territory Malaysia

职位描述

岗位职责

1. Key Responsibilities1.1 AI Cluster Network Architecture Design & Deployment
  • Design and deploy AI cluster network architectures based on NVIDIA Spectrum-X Ethernet and NVIDIA Quantum InfiniBand technologies.
  • Build Leaf-Spine and Clos network architectures covering GPU compute networks, storage networks, management networks, and cross-site connectivity.
  • Develop standard reference architectures for single-tenant and multi-tenant GPUaaS environments, including network segmentation, routing domains, and business zones.
  • Evaluate and select network hardware based on performance, cost, reliability, scalability, and security requirements.
1.2 RDMA, RoCEv2 & InfiniBand Operations and Optimization
  • Configure, operate, troubleshoot, and optimize RoCEv2 and InfiniBand lossless networks.
  • Perform in-depth RDMA performance tuning for distributed AI/ML workloads, including NCCL, MPI, and inter-GPU traffic.
  • Manage and optimize congestion control, ECN, PFC, QoS, and adaptive routing.
  • Troubleshoot communication bottlenecks caused by switches, NICs, DPUs, optical transceivers, firmware, and the Linux networking stack.
1.3 Cluster Integration & Performance Validation
  • Work closely with compute, storage, and data center teams to integrate and commission GPUs, NICs, DPUs, storage nodes, and related infrastructure.
  • Conduct network and cluster performance testing using tools such as ib_write_bw, iperf, NCCL test suites, and other benchmarking tools.
  • Establish performance baselines covering latency, bandwidth, packet loss, CRC errors, and other critical network metrics.
  • Support customer POCs, benchmark testing, technical validation, and production deployment.
1.4 Multi-Tenant Network Services
  • Implement tenant isolation using VLAN, VRF, VXLAN/EVPN, and related data center networking technologies.
  • Design and deploy customer connectivity solutions including dedicated circuits, cloud interconnects, IPsec/VPN, and other secure access services.
  • Operate network services including firewalls, load balancers, DNS, DHCP, and out-of-band management.
  • Develop standardized deployment and onboarding templates to improve tenant provisioning efficiency.
1.5 Automation, Monitoring & Operations
  • Develop network automation frameworks using Python/Go, Ansible, CI/CD, and Infrastructure as Code (IaC).
  • Implement Zero-Touch Provisioning (ZTP), configuration validation, automated deployment, and version rollback capabilities.
  • Build comprehensive telemetry and monitoring covering port status, optical power levels, congestion, RDMA metrics, and other critical infrastructure indicators.
  • Develop and maintain operational runbooks, escalation procedures, incident response processes, and Root Cause Analysis (RCA) templates.
  • Support 24×7 production operations and participate in on-call and emergency incident response.
1.6 Project Delivery & Technical Documentation
  • Participate in data center planning, including rack layout, network cabling, IP addressing, hardware BOMs, and infrastructure design.
  • Produce High-Level Design (HLD), Low-Level Design (LLD), Method of Procedure (MOP), test plans, test reports, and other technical documentation.
  • Provide technical support for major production incidents and participate in project deployments across the Asia-Pacific region.
  • Coordinate with data center operators, hardware vendors, customers, and internal technical teams to ensure successful project delivery.



2. Requirements2.1 Essential Requirements
  • Bachelor’s degree or above in Computer Science, Information Technology, Electronics, Telecommunications, or a related discipline.
  • Minimum 5 years of experience in large-scale data center network design, deployment, and troubleshooting.
  • Minimum 3 years of hands-on experience with InfiniBand, RoCEv2, and RDMA.
  • Proven experience supporting AI/GPU clusters, High-Performance Computing (HPC), or large-scale distributed computing environments.
2.2 Required Technical Skills
  • Strong knowledge of BGP, ECMP, VXLAN/EVPN, QoS, high availability, and modern data center networking technologies.
  • Hands-on experience with NVIDIA/Mellanox networking products, including:
  • NVIDIA Spectrum switches
  • NVIDIA ConnectX NICs
  • NVIDIA BlueField DPUs
  • Familiarity with networking equipment from major vendors such as Arista, Cisco, and Juniper.
  • Strong Linux networking troubleshooting skills using tools such as tcpdump, ethtool, and related utilities.
  • Experience with network monitoring and observability platforms such as Prometheus, Grafana, and ELK.
2.3 Preferred / Additional Qualifications
  • Hands-on experience deploying NVIDIA DGX, HGX, GB-series AI systems, Spectrum-X, or Quantum InfiniBand.
  • Knowledge of NCCL, GPUDirect RDMA, and large-scale AI/ML workload traffic patterns, including MoE workloads.
  • Familiarity with Kubernetes networking, Slurm, parallel file systems, NVMe-oF, and related AI/HPC technologies.
  • Experience with Python/Go, Ansible, CI/CD, and network automation.
  • NVIDIA, Arista, or other relevant vendor certifications are an advantage.



3. Performance Indicators & Additional Requirements3.1 Key Performance Indicators
  • Deliver network projects on schedule with complete and accurate cabling, configuration, testing, and acceptance documentation.
  • Ensure RDMA and NCCL performance meets defined customer and production requirements.
  • Respond rapidly to network incidents and accurately identify root causes, with complete RCA documentation and corrective action plans.
  • Continuously improve network automation, operational efficiency, tenant security, and network isolation.
3.2 Additional Requirements
  • Willingness to travel frequently across Asia-Pacific countries.
  • Able to support scheduled night-time maintenance, emergency troubleshooting, and on-call operations when required.
  • Comfortable working in a fast-paced project environment with multiple stakeholders.
  • Strong communication and coordination skills when working with data center operators, hardware vendors, enterprise customers, and internal engineering teams.



4. Job Summary

The AI Cluster Network Engineer is a key technical role responsible for designing, deploying, optimizing, and operating high-performance networks for AI/GPU computing and GPUaaS environments.

The role is deeply focused on NVIDIA’s advanced networking technologies, including Spectrum-X Ethernet, Quantum InfiniBand, RoCEv2, RDMA, and high-performance cluster networking. The successful candidate will combine strong traditional data center networking expertise with hands-on experience in AI/HPC networking and distributed workload optimization.

In addition to network architecture and performance tuning, the role covers multi-tenant network isolation, automation, observability, production operations, customer POCs, and project delivery across the Asia-Pacific region.

This position is ideal for a senior network engineer who wants to specialize in AI infrastructure, GPU clusters, high-performance computing, and next-generation data center networking, with a strong focus on commercial, production-grade GPUaaS deployments.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多