jobs in Raydian Cloud

Raydian Cloud Hiring! Full Time Senior Network Engineer in Federal Territory - Ricebowl

Senior Network Engineer

Raydian Cloud

Undisclosed

KL City, Federal Territory

Share
Save

Working Location

  • SMART Tunnel Kuala Lumpur Federal Territory Malaysia

Job Description

Responsibilities

1. Key Responsibilities1.1 AI Cluster Network Architecture Design & Deployment
  • Design and deploy AI cluster network architectures based on NVIDIA Spectrum-X Ethernet and NVIDIA Quantum InfiniBand technologies.
  • Build Leaf-Spine and Clos network architectures covering GPU compute networks, storage networks, management networks, and cross-site connectivity.
  • Develop standard reference architectures for single-tenant and multi-tenant GPUaaS environments, including network segmentation, routing domains, and business zones.
  • Evaluate and select network hardware based on performance, cost, reliability, scalability, and security requirements.
1.2 RDMA, RoCEv2 & InfiniBand Operations and Optimization
  • Configure, operate, troubleshoot, and optimize RoCEv2 and InfiniBand lossless networks.
  • Perform in-depth RDMA performance tuning for distributed AI/ML workloads, including NCCL, MPI, and inter-GPU traffic.
  • Manage and optimize congestion control, ECN, PFC, QoS, and adaptive routing.
  • Troubleshoot communication bottlenecks caused by switches, NICs, DPUs, optical transceivers, firmware, and the Linux networking stack.
1.3 Cluster Integration & Performance Validation
  • Work closely with compute, storage, and data center teams to integrate and commission GPUs, NICs, DPUs, storage nodes, and related infrastructure.
  • Conduct network and cluster performance testing using tools such as ib_write_bw, iperf, NCCL test suites, and other benchmarking tools.
  • Establish performance baselines covering latency, bandwidth, packet loss, CRC errors, and other critical network metrics.
  • Support customer POCs, benchmark testing, technical validation, and production deployment.
1.4 Multi-Tenant Network Services
  • Implement tenant isolation using VLAN, VRF, VXLAN/EVPN, and related data center networking technologies.
  • Design and deploy customer connectivity solutions including dedicated circuits, cloud interconnects, IPsec/VPN, and other secure access services.
  • Operate network services including firewalls, load balancers, DNS, DHCP, and out-of-band management.
  • Develop standardized deployment and onboarding templates to improve tenant provisioning efficiency.
1.5 Automation, Monitoring & Operations
  • Develop network automation frameworks using Python/Go, Ansible, CI/CD, and Infrastructure as Code (IaC).
  • Implement Zero-Touch Provisioning (ZTP), configuration validation, automated deployment, and version rollback capabilities.
  • Build comprehensive telemetry and monitoring covering port status, optical power levels, congestion, RDMA metrics, and other critical infrastructure indicators.
  • Develop and maintain operational runbooks, escalation procedures, incident response processes, and Root Cause Analysis (RCA) templates.
  • Support 24×7 production operations and participate in on-call and emergency incident response.
1.6 Project Delivery & Technical Documentation
  • Participate in data center planning, including rack layout, network cabling, IP addressing, hardware BOMs, and infrastructure design.
  • Produce High-Level Design (HLD), Low-Level Design (LLD), Method of Procedure (MOP), test plans, test reports, and other technical documentation.
  • Provide technical support for major production incidents and participate in project deployments across the Asia-Pacific region.
  • Coordinate with data center operators, hardware vendors, customers, and internal technical teams to ensure successful project delivery.



2. Requirements2.1 Essential Requirements
  • Bachelor’s degree or above in Computer Science, Information Technology, Electronics, Telecommunications, or a related discipline.
  • Minimum 5 years of experience in large-scale data center network design, deployment, and troubleshooting.
  • Minimum 3 years of hands-on experience with InfiniBand, RoCEv2, and RDMA.
  • Proven experience supporting AI/GPU clusters, High-Performance Computing (HPC), or large-scale distributed computing environments.
2.2 Required Technical Skills
  • Strong knowledge of BGP, ECMP, VXLAN/EVPN, QoS, high availability, and modern data center networking technologies.
  • Hands-on experience with NVIDIA/Mellanox networking products, including:
  • NVIDIA Spectrum switches
  • NVIDIA ConnectX NICs
  • NVIDIA BlueField DPUs
  • Familiarity with networking equipment from major vendors such as Arista, Cisco, and Juniper.
  • Strong Linux networking troubleshooting skills using tools such as tcpdump, ethtool, and related utilities.
  • Experience with network monitoring and observability platforms such as Prometheus, Grafana, and ELK.
2.3 Preferred / Additional Qualifications
  • Hands-on experience deploying NVIDIA DGX, HGX, GB-series AI systems, Spectrum-X, or Quantum InfiniBand.
  • Knowledge of NCCL, GPUDirect RDMA, and large-scale AI/ML workload traffic patterns, including MoE workloads.
  • Familiarity with Kubernetes networking, Slurm, parallel file systems, NVMe-oF, and related AI/HPC technologies.
  • Experience with Python/Go, Ansible, CI/CD, and network automation.
  • NVIDIA, Arista, or other relevant vendor certifications are an advantage.



3. Performance Indicators & Additional Requirements3.1 Key Performance Indicators
  • Deliver network projects on schedule with complete and accurate cabling, configuration, testing, and acceptance documentation.
  • Ensure RDMA and NCCL performance meets defined customer and production requirements.
  • Respond rapidly to network incidents and accurately identify root causes, with complete RCA documentation and corrective action plans.
  • Continuously improve network automation, operational efficiency, tenant security, and network isolation.
3.2 Additional Requirements
  • Willingness to travel frequently across Asia-Pacific countries.
  • Able to support scheduled night-time maintenance, emergency troubleshooting, and on-call operations when required.
  • Comfortable working in a fast-paced project environment with multiple stakeholders.
  • Strong communication and coordination skills when working with data center operators, hardware vendors, enterprise customers, and internal engineering teams.



4. Job Summary

The AI Cluster Network Engineer is a key technical role responsible for designing, deploying, optimizing, and operating high-performance networks for AI/GPU computing and GPUaaS environments.

The role is deeply focused on NVIDIA’s advanced networking technologies, including Spectrum-X Ethernet, Quantum InfiniBand, RoCEv2, RDMA, and high-performance cluster networking. The successful candidate will combine strong traditional data center networking expertise with hands-on experience in AI/HPC networking and distributed workload optimization.

In addition to network architecture and performance tuning, the role covers multi-tenant network isolation, automation, observability, production operations, customer POCs, and project delivery across the Asia-Pacific region.

This position is ideal for a senior network engineer who wants to specialize in AI infrastructure, GPU clusters, high-performance computing, and next-generation data center networking, with a strong focus on commercial, production-grade GPUaaS deployments.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More