Job Role: Mid/Senior AI Network & Security Engineer
Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service
Location: Kulai, Johor, Malaysia
Job Type: Full Time – On Site
Experience: 8+ years of relevant network engineering experience, preferably within data-center, cloud, service-provider or large-scale infrastructure environments.
JOB DESCRIPTION
AI & Data Center Network Infrastructure
- Design, implement, operate and maintain highly available network infrastructure supporting AI/GPU clusters, data centres and AI Cloud services.
- Support high-performance GPU network fabrics including NVIDIA InfiniBand and Ethernet/RoCE-based architectures.
- Design and operate data-center networking technologies including routing, switching, VLANs, BGP, EVPN/VXLAN and high-availability architectures.
- Configure and support network infrastructure across compute, storage, management, out-of-band (OOB), customer and external connectivity networks.
- Support LAN, WAN, VPN and private interconnect connectivity between data centres, customers, partners and cloud environments.
- Participate in network architecture, capacity planning and infrastructure expansion activities for new AI/GPU clusters.
Network Operations & Performance
- Monitor network and AI fabric availability, throughput, latency, utilisation, errors and congestion to ensure infrastructure meets performance and SLA requirements.
- Troubleshoot complex connectivity and performance issues across switches, ConnectX NICs, DPUs and SuperNICs, servers, host networking and GPU workloads.
- Work with Systems, Platform and Operations teams to identify network-related issues impacting distributed GPU workloads.
- Develop expertise in AI networking technologies including InfiniBand, RDMA, RoCEv2, lossless Ethernet, PFC, ECN, QoS and congestion management.
- Support AI fabric monitoring and management platforms such as NVIDIA UFM, NetQ or equivalent tools.
- Support network validation, commissioning and performance testing for new GPU clusters and infrastructure deployments.
Network Security
- Design, implement and operate network security infrastructure including firewalls, VPNs, ACLs, segmentation, NAT, IPS and secure connectivity.
- Configure and manage Fortinet/FortiGate or equivalent enterprise firewall platforms.
- Implement appropriate network segmentation across compute, management, storage, OOB, customer and external-facing environments.
- Support security hardening, vulnerability remediation and compliance requirements.
- Work closely with Security and Risk teams during security incidents, assessments and audits.
Network Automation
- Develop and maintain network automation using technologies such as Python, Ansible, APIs, DCIM and Infrastructure as Code.
- Automate configuration deployment, backups, compliance validation, provisioning and routine operational activities.
- Support CI/CD and controlled network change processes to improve consistency, reliability and auditability.
- Work with platform teams to integrate network and fabric telemetry into monitoring platforms.
Operations & Reliability
- Provide technical leadership for complex network incidents, outages and performance degradation.
- Perform root-cause analysis and drive corrective and preventive actions.
- Analyse network performance trends, capacity, hardware utilisation and growth requirements.
- Plan and test redundancy, failover and recovery mechanisms.
- Participate in change management, maintenance, upgrades and lifecycle management activities.
- Participate in the operational standby/on-call roster supporting 24×7 AI Cloud services.
- Develop and maintain HLD/LLD, network diagrams, SOPs, MOPs, EOPs, troubleshooting runbooks and technical documentation.
JOB REQUIREMENTS
- Bachelor's degree in Network Engineering, Computer Science, Information Technology or related discipline, or equivalent practical experience.
- 8+ years of relevant network engineering experience, preferably within data-center, cloud, service-provider or large-scale infrastructure environments.
- Strong hands-on knowledge of TCP/IP, routing, switching, VLANs, BGP and network redundancy/high availability.
- Strong experience designing, implementing and troubleshooting production network infrastructure.
- Hands-on experience with enterprise/data-center switching and routing platforms such as Juniper, Cisco, Arista or equivalent.
- Experience with enterprise firewall technologies, network segmentation, ACLs, VPNs and security policies.
- Solid understanding of advanced networking technologies, particularly those related to AI would be highly advantageous.
- Strong communication skills, both written and verbal.
- Excellent problem-solving and analytical skills.
- Ability to work independently and as part of a team.
- Willingness to work site-based in Johor and to participate in a 24×7 escalation roster.
Key Competencies
- Hands-on experience with NVIDIA InfiniBand, Spectrum Ethernet Platform, and/or RDMA over Converged Ethernet (RoCE) preferred.
- Good understanding of Linux and host networking.
- Enterprise & AI network delivery (LAN, WAN, WLAN, VPN).
- Secure connectivity (site-to-site VPN, MPLS, private interconnects).
- Network automation & NetDevOps (IaC, Ansible, Python, CI/CD).
- Network security (firewalls, VLANs, ACLs, zero-trust).
- Operations, monitoring & incident response.
- AI networking tech (InfiniBand, Spectrum, RoCE).
- Linux networking fundamentals.
- Cross-team collaboration & mentoring.
- Certifications such as NVIDIA-Certified InfiniBand or Networking Professional, JNCIP or JNCIE, CCNP or CCIE, Fortinet NSE 4 and above, CISSP or CISM.