jobs in Tencent

Tencent Hiring! Full Time AI Infrastructure Network Engineer in - Ricebowl

AI Infrastructure Network Engineer

Tencent

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

About the Role


We are looking for an experienced AI Infrastructure Network Engineer / Architect to support the planning, delivery, acceptance, and ongoing operations of large-scale GPU cluster infrastructure for our internal AI business teams.


The AI Compute Centre sits within Tencent's Overseas IT department, acting as the bridge between internal AI infrastructure demand and the external resources that fulfill it. We own the full lifecycle of AI compute clusters — requirement gathering, capacity planning, architecture review, delivery coordination, and day-to-day operations — across regions worldwide.


In this role, you will act as the key technical bridge between internal AI business stakeholders and external GPU cluster vendors, network equipment vendors, colocation partners, and managed infrastructure providers. You will be responsible for understanding business and technical requirements, translating them into clear infrastructure and network requirements, overseeing vendor design and delivery, validating the final implementation, and ensuring the cluster is successfully handed over to business users and operated reliably in production.


This role is ideal for someone who has strong data center networking and AI infrastructure knowledge, hands-on experience with GPU cluster networks, and the ability to coordinate across business teams, engineering teams, and external vendors.


Key Responsibilities

  • Work closely with internal AI business teams, AI platform teams, and infrastructure stakeholders to understand AI workload requirements, including training, inference, data access, storage, bandwidth, latency, scalability, and reliability requirements.
  • Translate business and platform requirements into clear technical requirements for GPU cluster infrastructure, especially networking architecture, interconnect design, capacity planning, and operational requirements.
  • Engage with external vendors, including GPU cluster providers, network equipment vendors, colocation providers, and system integrators, to review proposed solutions and ensure they meet business and technical expectations.
  • Review and validate GPU cluster network architecture, including InfiniBand / RoCE networks, Spine-Leaf / Clos topologies, front-end service networks, storage networks, and out-of-band management networks.
  • Participate in the planning of foundational network resources, including IP addressing, VLAN/VXLAN, routing, bandwidth capacity, network segmentation, and management access.
  • Review vendor-provided architecture documents, topology diagrams, implementation plans, configuration standards, test plans, and operational runbooks.
  • Monitor and drive vendor delivery progress, identify technical risks or delivery gaps, and coordinate corrective actions to ensure the cluster is delivered on time and according to requirements.
  • Define and execute acceptance criteria for GPU cluster delivery, including network connectivity, bandwidth, latency, redundancy, congestion control, fault tolerance, GPU-to-GPU communication performance, and cluster-level benchmark validation.
  • Support performance validation and troubleshooting for AI training and inference workloads, including network-related issues affecting distributed training, collective communication, storage access, or application performance.
  • Coordinate the handover of validated GPU clusters to internal business and platform teams, including documentation, knowledge transfer, operational procedures, and post-handover support.
  • Own or support daily operations and maintenance of AI infrastructure networks, including incident response, troubleshooting, configuration review, capacity monitoring, change management, and continuous improvement.
  • Develop and maintain technical documentation, including HLD/LLD, network topology diagrams, configuration baselines, acceptance reports, SOPs, and operational playbooks.


Qualifications

  • Bachelor's degree or above in Computer Science, Telecommunications, Electrical Engineering, Information Technology, or a related technical field.
  • 5+ years of experience in data center networking, cloud infrastructure networking, backbone networking, or large-scale infrastructure delivery.
  • Solid understanding of AI infrastructure and GPU cluster networking requirements, especially for distributed training and high-performance computing workloads.
  • Familiarity with high-performance GPU cluster interconnect technologies such as InfiniBand and/or RoCEv2.
  • Good understanding of AI cluster components, including GPU servers, high-speed networking, storage systems, management networks, and cluster orchestration platforms.
  • Experience working with infrastructure vendors, system integrators, colocation providers, or cloud infrastructure providers.
  • Ability to review and challenge vendor designs, implementation plans, test results, and operational documents from both architecture and production-readiness perspectives.
  • Experience defining or executing infrastructure acceptance tests, including network performance, redundancy, failure recovery, connectivity, and stability validation.
  • Strong troubleshooting skills across L2/L3 networking, Linux networking, TCP/IP, routing, DNS, firewall/security rules, and data center connectivity.
  • Strong project coordination skills, with the ability to track vendor delivery progress, identify risks, drive issue resolution, and communicate clearly with both technical and non-technical stakeholders.
  • Excellent written and verbal communication skills, with the ability to translate business needs into technical requirements and explain technical trade-offs to business teams.
  • Self-motivated, responsible, detail-oriented, and comfortable working in a fast-paced environment with multiple internal and external stakeholders.
  • Customer-oriented mindset and strong ownership in ensuring business teams receive stable, performant, and production-ready AI infrastructure.


Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More