About the Client
Our Client is a publicly traded artificial intelligence (AI) and high-performance computing (HPC) infrastructure provider.
About the Role
- Lead the end-to-end hardware architecture design and planning for AI Computing Centres (AI Data Centres), encompassing compute, networking, storage, rack, power, and cooling infrastructure.
- Design and optimize GPU cluster architectures, including multi-GPU interconnects (NVLink/NVSwitch), high-speed inter-node networking (InfiniBand/RoCE), and non-blocking network topologies such as Spine-Leaf and Fat-Tree.
- Define hardware standards, technical specifications, and deployment architectures for large-scale AI and high-performance computing environments.
- Evaluate, select, and validate servers and key hardware components, including CPUs, GPUs, memory, storage, network interface cards (NICs), and power systems.
- Develop and maintain Bill of Materials (BOM), sizing models, capacity plans, and hardware configuration standards.
- Perform hardware benchmarking, performance testing, compatibility validation, and capacity planning to ensure optimal infrastructure performance.
- Provide technical leadership for hardware deployment activities, including rack installation, cabling, commissioning, power-on testing, acceptance testing, and troubleshooting.
- Collaborate with data centre, network, facilities, and operations teams to ensure successful implementation of AI infrastructure solutions.
- Manage hardware vendors and technology partners to ensure solution quality, delivery timelines, and technical compliance.
- Drive architecture reviews and recommend technology improvements to enhance performance, scalability, reliability, and cost efficiency.
- Track industry trends and emerging technologies including next-generation GPUs/NPUs, Co-Packaged Optics (CPO), liquid cooling solutions, and AI computing infrastructure innovations.
- Support the development of technical proposals, solution documentation, implementation plans, and customer presentations.
What We're Looking For
- Bachelor's Degree in Computer Science, Electronics Engineering, Communications Engineering, Automation, or a related discipline.
- Minimum 5 years of experience in server, data centre infrastructure, hardware architecture, or AI computing solutions.
- Strong understanding of server hardware architecture, including CPU, GPU, memory, PCIe, NVLink, and accelerator technologies.
- Hands-on experience with GPU-based platforms such as NVIDIA HGX and multi-GPU server architectures.
- Strong knowledge of data centre networking technologies, including InfiniBand, RoCE, Ethernet fabrics, and high-performance network design.
- Experience designing scalable, non-blocking network architectures for AI and HPC environments.
- Understanding of data centre facilities, including power distribution, cooling systems, liquid cooling technologies, rack design, and high-density deployment considerations.
- Proven experience in hardware sizing, component selection, performance testing, and infrastructure capacity planning.
- Ability to design and deliver end-to-end hardware architecture solutions and BOM recommendations.
- Strong analytical, problem-solving, vendor management, stakeholder management, and technical documentation skills.
- Experience working across cross-functional teams, including engineering, operations, facilities, and external vendors.
What's Next
If you are ready to take the next step in your career journey within a supportive global team environment-where your skills are valued and developed-this opportunity awaits your application. Click APPLY or email your resume to *************
Desired Skills and Experience
#GPU #NVIDIA #infrastructurecluster #NVLink #NVSwitch #Infradesign #architect #design #GPUCluster #DataCentre #InfiniBand #RoCE #architecture
Do note that we will only be in touch if your application is shortlisted.
Robert Walters (Singapore) Pte Ltd
ROC No.: 199706961E | EA Licence No.: 03C5451
EA Registration No.: R1324990 Neha Singh