- Kuala Lumpur Federal Territory Malaysia
工作地点
职位描述
岗位职责
Location: Malaysia
Travel Requirement: Willingness to travel internationally as required
Language Requirements: English, Mandarin and other languages are an advantage
Core AI Engineering Competencies
AI Compute Cluster Deployment & Operations
• Hands-on experience in deploying, configuring, and operating GPU clusters ranging from hundreds to thousands of GPUs.
• Familiarity with NVIDIA GPUs and mainstream AI server hardware.
• Experience in GPU server installation, configuration, troubleshooting, and system optimisation.
High-Speed Networking & Storage
• Familiarity with 100G / 200G / 400G high-speed networking architectures, including InfiniBand (IB) and RoCE.
• Basic understanding of NVMe-based storage systems.
• Ability to perform basic network connectivity testing, troubleshooting, and fault isolation.
Linux & Cloud-Native Environment
• Strong proficiency in common Linux commands and system administration.
• Basic hands-on experience with Docker and Kubernetes.
• Ability to work with algorithm and AI engineering teams to build, configure, and maintain AI training environments.
Data Centre Infrastructure
• Understanding of data centre / IDC infrastructure, including:
- Rack planning and deployment
- UPS systems
- Power distribution
- Cooling infrastructure
- Liquid cooling and air cooling systems
Key Responsibilities
• Responsible for the installation, commissioning, deployment, go-live, and daily operations of AI compute centres, GPU servers, and supporting infrastructure.
• Participate in large-scale GPU cluster deployment, including network configuration, storage mounting, system environment deployment, and infrastructure troubleshooting.
• Assist in resolving basic hardware and software issues encountered during AI training workloads.
• Work closely with senior architects and technical leads on equipment selection, technical solution implementation, and on-site project deployment.
• Establish and maintain cluster monitoring, alerting, and operational procedures.
• Maintain operational logs and participate in incident reviews, root-cause analysis, and process optimisation.
• Coordinate with domestic and international suppliers, data centre operators, vendors, and engineering teams to ensure on-time project delivery and stable infrastructure operations.
Experience Requirements
• 3–5 years of experience in data centre operations, server hardware, networking, cloud computing, or related infrastructure operations.
• Experience in AI compute centres, GPU clusters, intelligent computing projects, or large-scale data centre projects is highly preferred.
• Experience with NVIDIA GPU infrastructure is an advantage.
Technical Background
• Degree or diploma in Computer Science, Information Technology, Computer Networking, Telecommunications, Electronic Engineering, or a related field.
• Strong foundation in Linux system administration and network troubleshooting.
• Good understanding of server hardware, networking, storage, and data centre infrastructure.
• Strong hands-on troubleshooting and problem-solving capabilities.
Execution & Responsibility
• Strong hands-on capability with the ability to independently undertake on-site implementation, commissioning, troubleshooting, and incident handling.
• Strong sense of ownership, responsibility, and execution.
• Ability to communicate and collaborate effectively across technical, engineering, vendor, and project teams.
Resilience & Adaptability
• Willingness to work in high-intensity project environments and respond to urgent infrastructure issues when required.
• Willingness to undertake frequent international travel for data centre and AI infrastructure projects.
• Strong interest in AI compute infrastructure and a commitment to continuously developing technical knowledge.
International Exposure & Career Opportunity
• Direct involvement in international AI compute centre and data centre projects.
• Opportunity to participate in the deployment and operation of large-scale GPU clusters, including projects ranging from hundreds to potentially 10,000+ GPUs.
• Gain hands-on experience across GPU infrastructure, high-speed networking, AI compute, data centre infrastructure, cloud platforms, and large-scale AI deployments.
• Build international project experience across Southeast Asia, the Middle East, and other global markets.
重要安全守则
申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。