We’re Hiring: AI Data Center Operations (Core Linux & Containers)
We are seeking an experienced Senior SME/Lead – AI Data Center Operations with strong hands-on experience in AI Data Center Operations, Linux, Containers, and NVIDIA AI infrastructure.
Location: Kuala Lumpur, Malaysia
Work Model: Hybrid
Project Duration: 6–8 Months
️ Relocation: Candidates must be willing to relocate to Malaysia
Qualifications
- 7+ years of hands-on AI Data Center Operations & Maintenance experience
- Strong experience with Linux & Containers in AI Data Center environments
- Hands-on experience supporting large-scale NVIDIA AI clusters
- Direct experience with NVIDIA GB200 / Blackwell deployment or production operations
- Experience handling GPU / Node failures and troubleshooting NVLink / NVSwitch
- Experience managing cooling & power incidents
- Experience handling capacity & performance issues
- Experience with firmware / driver lifecycle management
- Experience with security remediation
- Experience with resilience & Disaster Recovery (DR)
- Strong understanding of escalation & governance processes
- Practical experience developing or working with operational runbooks
- Experience with 20 MW AI Data Center environments is preferred
- Understanding of manpower requirements for 24×7 operations coverage
️ Key Responsibilities
- Lead day-to-day AI Data Center Operations & Maintenance activities
- Monitor and troubleshoot NVIDIA AI infrastructure, GPU nodes, Linux, and container environments
- Handle and resolve critical incidents involving GPU, NVLink/NVSwitch, cooling, power, and system failures
- Drive incident management, root cause analysis (RCA), and problem resolution
- Manage firmware, drivers, patches, and security remediation across the infrastructure
- Develop, maintain, and continuously improve operational runbooks and SOPs
- Support capacity planning, performance optimization, resilience, and Disaster Recovery (DR)
- Coordinate escalations, technical governance, and cross-functional teams during critical incidents
- Support operational planning for 24×7 AI Data Center coverage
- Ensure reliable and efficient Hardware O&M across the AI Data Center environment