jobs in Getrosoft

全职 AI Data Center Operations (Core Linux - Containers) 工作, 薪水, Getrosoft Federal Territory 公司招聘中 - Ricebowl

AI Data Center Operations (Core Linux - Containers)

Getrosoft

KL City, Federal Territory

分享
保存

工作地点

  • Kuala Lumpur Federal Territory Malaysia

职位描述

岗位职责

We’re Hiring: AI Data Center Operations (Core Linux & Containers)


We are seeking an experienced Senior SME/Lead – AI Data Center Operations with strong hands-on experience in AI Data Center Operations, Linux, Containers, and NVIDIA AI infrastructure.


Location: Kuala Lumpur, Malaysia

Work Model: Hybrid

Project Duration: 6–8 Months

️ Relocation: Candidates must be willing to relocate to Malaysia


Qualifications

  • 7+ years of hands-on AI Data Center Operations & Maintenance experience
  • Strong experience with Linux & Containers in AI Data Center environments
  • Hands-on experience supporting large-scale NVIDIA AI clusters
  • Direct experience with NVIDIA GB200 / Blackwell deployment or production operations
  • Experience handling GPU / Node failures and troubleshooting NVLink / NVSwitch
  • Experience managing cooling & power incidents
  • Experience handling capacity & performance issues
  • Experience with firmware / driver lifecycle management
  • Experience with security remediation
  • Experience with resilience & Disaster Recovery (DR)
  • Strong understanding of escalation & governance processes
  • Practical experience developing or working with operational runbooks
  • Experience with 20 MW AI Data Center environments is preferred
  • Understanding of manpower requirements for 24×7 operations coverage


️ Key Responsibilities

  • Lead day-to-day AI Data Center Operations & Maintenance activities
  • Monitor and troubleshoot NVIDIA AI infrastructure, GPU nodes, Linux, and container environments
  • Handle and resolve critical incidents involving GPU, NVLink/NVSwitch, cooling, power, and system failures
  • Drive incident management, root cause analysis (RCA), and problem resolution
  • Manage firmware, drivers, patches, and security remediation across the infrastructure
  • Develop, maintain, and continuously improve operational runbooks and SOPs
  • Support capacity planning, performance optimization, resilience, and Disaster Recovery (DR)
  • Coordinate escalations, technical governance, and cross-functional teams during critical incidents
  • Support operational planning for 24×7 AI Data Center coverage
  • Ensure reliable and efficient Hardware O&M across the AI Data Center environment

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多