jobs in Getrosoft

Getrosoft Hiring! Full Time AI Data Center Operations (Core Linux - Containers) in Federal Territory - Ricebowl

AI Data Center Operations (Core Linux - Containers)

Getrosoft

KL City, Federal Territory

Share
Save

Working Location

  • Kuala Lumpur Federal Territory Malaysia

Job Description

Responsibilities

We’re Hiring: AI Data Center Operations (Core Linux & Containers)


We are seeking an experienced Senior SME/Lead – AI Data Center Operations with strong hands-on experience in AI Data Center Operations, Linux, Containers, and NVIDIA AI infrastructure.


Location: Kuala Lumpur, Malaysia

Work Model: Hybrid

Project Duration: 6–8 Months

️ Relocation: Candidates must be willing to relocate to Malaysia


Qualifications

  • 7+ years of hands-on AI Data Center Operations & Maintenance experience
  • Strong experience with Linux & Containers in AI Data Center environments
  • Hands-on experience supporting large-scale NVIDIA AI clusters
  • Direct experience with NVIDIA GB200 / Blackwell deployment or production operations
  • Experience handling GPU / Node failures and troubleshooting NVLink / NVSwitch
  • Experience managing cooling & power incidents
  • Experience handling capacity & performance issues
  • Experience with firmware / driver lifecycle management
  • Experience with security remediation
  • Experience with resilience & Disaster Recovery (DR)
  • Strong understanding of escalation & governance processes
  • Practical experience developing or working with operational runbooks
  • Experience with 20 MW AI Data Center environments is preferred
  • Understanding of manpower requirements for 24×7 operations coverage


️ Key Responsibilities

  • Lead day-to-day AI Data Center Operations & Maintenance activities
  • Monitor and troubleshoot NVIDIA AI infrastructure, GPU nodes, Linux, and container environments
  • Handle and resolve critical incidents involving GPU, NVLink/NVSwitch, cooling, power, and system failures
  • Drive incident management, root cause analysis (RCA), and problem resolution
  • Manage firmware, drivers, patches, and security remediation across the infrastructure
  • Develop, maintain, and continuously improve operational runbooks and SOPs
  • Support capacity planning, performance optimization, resilience, and Disaster Recovery (DR)
  • Coordinate escalations, technical governance, and cross-functional teams during critical incidents
  • Support operational planning for 24×7 AI Data Center coverage
  • Ensure reliable and efficient Hardware O&M across the AI Data Center environment

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More