We are seeking a highly experienced and results-driven professional to lead the operations and management of our AI GPU Data Centre. The successful candidate will be responsible for ensuring the reliability, performance, scalability, and security of mission-critical infrastructure supporting AI, machine learning, and high-performance computing (HPC) workloads.
Key Responsibilities
- Lead and manage the overall operations of the AI GPU Data Centre, ensuring high availability, reliability, and security.
- Oversee GPU clusters, servers, storage, networking, power, and cooling infrastructure.
- Manage capacity planning, infrastructure expansion, upgrades, and performance optimization.
- Lead incident management, troubleshooting, disaster recovery, and business continuity initiatives.
- Ensure compliance with operational standards, security requirements, and SLAs.
- Manage budgets, vendors, contracts, and procurement activities.
- Lead, mentor, and develop the Data Centre Operations team.
- Collaborate with AI, cloud, engineering, and infrastructure teams to support business growth.
Key Requirements
- Bachelor's Degree in IT, Computer Science, Engineering, or related field.
- Minimum 10 years of experience in Data Centre Operations, Critical Facilities, Cloud Infrastructure, AI, or HPC environments.
- At least 5 years of leadership or managerial experience.
- Strong knowledge of AI GPU infrastructure, servers, storage, networking, virtualization, and data centre facilities.
- Experience managing mission-critical or hyperscale data centre environments.
- Familiarity with NVIDIA GPU platforms, AI workloads, cloud infrastructure, and HPC environments is an advantage.
- Strong leadership, project management, and stakeholder management skills.
Pay: RM20,000.00 - RM25,000.00 per month
Benefits:
Work Location: In person