jobs in Oxydata Software

全职 Senior Data Centre Operations Engineer 工作, 薪水, Oxydata Software Johor 公司招聘中 - Ricebowl

Senior Data Centre Operations Engineer

Undisclosed
分享
保存

工作地点

  • Senai Johor Malaysia

职位描述

岗位职责

Senior Data Centre Operations Engineer

Location: Senai, Johor, Malaysia
Work Mode: Onsite

Employment type: Permanent

Our client is a leading specialist in the repair and maintenance of high-end AI computing infrastructure, with a state-of-the-art facility located in Johor Bahru, Malaysia. They are dedicated to providing mission-critical support and have established a reputation for precision and reliability in the Southeast Asian market. With a strong focus on transparency and accountability, they ensure that every repair process is documented and approved by clients, maintaining a high standard of service excellence.

We are seeking an experienced Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.

Responsibilities

  • Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
  • Monitor server health and review IPMI, BMC, and operating system logs.
  • Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
  • Manage RAID configurations and monitor SSD and NVMe health.
  • Troubleshoot GPU servers and replace faulty hardware components.
  • Support GPU cluster performance, network topology, and system stability.
  • Collaborate with networking, storage, and virtualisation teams to resolve technical issues.
  • Automate routine tasks such as firmware upgrades, inspections, and hardware alert management.
  • Prepare technical guides, troubleshooting documentation, and standard operating procedures.
  • Coordinate with hardware vendors and manage RMAs and spare parts.
  • Monitor rack power, temperature, and air-cooling conditions.

Requirements

Must-have:

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
  • Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
  • At least 5-7 years of experience in server operations.
  • Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
  • Strong Linux administration and troubleshooting knowledge.
  • Good hardware troubleshooting and problem-solving skills.
  • Ability to work effectively with internal teams and external vendors.
  • Good technical communication skills in English.

Nice-to-have:

  • Experience operating large-scale GPU clusters.
  • Experience with Shell or Ansible automation.
  • Familiarity with InfiniBand, RoCE, RDMA networking and optical modules.
  • Exposure to air-cooled server environments.
  • Experience with liquid-cooled servers.
  • Knowledge of Ceph, KVM, VMware or hyper-converged infrastructure.
  • Relevant certifications such as RHCE, RHCA, CompTIA Server+ or server vendor certifications.

Education:

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.

Why Join Us

  • Be part of a dynamic team at the forefront of data centre technology, supporting mission-critical AI infrastructure for leading enterprises.
  • Opportunity to work with advanced hardware, collaborate with skilled professionals, and contribute to the reliability of high-performance computing environments across the region.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多