jobs in INF Tech

全职 Datacenter IT Infrastructure Operations Expert 工作, 薪水, INF Tech 公司招聘中 - Ricebowl

Datacenter IT Infrastructure Operations Expert

INF Tech

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

Responsibilities

1. Own day-to-day availability, health inspection and hardware fault closure for GPU servers, CPU servers, network equipment and storage systems in the data hall.

2. Own rack-and-stack, structured cabling and optics management, together with in-rack rPDU and rack-level power distribution and capacity management.

3. Own liquid-cooling CDU and secondary-loop operations: runtime monitoring, coolant and filter maintenance, leak detection and emergency response, and quick-disconnect work procedures.

4. Own firmware and driver baseline management and the full vendor RMA cycle — fault evidence capture, ticket submission, part replacement and repair verification.

5. Own NOC / ECC monitoring operations and incident management: 24×7 coverage, alert tiering and escalation, and command of major incident response and post-incident review.

6. Run operations across multiple sites or projects in parallel, allocating people, spares and budget; manage the frontline team and vendors, and build SOPs, emergency plans and spare-parts strategy.


Requirements

1. Bachelor's degree or above in Computer Science, Electronics, Automation or a related field; 8+ years in datacenter IT equipment operations, including 3+ years in an operations management or team-lead role.

2. Strong hands-on data hall experience: rack-and-stack, cabling, part replacement and on-site fault handling for servers, network and storage equipment.

3. Solid knowledge of server hardware and out-of-band management (BMC / IPMI / Redfish), with command of firmware upgrade and fleet-wide operations methods.

4. Hands-on GPU server hardware operations: fault diagnosis and replacement of GPU boards, NVLink / NVSwitch modules, and power and cooling components.

5. Hands-on field operations for network and storage equipment: switch commissioning and configuration backup, optics and high-speed link (InfiniBand / RoCE) fault isolation.

6. Hands-on liquid-cooling CDU and secondary-loop operations, including coolant and water-quality management, leak detection and emergency response procedures.

7. NOC / ECC build-out or management experience; familiarity with ITIL or an equivalent framework and practical incident, problem and change management.

8. Proven ability to run multiple projects in parallel with strong incident command and cross-team coordination; resilient under pressure and able to support 24×7 escalation.


Preferred

• Field operations and hardware fault-handling experience with large-scale GPU clusters (1,000+ GPUs).

• End-to-end experience with liquid-cooled racks (direct-to-chip / DLC) from deployment and commissioning through to volume operations.

• Linux fundamentals and scripting to build automated inspection or fleet diagnostic tooling.

• ITIL, PMP or a vendor hardware-maintenance certification.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多