- Singapore
工作地点
职位描述
岗位职责
1. Own day-to-day availability, health inspection and hardware fault closure for GPU servers, CPU servers, network equipment and storage systems in the data hall.
2. Own rack-and-stack, structured cabling and optics management, together with in-rack rPDU and rack-level power distribution and capacity management.
3. Own liquid-cooling CDU and secondary-loop operations: runtime monitoring, coolant and filter maintenance, leak detection and emergency response, and quick-disconnect work procedures.
4. Own firmware and driver baseline management and the full vendor RMA cycle — fault evidence capture, ticket submission, part replacement and repair verification.
5. Own NOC / ECC monitoring operations and incident management: 24×7 coverage, alert tiering and escalation, and command of major incident response and post-incident review.
6. Run operations across multiple sites or projects in parallel, allocating people, spares and budget; manage the frontline team and vendors, and build SOPs, emergency plans and spare-parts strategy.
1. Bachelor's degree or above in Computer Science, Electronics, Automation or a related field; 8+ years in datacenter IT equipment operations, including 3+ years in an operations management or team-lead role.
2. Strong hands-on data hall experience: rack-and-stack, cabling, part replacement and on-site fault handling for servers, network and storage equipment.
3. Solid knowledge of server hardware and out-of-band management (BMC / IPMI / Redfish), with command of firmware upgrade and fleet-wide operations methods.
4. Hands-on GPU server hardware operations: fault diagnosis and replacement of GPU boards, NVLink / NVSwitch modules, and power and cooling components.
5. Hands-on field operations for network and storage equipment: switch commissioning and configuration backup, optics and high-speed link (InfiniBand / RoCE) fault isolation.
6. Hands-on liquid-cooling CDU and secondary-loop operations, including coolant and water-quality management, leak detection and emergency response procedures.
7. NOC / ECC build-out or management experience; familiarity with ITIL or an equivalent framework and practical incident, problem and change management.
8. Proven ability to run multiple projects in parallel with strong incident command and cross-team coordination; resilient under pressure and able to support 24×7 escalation.
• Field operations and hardware fault-handling experience with large-scale GPU clusters (1,000+ GPUs).
• End-to-end experience with liquid-cooled racks (direct-to-chip / DLC) from deployment and commissioning through to volume operations.
• Linux fundamentals and scripting to build automated inspection or fleet diagnostic tooling.
• ITIL, PMP or a vendor hardware-maintenance certification.
重要安全守则
申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。