jobs in INF Tech

INF Tech Hiring! Full Time Datacenter IT Infrastructure Operations Expert in - Ricebowl

Datacenter IT Infrastructure Operations Expert

INF Tech

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

Responsibilities

1. Own day-to-day availability, health inspection and hardware fault closure for GPU servers, CPU servers, network equipment and storage systems in the data hall.

2. Own rack-and-stack, structured cabling and optics management, together with in-rack rPDU and rack-level power distribution and capacity management.

3. Own liquid-cooling CDU and secondary-loop operations: runtime monitoring, coolant and filter maintenance, leak detection and emergency response, and quick-disconnect work procedures.

4. Own firmware and driver baseline management and the full vendor RMA cycle — fault evidence capture, ticket submission, part replacement and repair verification.

5. Own NOC / ECC monitoring operations and incident management: 24×7 coverage, alert tiering and escalation, and command of major incident response and post-incident review.

6. Run operations across multiple sites or projects in parallel, allocating people, spares and budget; manage the frontline team and vendors, and build SOPs, emergency plans and spare-parts strategy.


Requirements

1. Bachelor's degree or above in Computer Science, Electronics, Automation or a related field; 8+ years in datacenter IT equipment operations, including 3+ years in an operations management or team-lead role.

2. Strong hands-on data hall experience: rack-and-stack, cabling, part replacement and on-site fault handling for servers, network and storage equipment.

3. Solid knowledge of server hardware and out-of-band management (BMC / IPMI / Redfish), with command of firmware upgrade and fleet-wide operations methods.

4. Hands-on GPU server hardware operations: fault diagnosis and replacement of GPU boards, NVLink / NVSwitch modules, and power and cooling components.

5. Hands-on field operations for network and storage equipment: switch commissioning and configuration backup, optics and high-speed link (InfiniBand / RoCE) fault isolation.

6. Hands-on liquid-cooling CDU and secondary-loop operations, including coolant and water-quality management, leak detection and emergency response procedures.

7. NOC / ECC build-out or management experience; familiarity with ITIL or an equivalent framework and practical incident, problem and change management.

8. Proven ability to run multiple projects in parallel with strong incident command and cross-team coordination; resilient under pressure and able to support 24×7 escalation.


Preferred

• Field operations and hardware fault-handling experience with large-scale GPU clusters (1,000+ GPUs).

• End-to-end experience with liquid-cooled racks (direct-to-chip / DLC) from deployment and commissioning through to volume operations.

• Linux fundamentals and scripting to build automated inspection or fleet diagnostic tooling.

• ITIL, PMP or a vendor hardware-maintenance certification.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More