This role focuses on the operational execution of GPU, network, and compute infrastructure required to support large-scale AI workloads. You will help formalize how AI infrastructure build operations are delivered at sites such as JBP/FYV, including the type of work being performed, the ownership boundaries, and which activities should remain within the build operations function versus being handled by partner teams, vendors, or other Oracle organizations.
You will work closely with internal teams including TPM, DCBT, DCO, DCFE, DCSO, Capacity Planning, Network, Logistics, and Deployment teams, as well as external data center partners and vendors. The role requires a strong understanding of mission-critical data center operations, AI infrastructure, physical deployment workflows, incident response, and vendor coordination.
Responsibilities
Main responsibilities include:
- Define and manage the build operations model for GPU, network, compute, power, cooling, and related infrastructure.
- Create clear processes, ownership, SLAs, reporting, escalation paths, and operational standards.
- Oversee physical deployment from material delivery, rack installation, cabling, TCS connection, power-on, validation, and handover to operations.
- Ensure deployment work is completed safely, consistently, and in line with Oracle standards and project timelines.
- Identify deployment blockers and coordinate resolution with Oracle teams, vendors, logistics providers, and data center partners.
- Mentor and develop DCTs, shift leads, site leads, and fly-in support teams across JAPAC.
- Provide technical operations guidance on rack layout, cabling, power, cooling, serviceability, logistics flow, and build sequencing.
- Manage vendors and partners to ensure they meet Oracle’s safety, quality, timeline, and process requirements.
- Improve material movement, receiving, staging, inventory tracking, store room processes, and spare parts readiness.
- Support steady-state operations, including maintenance, break/fix, incident response, escalation, and service restoration.
- Capture lessons learned from deployments and incidents to improve future build execution.
- Use tools such as Jira, Confluence, dashboards, and reports to track progress, risks, blockers, and operational readiness.
- Drive continuous improvement in deployment quality, execution efficiency, hardware availability, and time to market.
Overall, the role ensures Oracle’s AI infrastructure builds in JAPAC are delivered safely, efficiently, and consistently while building strong regional execution capability.
Key Requirements
- 12+ years of experience in Data Center Operations, Infrastructure Deployment, or AI Infrastructure Support, preferably in hyperscale or cloud environments.
- Strong expertise in GPU, compute, networking, and AI infrastructure deployment and operations.
- Experience managing mission-critical data center build, installation, deployment, and operational execution.
- Proven ability to lead cross-functional teams and collaborate with TPM, Data Center Operations, Network, Capacity Planning, Logistics, and external vendors.
- Strong knowledge of incident management, vendor management, operational readiness, and service deliveryin large-scale data center environments.
- Excellent leadership, stakeholder management, and project execution skills with experience driving complex infrastructure programs.