We are looking for a hands-on System Engineer – Infrastructure to support and optimise large-scale GPU and CPU infrastructure powering AI workloads, large model training and high-performance computing environments.
You will be responsible for deploying, maintaining and troubleshooting GPU/CPU servers while ensuring infrastructure reliability, performance and operational readiness.
Key Responsibilities
- Deploy, configure and maintain GPU and CPU servers across large-scale compute clusters.
- Optimise BIOS, firmware and operating system configurations for AI and HPC workloads.
- Perform server health monitoring, hardware diagnostics, firmware upgrades and lifecycle management.
- Support cluster deployments across multiple racks and coordinate with hardware vendors and system integrators.
- Manage server provisioning, OS deployment, GPU/NIC driver installation and system hardening.
- Conduct system validation, burn-in testing and workload benchmarking.
- Monitor system health, investigate failures and perform root cause analysis.
- Work closely with networking, storage and DevOps teams to ensure end-to-end infrastructure performance.
- Develop scripts and automation to improve infrastructure deployment, monitoring and remediation.
- Maintain technical documentation including system configurations, rack layouts, cabling and operational procedures.
- Provide L2/L3 support and participate in on-call activities.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering or a related technical field.
- At least 3 years of hands-on experience managing server infrastructure in HPC, AI, GPU cluster or data centre environments.
- Strong experience with Linux systems, system tuning and performance optimisation.
- Hands-on experience with GPU/CPU servers and bare-metal infrastructure.
- Knowledge of server provisioning technologies such as IPMI, PXE, Redfish or BMC.
- Familiarity with monitoring tools such as Prometheus and Grafana.
- Basic knowledge of Kubernetes or containerised environments.
- Experience with server hardware, GPU platforms and infrastructure troubleshooting.
Pay: RM8,000.00 - RM15,000.00 per month
Application Question(s):
- How many years of experience do you have in server or infrastructure engineering?
- Do you have hands-on experience with GPU/CPU servers, HPC or AI cluster environments?
- Do you have experience working with Linux and bare-metal servers?
Work Location: In person