Job Role: System Engineer - Infrastructure (AI & HPC Systems)
Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service.
Location: Kulai, Johor, Malaysia
Job Type: Full Time – On Site
Experience: 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
Applicant: Local Malaysian citizens only
JOB DESCRIPTION
- Deploy and manage GPU and CPU servers in GPU clusters, ensuring optimized BIOS, firmware, and OS configurations for high-performance AI workloads.
- Maintain physical and virtual infrastructure, including server health monitoring, firmware upgrades, and hardware-level diagnostics.
- Support large-scale cluster deployments across multiple racks; coordinate with hardware vendors and integrators on system delivery, RMA, and maintenance.
- Implement system provisioning processes including automation for OS flashing, GPU/NIC driver installs, and baseline system hardening.
- Conduct system-level validation, burn-in tests, and workload benchmarking for cluster readiness.
- Monitor system health, track failure patterns, and drive corrective actions including hardware replacements and root cause analysis.
- Work closely with networking, storage, and DevOps teams to ensure end-to-end performance and service quality.
- Write scripts and automation tools to streamline infrastructure setup, monitoring, alerting, and remediation.
- Maintain technical documentation for rack layouts, cabling diagrams, system configs, and operational procedures.
- Participate in on-call rotations and provide L2/L3 support for system-related incidents.
JOB REQUIREMENTS
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
- 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
- Strong technical expertise in configuring, deploying, and maintaining bare metal servers, GPU nodes, and CPU-based systems.
- Deep understanding of Linux system internals, kernel tuning, and performance optimization specific to compute-heavy workloads.
- Familiarity with server provisioning and orchestration tools, such as IPMI, PXE boot, Redfish, or BMC tooling.
- Basic understanding in monitoring (e.g., Prometheus, Grafana) and centralized logging tools.
- Basic familiarity with Kubernetes or container-based environments.
Desired Skills
- Strong interpersonal skills, with a proven ability to develop professional relationships across business and technical teams.
- Hands-on experience with server vendors and GPU platforms.
- Hands-on experience with storage systems (e.g., NVMe, SAN, NAS), networking concepts, and protocols (e.g., TCP/IP, RDMA) will be advantageous.
- Knowledgeable in operating ticketing system and trouble shooting process in CPU/GPU cluster.
- Excellent documentation skills to effectively articulate technical designs, issues, procedures, and assessments.
- Strong understanding of GPU architectures, virtualization technologies, and bare metal provisioning is a plus.
- Strong analytical and troubleshooting skills with a customer-centric approach.
Benefits:
- Opportunities for promotion
- Professional development
Application Question(s):
- Are you a local Malaysian citizen?
Experience:
- High Performance Computing: 3 years (Required)
- GPU Cluster: 3 years (Required)
- AI Cluster: 3 years (Required)
- CPU Cluster: 3 years (Required)
Work Location: In person