- Kulai Johor Malaysia
Working Location
Job Description
Responsibilities
Job Role: System Engineer - Infrastructure (AI & HPC Systems)
Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service
Location: Kulai, Johor, Malaysia
Job Type: Full Time – On Site
Experience: 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
JOB DESCRIPTION
• Deploy and manage GPU and CPU servers in GPU clusters, ensuring optimized BIOS, firmware, and OS configurations for high-performance AI workloads.
• Maintain physical and virtual infrastructure, including server health monitoring, firmware upgrades, and hardware-level diagnostics.
• Support large-scale cluster deployments across multiple racks; coordinate with hardware vendors and integrators on system delivery, RMA, and maintenance.
• Implement system provisioning processes including automation for OS flashing, GPU/NIC driver installs, and baseline system hardening.
• Conduct system-level validation, burn-in tests, and workload benchmarking for cluster readiness.
• Monitor system health, track failure patterns, and drive corrective actions including hardware replacements and root cause analysis.
• Work closely with networking, storage, and DevOps teams to ensure end-to-end performance and service quality.
• Write scripts and automation tools to streamline infrastructure setup, monitoring, alerting, and remediation.
• Maintain technical documentation for rack layouts, cabling diagrams, system configs, and operational procedures.
• Participate in on-call rotations and provide L2/L3 support for system-related incidents.
JOB REQUIREMENTS
• Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
• 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
• Strong technical expertise in configuring, deploying, and maintaining bare metal servers, GPU nodes, and CPU-based systems.
• Deep understanding of Linux system internals, kernel tuning, and performance optimization specific to compute-heavy workloads.
• Familiarity with server provisioning and orchestration tools, such as IPMI, PXE boot, Redfish, or BMC tooling.
• Basic understanding in monitoring (e.g., Prometheus, Grafana) and centralized logging tools.
• Basic familiarity with Kubernetes or container-based environments.
Desired Skills
• Strong interpersonal skills, with a proven ability to develop professional relationships across business and technical teams.
• Hands-on experience with server vendors and GPU platforms.
• Hands-on experience with storage systems (e.g., NVMe, SAN, NAS), networking concepts, and protocols (e.g., TCP/IP, RDMA) will be advantageous.
• Knowledgeable in operating ticketing system and trouble shooting process in CPU/GPU cluster.
• Excellent documentation skills to effectively articulate technical designs, issues, procedures, and assessments.
• Strong understanding of GPU architectures, virtualization technologies, and bare metal provisioning is a plus.
• Strong analytical and troubleshooting skills with a customer-centric approach.
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.