- Kulai Johor Malaysia
工作地点
职位描述
岗位职责
Job Role: System Engineer - Infrastructure (AI & HPC Systems)
Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service
Location: Kulai, Johor, Malaysia
Job Type: Full Time – On Site
Experience: 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
JOB DESCRIPTION
• Deploy and manage GPU and CPU servers in GPU clusters, ensuring optimized BIOS, firmware, and OS configurations for high-performance AI workloads.
• Maintain physical and virtual infrastructure, including server health monitoring, firmware upgrades, and hardware-level diagnostics.
• Support large-scale cluster deployments across multiple racks; coordinate with hardware vendors and integrators on system delivery, RMA, and maintenance.
• Implement system provisioning processes including automation for OS flashing, GPU/NIC driver installs, and baseline system hardening.
• Conduct system-level validation, burn-in tests, and workload benchmarking for cluster readiness.
• Monitor system health, track failure patterns, and drive corrective actions including hardware replacements and root cause analysis.
• Work closely with networking, storage, and DevOps teams to ensure end-to-end performance and service quality.
• Write scripts and automation tools to streamline infrastructure setup, monitoring, alerting, and remediation.
• Maintain technical documentation for rack layouts, cabling diagrams, system configs, and operational procedures.
• Participate in on-call rotations and provide L2/L3 support for system-related incidents.
JOB REQUIREMENTS
• Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
• 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
• Strong technical expertise in configuring, deploying, and maintaining bare metal servers, GPU nodes, and CPU-based systems.
• Deep understanding of Linux system internals, kernel tuning, and performance optimization specific to compute-heavy workloads.
• Familiarity with server provisioning and orchestration tools, such as IPMI, PXE boot, Redfish, or BMC tooling.
• Basic understanding in monitoring (e.g., Prometheus, Grafana) and centralized logging tools.
• Basic familiarity with Kubernetes or container-based environments.
Desired Skills
• Strong interpersonal skills, with a proven ability to develop professional relationships across business and technical teams.
• Hands-on experience with server vendors and GPU platforms.
• Hands-on experience with storage systems (e.g., NVMe, SAN, NAS), networking concepts, and protocols (e.g., TCP/IP, RDMA) will be advantageous.
• Knowledgeable in operating ticketing system and trouble shooting process in CPU/GPU cluster.
• Excellent documentation skills to effectively articulate technical designs, issues, procedures, and assessments.
• Strong understanding of GPU architectures, virtualization technologies, and bare metal provisioning is a plus.
• Strong analytical and troubleshooting skills with a customer-centric approach.
重要安全守则
申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。