jobs in Neuron Solutions Sdn Bhd

全职 System Engineer - Infrastructure (AI - HPC Systems) 工作, 薪水, Neuron Solutions Johor 公司招聘中 - Ricebowl

System Engineer - Infrastructure (AI - HPC Systems)

Neuron Solutions Sdn Bhd

Undisclosed
分享
保存

工作地点

  • Kulai Johor Malaysia

职位描述

岗位职责

Job Role: System Engineer - Infrastructure (AI & HPC Systems)

Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service

Location: Kulai, Johor, Malaysia

Job Type: Full Time – On Site

Experience: 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.


JOB DESCRIPTION

• Deploy and manage GPU and CPU servers in GPU clusters, ensuring optimized BIOS, firmware, and OS configurations for high-performance AI workloads.

• Maintain physical and virtual infrastructure, including server health monitoring, firmware upgrades, and hardware-level diagnostics.

• Support large-scale cluster deployments across multiple racks; coordinate with hardware vendors and integrators on system delivery, RMA, and maintenance.

• Implement system provisioning processes including automation for OS flashing, GPU/NIC driver installs, and baseline system hardening.

• Conduct system-level validation, burn-in tests, and workload benchmarking for cluster readiness.

• Monitor system health, track failure patterns, and drive corrective actions including hardware replacements and root cause analysis.

• Work closely with networking, storage, and DevOps teams to ensure end-to-end performance and service quality.

• Write scripts and automation tools to streamline infrastructure setup, monitoring, alerting, and remediation.

• Maintain technical documentation for rack layouts, cabling diagrams, system configs, and operational procedures.

• Participate in on-call rotations and provide L2/L3 support for system-related incidents.


JOB REQUIREMENTS

• Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.

• 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.

• Strong technical expertise in configuring, deploying, and maintaining bare metal servers, GPU nodes, and CPU-based systems.

• Deep understanding of Linux system internals, kernel tuning, and performance optimization specific to compute-heavy workloads.

• Familiarity with server provisioning and orchestration tools, such as IPMI, PXE boot, Redfish, or BMC tooling.

• Basic understanding in monitoring (e.g., Prometheus, Grafana) and centralized logging tools.

• Basic familiarity with Kubernetes or container-based environments.


Desired Skills

• Strong interpersonal skills, with a proven ability to develop professional relationships across business and technical teams.

• Hands-on experience with server vendors and GPU platforms.

• Hands-on experience with storage systems (e.g., NVMe, SAN, NAS), networking concepts, and protocols (e.g., TCP/IP, RDMA) will be advantageous.

• Knowledgeable in operating ticketing system and trouble shooting process in CPU/GPU cluster.

• Excellent documentation skills to effectively articulate technical designs, issues, procedures, and assessments.

• Strong understanding of GPU architectures, virtualization technologies, and bare metal provisioning is a plus.

• Strong analytical and troubleshooting skills with a customer-centric approach.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多