jobs in Verinon

全职 System Engineer - Infrastructure (AI - HPC Systems) 工作, 薪水, Verinon Johor 公司招聘中 - Ricebowl

System Engineer - Infrastructure (AI - HPC Systems)

Verinon

Undisclosed
分享
保存

工作地点

  • Kulai Johor Malaysia

职位描述

岗位职责

Job Role: System Engineer - Infrastructure (AI & HPC Systems)

Employer: Company that provides specialized, massive-scale GPU-based accelerated computing and AI infrastructure-as-a-service.

Location: Kulai, Johor, Malaysia

Job Type: Full Time – On Site

Experience: 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.

Applicant: Local Malaysian citizens only

JOB DESCRIPTION

  • Deploy and manage GPU and CPU servers in GPU clusters, ensuring optimized BIOS, firmware, and OS configurations for high-performance AI workloads.
  • Maintain physical and virtual infrastructure, including server health monitoring, firmware upgrades, and hardware-level diagnostics.
  • Support large-scale cluster deployments across multiple racks; coordinate with hardware vendors and integrators on system delivery, RMA, and maintenance.
  • Implement system provisioning processes including automation for OS flashing, GPU/NIC driver installs, and baseline system hardening.
  • Conduct system-level validation, burn-in tests, and workload benchmarking for cluster readiness.
  • Monitor system health, track failure patterns, and drive corrective actions including hardware replacements and root cause analysis.
  • Work closely with networking, storage, and DevOps teams to ensure end-to-end performance and service quality.
  • Write scripts and automation tools to streamline infrastructure setup, monitoring, alerting, and remediation.
  • Maintain technical documentation for rack layouts, cabling diagrams, system configs, and operational procedures.
  • Participate in on-call rotations and provide L2/L3 support for system-related incidents.

JOB REQUIREMENTS

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.
  • 3+ years of hands-on experience managing large-scale GPU/CPU server infrastructure in high-performance computing (HPC), AI clusters, or data center environments.
  • Strong technical expertise in configuring, deploying, and maintaining bare metal servers, GPU nodes, and CPU-based systems.
  • Deep understanding of Linux system internals, kernel tuning, and performance optimization specific to compute-heavy workloads.
  • Familiarity with server provisioning and orchestration tools, such as IPMI, PXE boot, Redfish, or BMC tooling.
  • Basic understanding in monitoring (e.g., Prometheus, Grafana) and centralized logging tools.
  • Basic familiarity with Kubernetes or container-based environments.

Desired Skills

  • Strong interpersonal skills, with a proven ability to develop professional relationships across business and technical teams.
  • Hands-on experience with server vendors and GPU platforms.
  • Hands-on experience with storage systems (e.g., NVMe, SAN, NAS), networking concepts, and protocols (e.g., TCP/IP, RDMA) will be advantageous.
  • Knowledgeable in operating ticketing system and trouble shooting process in CPU/GPU cluster.
  • Excellent documentation skills to effectively articulate technical designs, issues, procedures, and assessments.
  • Strong understanding of GPU architectures, virtualization technologies, and bare metal provisioning is a plus.
  • Strong analytical and troubleshooting skills with a customer-centric approach.

Benefits:

  • Opportunities for promotion
  • Professional development

Application Question(s):

  • Are you a local Malaysian citizen?

Experience:

  • High Performance Computing: 3 years (Required)
  • GPU Cluster: 3 years (Required)
  • AI Cluster: 3 years (Required)
  • CPU Cluster: 3 years (Required)

Work Location: In person

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多