jobs in Fujitsu

全职 Systems Engineer 工作, 薪水, Fujitsu 公司招聘中 - Ricebowl

Undisclosed

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

Job Location: Singapore

Location Flexibility: Primary Location Only

Req Id: 11076

Posting Start Date: 8/7/26

Responsibilites

  • Manage the day-to-day operations of the HPE Cray EX supercomputing environment, ensuring high availability, stability, performance, and reliability of HPC services.
  • Administer and maintain HPE Cluster Manager (HPCM) for cluster provisioning, monitoring, health management, software deployment, and lifecycle management.
  • Manage AMD-based HPE Cray EX compute infrastructure delivering up to 10 PFLOPS of computational performance, ensuring optimal resource utilization and system efficiency.
  • Administer and optimize HPE ClusterStor Lustre parallel file system with over 10 PB of storage capacity, ensuring high-performance I/O, data integrity, and storage availability.
  • Manage IBM Storage Scale (formerly GPFS) parallel file system with over 15 PB of storage capacity, including performance tuning, capacity planning, and filesystem maintenance.
  • Configure, administer, and maintain the PBS Professional workload manager, including queue configuration, scheduling policies, fair-share management, resource allocation, and job troubleshooting.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Provide technical support to a diverse community of researchers, scientists, engineers, and academic users by troubleshooting application, storage, scheduler, and system-related issues.
  • Assist users in optimizing HPC applications through performance analysis, job scheduling best practices, parallel computing techniques, and efficient resource utilization.
  • Conduct user onboarding sessions, technical workshops, and training programs on HPC environment usage, job submission, parallel file systems, and cluster best practices.
  • Perform software installation, upgrades, patch management, and validation for HPC operating systems, middleware, compilers, MPI libraries, and scientific applications.
  • Develop and maintain automation scripts using Shell, Python, or similar scripting languages to streamline system administration, monitoring, reporting, and operational tasks.
  • Maintain comprehensive operational documentation, standard operating procedures (SOPs), architecture diagrams, and technical knowledge base articles.
  • Participate in incident response, planned maintenance activities, disaster recovery exercises, and root cause analysis to ensure continuous improvement of HPC infrastructure.
  • Ensure adherence to security policies, operational standards, and best practices while maintaining a secure and highly available HPC environment.
  • Continuously evaluate emerging HPC technologies and recommend improvements to enhance system performance, scalability, reliability, and operational efficiency.

Requirements

  • 5–7 years of hands-on experience administering High Performance Computing (HPC) environments in enterprise, research, or academic organizations.
  • Familiarity with configuration management and automation tools such as Ansible, xCAT, Bright Cluster Manager, or Infrastructure-as-Code solutions is an advantage.
  • Understanding of GPU computing technologies (NVIDIA CUDA, AMD ROCm) and accelerator-based HPC environments is an added advantage.

Relocation Supported: No

Visa Sponsorship Approved: No

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多