jobs in Fujitsu

Full Time Systems Engineer Jobs, in Fujitsu - Ricebowl

Undisclosed

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

Job Location: Singapore

Location Flexibility: Primary Location Only

Req Id: 11076

Posting Start Date: 8/7/26

Responsibilites

  • Manage the day-to-day operations of the HPE Cray EX supercomputing environment, ensuring high availability, stability, performance, and reliability of HPC services.
  • Administer and maintain HPE Cluster Manager (HPCM) for cluster provisioning, monitoring, health management, software deployment, and lifecycle management.
  • Manage AMD-based HPE Cray EX compute infrastructure delivering up to 10 PFLOPS of computational performance, ensuring optimal resource utilization and system efficiency.
  • Administer and optimize HPE ClusterStor Lustre parallel file system with over 10 PB of storage capacity, ensuring high-performance I/O, data integrity, and storage availability.
  • Manage IBM Storage Scale (formerly GPFS) parallel file system with over 15 PB of storage capacity, including performance tuning, capacity planning, and filesystem maintenance.
  • Configure, administer, and maintain the PBS Professional workload manager, including queue configuration, scheduling policies, fair-share management, resource allocation, and job troubleshooting.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Provide technical support to a diverse community of researchers, scientists, engineers, and academic users by troubleshooting application, storage, scheduler, and system-related issues.
  • Assist users in optimizing HPC applications through performance analysis, job scheduling best practices, parallel computing techniques, and efficient resource utilization.
  • Conduct user onboarding sessions, technical workshops, and training programs on HPC environment usage, job submission, parallel file systems, and cluster best practices.
  • Perform software installation, upgrades, patch management, and validation for HPC operating systems, middleware, compilers, MPI libraries, and scientific applications.
  • Develop and maintain automation scripts using Shell, Python, or similar scripting languages to streamline system administration, monitoring, reporting, and operational tasks.
  • Maintain comprehensive operational documentation, standard operating procedures (SOPs), architecture diagrams, and technical knowledge base articles.
  • Participate in incident response, planned maintenance activities, disaster recovery exercises, and root cause analysis to ensure continuous improvement of HPC infrastructure.
  • Ensure adherence to security policies, operational standards, and best practices while maintaining a secure and highly available HPC environment.
  • Continuously evaluate emerging HPC technologies and recommend improvements to enhance system performance, scalability, reliability, and operational efficiency.

Requirements

  • 5–7 years of hands-on experience administering High Performance Computing (HPC) environments in enterprise, research, or academic organizations.
  • Familiarity with configuration management and automation tools such as Ansible, xCAT, Bright Cluster Manager, or Infrastructure-as-Code solutions is an advantage.
  • Understanding of GPU computing technologies (NVIDIA CUDA, AMD ROCm) and accelerator-based HPC environments is an added advantage.

Relocation Supported: No

Visa Sponsorship Approved: No

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More