We are partnering with an industry leader in Singapore to hire experienced HPC Engineers to support the implementation, configuration, maintenance, and operation of High-Performance Computing (HPC) clusters, parallel file systems, and associated infrastructure. This is a hands-on technical role responsible for ensuring the reliability, performance, and availability of HPC environments.
Responsibilities
- Operate and maintain High-Performance Computing (HPC) clusters in accordance with defined service level agreements (SLAs).
- Design, develop, recommend, and implement new or enhanced system software, utilities, and automation processes.
- Perform system analysis, configuration management, and performance tuning for large-scale Linux-based compute environments.
- Support multiple projects independently with minimal supervision.
- Monitor and tune HPC systems to optimize performance across standalone and multi-tier environments.
- Diagnose hardware and operating system issues and implement long-term remediation solutions.
- Manage storage, backup, and disaster recovery processes to maintain data integrity.
- Implement and enforce security controls and infrastructure best practices.
- Conduct user training sessions, researcher workshops, project onboarding sessions, review workshops, and workshops related to cloud resources for Research HPC.
- Develop and maintain Bash scripts and automation tools to improve operational efficiency.
- Perform OS patching, kernel tuning, and lifecycle maintenance for RHEL-based Linux systems.
- Maintain technical documentation, standard operating procedures (SOPs), and operational runbooks.
- Collaborate with internal teams, leadership, and customers to resolve escalations and deliver support services.
- Contribute to process improvements, system enhancements, and technology modernization initiatives.
- Evaluate emerging HPC, cloud, and automation technologies to support organizational objectives and performance requirements.
Requirements
- Experience operating HPC clusters, including job schedulers, cluster management platforms, parallel programming libraries, and parallel file systems.
- Minimum 5 years of Linux/UNIX system administration experience, preferably in RHEL-based environments.
- Experience configuring, supporting, and maintaining systems hardware.
- Experience configuring, supporting, and using hypervisors and virtual machines across AWS, Azure, VMware, OpenStack, OpenShift, and other cloud platforms.
- Experience configuring, supporting, and using job schedulers such as PBS, PBS-Pro, and Slurm.
- Experience configuring, supporting, and using monitoring tools such as Nagios, Grafana, Prometheus, and Ganglia.
- Experience configuring, supporting, and using file systems including NFSv4, GPFS, BeeGFS, and Lustre.
- Experience configuring, supporting, and using high-speed, low-latency networking technologies such as InfiniBand.
- Experience configuring, supporting, and using identity management solutions such as Active Directory (AD), LDAP, and CentraDB.
- Experience configuring, supporting, and using parallel programming libraries including MPI, MPICH, OpenMPI, Intel MPI, and OpenMP.
- Experience working with programming languages and development tools such as GCC, ICC, Python, and R.