Talent Link by e2i is a programme to match candidates to job opportunities offered by e2i’s Industry Partners. Applicable for Singaporeans and Singapore Permanent Residents only.
This job opportunity is from our Industry Partner in the Data Centre sector.
Job Descriptions
- Responsible for ensuring the continuous availability, performance, and reliability of application infrastructure hosted within the data centre.
- Manage and maintain physical and virtual servers supporting business-critical applications.
- Monitor system health, capacity, and performance to ensure high availability.
- Manage and review the administration of operating systems, virtualization platforms, storage, and server hardware to ensure maintenance activities are completed accurately and on schedule.
- Manage and review the maintenance of UPS systems, server racks, and environmental monitoring within the data centre.
- Oversee routine maintenance, patching, firmware upgrades, and infrastructure lifecycle management.
- Review and verify change requests and service requests, ensuring implementation is completed correctly and within scheduled engineering windows.
- Lead incident response, troubleshooting, and recovery activities for server and infrastructure-related issues.
- Develop disaster recovery procedures and conduct failover and restoration testing.
- Coordinate with vendors and stakeholders during planned maintenance and incident resolution.
- Prepare operational reports, technical documentation, and standard operating procedures.
- Participate in 24/7 standby or on-call support for critical infrastructure incidents.
Job Requirements
- Diploma or Degree in Computer Science, Information Technology, Engineering, or a related field.
- Preferably 6 to 10 years of relevant experience in infrastructure engineering or systems administration.
- Experience administering Red Hat Enterprise Linux servers.
- Experience with container technologies such as Podman.
- Knowledge of PostgreSQL database administration.
- Familiarity with virtualization technologies such as VMware or Hyper-V.
- Experience with enterprise storage, backup solutions, and data centre operations.
- Strong troubleshooting, analytical, and root cause analysis skills.
- Good communication, interpersonal, and documentation skills.
- Able to work independently and collaboratively in a team environment.
- Willing to participate in 24/7 standby or on-call support for critical system incidents.