ROLE OVERVIEW
The Executive, System Engineer serves as the primary hands-on engineer for enterprise Linux and cloud platforms, while also providing dependable secondary support for Windows infrastructure. The role is responsible for their administration, monitoring, maintenance, security, automation, and operational resilience across data center, AWS, and private cloud environments, with particular focus on Red Hat Enterprise Linux (RHEL) and Amazon Linux. The role also provides versatile support for Windows Server, Active Directory, VMware, storage, backup, and disaster recovery services. Reporting to the Asst. Manager- System Engineering, the incumbent will resolve incidents, execute approved changes, maintain technical documentation, and ensure Linux services remain secure, stable, recoverable, and compliant.
KEY RESPONSIBILITIES
Linux, AWS & Private Cloud Engineering
- Administer, monitor, patch, troubleshoot, and optimize RHEL and Amazon Linux servers across physical, VMware, AWS, and private cloud environments.
- Manage Linux operating system installation, configuration, upgrades, repositories, packages, services, filesystems, LVM, permissions, sudo, SSH, and scheduled jobs.
- Provision and support workloads, templates, lifecycle management, capacity, availability, and integration with virtualization, storage, network, and identity services. and related AWS services and TIME Cloud Services.
- Support core Linux services and integrations including DNS, NTP, SMTP relay, certificates, authentication, directory integration, and network configuration.
- Respond to incidents and service requests within agreed service levels, escalating complex issues with clear diagnostics and impact details.
Linux Security, Automation & Compliance
- Execute Linux patching, vulnerability remediation, OS hardening, access reviews, certificate renewals, and configuration changes in accordance with change-management procedures.
- Apply secure Linux configuration standards, investigate operating system and authentication logs, and support endpoint, vulnerability, and privileged-access controls.
- Develop and maintain Bash, Python, Ansible, or AWS Systems Manager automation for provisioning, configuration, health checks, patching, reporting, and repetitive operational tasks.
- Maintain accurate asset, configuration, access, patch, and operational records to support internal audits and regulatory requirements.
- Follow applicable IT policies, ITIL practices, BNM RMiT, MoF, and ISO 27001 controls relevant to assigned systems and activities.
Backup, Recovery & Resilience
- Operate and monitor Linux storage, filesystems, mounts, multipathing, EBS volumes, backup agents, and recovery processes across SAN/NAS, AWS, and private cloud environments.
- Perform scheduled backup verification, Linux file and system restoration tests, and disaster recovery exercises; document results and address identified gaps.
- Support Linux clustering, high availability, service recovery, and business continuity arrangements for mission-critical services, including after-hours activities when required.
Documentation & Continuous Improvement
- Create and maintain RHEL and Amazon Linux build standards, SOPs, runbooks, architecture records, configuration documents, troubleshooting guides, and operational checklists.
- Identify recurring Linux issues and opportunities for automation, standardization, performance tuning, capacity optimization, and service improvement.
- Prepare clear operational reports, incident updates, root-cause inputs, and technical recommendations for review by the Asst. Manager.
Windows & Cross-Platform Support
- Administer and troubleshoot Windows Server operating systems, including patching, services, event logs, performance, storage, access, and approved configuration changes.
- Support Active Directory, Group Policy, DNS, DHCP, certificates, service accounts, authentication, and Windows integration with Linux and cloud workloads.
- Provide supporting coverage for VMware, storage, backup, monitoring, and disaster recovery platforms to maintain team versatility and service continuity.
- Participate in infrastructure migrations, cloud adoption, security, and resilience projects, coordinating with application, network, cybersecurity, service desk, vendor, and business teams.
WHAT DOES IT TAKE TO BE SUCCESSFUL
Qualifications
- Bachelor's Degree or Diploma in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications such as RHCSA/RHCE, AWS Certified SysOps Administrator or Solutions Architect, VMware VCP, Microsoft certification, or ITIL 4 Foundation are advantageous.
- Strong working knowledge of RHEL and Amazon Linux administration, with exposure to AWS, private cloud, Windows, virtualization, identity, storage, backup, disaster recovery, and cybersecurity practices.
- Awareness of BNM RMiT, MoF, ISO 27001, or other governance requirements in a regulated environment is advantageous.
Work Experience
- Approximately 2-5 years of hands-on experience in Linux systems administration, systems engineering, or infrastructure operations.
- Practical experience supporting Red Hat Enterprise Linux and Amazon Linux in enterprise environments; exposure to other Linux distributions is advantageous.
- Experience administering Linux in hybrid infrastructure covering data center, VMware, AWS, and private cloud environments.
- Working experience supporting Windows Server and Active Directory in an enterprise environment.
- Experience with incident, problem, change, patch, backup, monitoring, and technical documentation processes.
- Experience in financial services or another regulated industry is advantageous.
Knowledge, Skills & Competencies
- Strong hands-on RHEL and Amazon Linux administration and troubleshooting skills, including systemd, RPM/YUM/DNF, repositories, filesystems, LVM, permissions, SSH, sudo, logging, and performance analysis.
- Working knowledge of TCP/IP, DNS, NTP, SMTP, certificates, storage, backup, clustering, high availability, disaster recovery, and infrastructure monitoring.
- Practical ability to use Bash, Python, Ansible, AWS Systems Manager, or similar automation and configuration-management tools to improve operational efficiency.
- Sound analytical skills, attention to detail, and a structured approach to incident diagnosis and root-cause analysis.
- Good understanding of Linux access control, authentication, OS hardening, vulnerability remediation, secure configuration, and operational risk.
- Competence in Windows Server and Active Directory administration, with versatility across VMware, storage, backup, monitoring, and hybrid-cloud technologies.
- Clear written and verbal communication skills, including concise technical documentation and stakeholder updates.
- Ability to prioritize work, collaborate across teams, take ownership of assigned tasks, and participate in after-hours support when required.