Role Overview
The Platform Operations Engineer supports on-premises and hybrid infrastructure platforms that underpin mission-critical systems. This role focuses on platform reliability, operational stability, and continuous infrastructure improvement within a government or regulated environment.
The incumbent will collaborate with internal teams and L1 vendor engineers to oversee day-to-day operational support, ensure timely resolution of incidents and issues, and provide technical analysis and recommendations when required.
The role also includes technical and security governance responsibilities, such as reviewing technical changes, supporting service request (SR) reviews and approvals, and ensuring infrastructure operations and changes align with operational, security, and architectural standards.
In addition, the role supports a multi-year technology refresh programme while maintaining stable day-to-day operations with minimal disruption. Responsibilities include coordinating technical infrastructure initiatives, tracking dependencies, managing risks and issues, and driving collaboration across internal stakeholders and vendors to support the organisation's transition towards cloud and hybrid infrastructure environments.
Key Responsibilities
Infrastructure Operations & Support
- Work with L1 vendor engineers and internal teams to oversee the day-to-day operations, support, and maintenance of critical infrastructure platforms across on-premises and hybrid environments.
- Provide technical oversight across compute, storage, virtualisation, operating systems, backup, disaster recovery (DR), high availability (HA), and related infrastructure platforms.
- Lead incident management for complex issues, including technical investigation, root cause analysis, recovery planning, and resolution tracking.
- Coordinate vendor activities to ensure operational tasks, incidents, service requests, and technical deliverables are completed effectively and within agreed timelines.
- Partner with application, network, security, and vendor teams to resolve platform-related issues and maintain runbooks, SOPs, governance artefacts, and supporting documentation.
Governance & Compliance
- Review and assess technical changes, service requests, implementation plans, and recovery approaches to ensure they are practical, supportable, and aligned with operational, security, and architectural requirements.
- Support technical and security governance activities, including change reviews, operational readiness assessments, risk management, and compliance-related activities.
Infrastructure Modernisation & Project Coordination
- Support infrastructure modernisation initiatives and technology refresh programmes while maintaining operational stability and service availability.
- Plan, coordinate, and track infrastructure activities across internal teams and vendors, including refresh, upgrade, migration, and related technical projects.
- Monitor project risks, issues, dependencies, action items, and status updates to support successful delivery and operational readiness.
Requirements
Experience & Competencies
- Relevant experience in infrastructure or platform operations, preferably within production or mission-critical environments.
- Strong understanding of infrastructure operations across virtualisation, operating systems, storage, backup, networking, and related platforms.
- Experience working within vendor-supported operating models and coordinating external support teams.
- Ability to participate in technical troubleshooting, problem management, and resolution of complex incidents.
- Experience supporting technical governance, security governance, change management, or service request review processes is advantageous.
- Experience supporting technical project delivery, including coordination, risk and issue tracking, dependency management, and stakeholder engagement.
- Experience working in government, public sector, regulated, or similarly governed environments is preferred.
- Data centre operations experience will be advantageous.
- Strong communication, stakeholder management, and cross-functional coordination skills.
Technical Requirements
- Experience supporting infrastructure platforms such as VMware, Hyper-V, Nutanix, or equivalent technologies.
- Experience working with Red Hat Linux and/or Windows Server environments.
- Familiarity with enterprise storage, backup, high availability (HA), and disaster recovery (DR) solutions.
- Strong understanding of monitoring, observability, logging, and alerting practices.
- Solid understanding of networking fundamentals, including TCP/IP, DNS, routing, firewalls, load balancing, and network segmentation.
- Familiarity with hybrid infrastructure environments, including AWS, Microsoft Azure, or equivalent cloud platforms.
Desired Technical Skills
- Experience with automation tools such as Ansible, Puppet, Chef, Terraform, or similar technologies.
- Familiarity with scripting languages such as Python, PowerShell, Bash, or equivalent.
- Familiarity with containerisation and orchestration platforms such as Docker and Kubernetes.
- Experience implementing Infrastructure as Code (IaC) for automated provisioning and configuration management.
Qualifications
- Degree in Computer Science, Engineering, Information Technology, or a related discipline; or equivalent relevant professional experience and qualifications.
Preferred Certifications
- VMware Certified Professional (VCP)
- Microsoft Certified: Windows Server
- Red Hat Certified Engineer (RHCE)
- ITIL 4 Foundation
- AWS Certified Associate or Microsoft Azure Associate certification (or equivalent)