Design, build and maintain infrastructure-as-code for our Sovereign AI platform using Terraform, Kubernetes and Docker.
Define and establish processes and standards for infrastructure provisioning, deployment and security — as an early team member, you'll be shaping how things get done, not just following an existing playbook.
Build and maintain CI/CD pipelines to support fast, reliable deployment of infrastructure and platform changes.
Operate and improve the reliability, scalability and performance of our cloud and on-prem infrastructure.
Implement monitoring, logging and alerting to support platform observability and operational readiness.
Apply security best practices across the infrastructure lifecycle — from image/container hardening and vulnerability scanning to access control and secrets management.
Support incident response and troubleshooting across infrastructure and deployment issues.
Work closely with the Head of Sovereign AI Infra and the rest of the team to continuously improve automation, tooling and platform maturity.
Contribute to securing AI/ML and high-performance computing (HPC) workloads as the platform evolves, including GPU infrastructure and model-serving environments.
Essential Experience
5–10 years of experience in a DevOps, Platform Engineering, SRE or DevSecOps role.
Strong, hands-on Linux skills — comfortable operating, troubleshooting and optimising at the OS and systems level.
Strong command of the open-source infrastructure ecosystem, with hands-on experience in Terraform, Kubernetes and Docker.
Hands-on experience with AWS or GCP in a production environment.
Experience building and maintaining CI/CD pipelines.