Architect and maintain enterprise-grade infrastructure observability, log aggregation, and self-healing patterns across all group application clusters. This includes configuring active runtime security instrumentation to detect system-call anomalies on production nodes, establishing metric-driven alerting thresholds, and configuring declarative GitOps controllers (Argo CD) to instantly detect and automatically correct cluster drift back to the designated Git configuration.
Own the reliability, scale, and disaster recovery strategies for the entire platform delivery lifecycle. The candidate will be responsible for defining and testing deterministic, zero-downtime rollback strategies—such as utilizing automated Git reverts to instantly return environments to previous stable states while managing cluster capacity, isolating failure domains, and ensuring the high availability of critical paths across Dev, SIT, UAT, and Production environments.
Deliver and operate golden paths that a new service can adopt in under a day, with CI/CD, observability, security gates and deployment wired in by default.
...