Configure and maintain monitoring thresholds, dashboards, alerts, logs, and escalation procedures; respond to alerts and service degradation before they become prolonged outages.
Investigate and resolve infrastructure incidents involving server availability, CPU/memory utilisation, disk/storage, network connectivity, operating systems, cloud services, and system performance.
Perform root cause analysis for outages and recurring technical issues, document findings, and implement or recommend corrective and preventive actions.
...