Job Overview
We are looking for an experienced Senior Observability Engineer to support and enhance enterprise monitoring and observability capabilities across hybrid infrastructure environments.
This is a hands-on technical role covering Windows, Linux, Azure, Kubernetes, monitoring, automation and incident response. You will help maintain highly available IT services by improving system visibility, identifying performance issues, strengthening alerting and automating operational activities.
The role is suited for an experienced infrastructure or SysOps professional who has progressed into cloud, observability, automation and reliability engineering within a large-scale 24/7 environment.
Key Responsibilities
Observability & Monitoring
- Manage and enhance monitoring solutions across on-premises and cloud infrastructure.
- Configure and support tools such as SolarWinds, Datadog, Azure Monitor, Grafana and Splunk.
- Build dashboards covering infrastructure health, application performance, logs, events and service availability.
- Improve monitoring effectiveness by tuning alerts, reducing unnecessary notifications and identifying service-impacting events.
- Support anomaly detection, event correlation and proactive identification of potential outages.
- Establish operational health checks, performance baselines and capacity monitoring.
Infrastructure & Cloud Operations
- Support enterprise Windows Server, Linux, Azure and Kubernetes environments.
- Monitor system availability, performance, capacity and overall platform health.
- Support high-availability and disaster recovery activities.
- Work with infrastructure and application teams to troubleshoot complex production issues.
- Support configuration management and Infrastructure-as-Code practices.
Automation & AIOps
- Develop automation using PowerShell, Bash, Python, Ansible and/or Terraform.
- Automate monitoring, health checks, reporting, configuration and remediation activities.
- Explore AI-assisted operations for log analysis, alert correlation, incident triage and operational reporting.
- Maintain reusable scripts, automation workflows and technical runbooks.
Incident & Problem Management
- Respond to critical infrastructure and application incidents in a 24/7 operational environment.
- Perform root cause analysis and recommend corrective actions.
- Support major incident reviews and identify opportunities to prevent recurring issues.
- Maintain escalation procedures, recovery documentation and operational runbooks.
Security & IT Service Management
- Support patching, vulnerability remediation, system hardening and access controls.
- Follow established incident, problem, change and configuration management processes.
- Support compliance requirements including PCI-DSS and ISO/IEC 27001.
- Maintain accurate infrastructure, service and configuration information within CMDB and related repositories.
- Work closely with infrastructure, applications, cybersecurity, vendors and managed service providers.
Requirements
- Minimum 8 years of experience in Systems Administration, Infrastructure Operations, SysOps, SRE or a related technical environment.
- Strong hands-on experience with Windows Server and Linux.
- Experience supporting Microsoft Azure, including IaaS, PaaS and cloud security.
- Experience with Kubernetes and containerized environments.
- Hands-on knowledge of monitoring/observability platforms such as SolarWinds, Datadog, Azure Monitor, Grafana or Splunk.
- Scripting or automation experience using PowerShell, Bash or Python.
- Exposure to Terraform, Ansible, CI/CD or Infrastructure as Code would be advantageous.
- Strong production troubleshooting, incident management and root cause analysis capabilities.
- Understanding of ITIL processes and enterprise change management.
- Knowledge of security and compliance requirements such as PCI-DSS and ISO/IEC 27001.
- Experience using AI/LLM tools for operational support, log analysis or incident investigation would be an advantage.
- Relevant certifications such as Microsoft Azure Administrator, AWS SysOps Administrator or CompTIA Security+ are advantageous.
- Comfortable supporting a business-critical environment where evening, weekend or on-call support may occasionally be required.
Please send your detailed resume in MS Word format to ************* with
- Education Level
- Working experiences
- Each employment background
- Reason for leaving each employment
- Last drawn salary
- Expected salary
- Date of availability