Key Responsibilities
Operate, monitor, and maintain enterprise platforms across development, test, staging, and production environments
Perform health checks, monitoring, troubleshooting, patching, upgrades, maintenance, and lifecycle activities
Support application and platform deployments and maintain operational documentation and runbooks
Monitor alerts, manage incidents, perform root-cause analysis, and support service restoration and problem management
Operate cloud and/or on-premises environments covering compute, storage, networking, containers, VMs, and platform services
Support AWS, Azure, or Google Cloud environments, backup, recovery, disaster recovery, and business continuity
Develop scripts and automation using Python, PowerShell, Bash, or similar technologies
Support CI/CD pipelines and Infrastructure as Code using Terraform, Ansible, or equivalent tools
Implement monitoring, logging, alerting, dashboards, and observability solutions
Support vulnerability remediation, security patching, system hardening, access controls, logging, and security assessments
Work with application, cybersecurity, network, database, DevOps, and vendor teams to resolve issues
Requirements
Essential
Degree/Diploma in Computer Science, IT, Engineering, or related discipline
3–6 years of relevant experience in platform operations, infrastructure, DevOps, systems administration, cloud operations, or production support
Hands-on production IT experience with Linux and/or Windows servers
Good understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls
Experience with monitoring, logging, incident management, troubleshooting, and IT service management
Experience with scripting/automation using Python, PowerShell, Bash, or similar
Good to Have
AWS, Azure, or Google Cloud experience
Kubernetes, Docker, or container technologies
CI/CD tools such as Jenkins, GitLab CI/CD, GitHub Actions, or Azure DevOps
Terraform, Ansible, or other Infrastructure as Code tools