jobs in GOLDTECH RESOURCES PTE LTD

全职 Reliability Engineer (Observability) 工作, 薪水, GOLDTECH RESOURCES PTE LTD 公司招聘中 - Ricebowl

Reliability Engineer (Observability)

GOLDTECH RESOURCES PTE LTD

Undisclosed

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

Job Overview

We are looking for an experienced Senior Observability Engineer to support and enhance enterprise monitoring and observability capabilities across hybrid infrastructure environments.

This is a hands-on technical role covering Windows, Linux, Azure, Kubernetes, monitoring, automation and incident response. You will help maintain highly available IT services by improving system visibility, identifying performance issues, strengthening alerting and automating operational activities.

The role is suited for an experienced infrastructure or SysOps professional who has progressed into cloud, observability, automation and reliability engineering within a large-scale 24/7 environment.


Key Responsibilities

Observability & Monitoring

  • Manage and enhance monitoring solutions across on-premises and cloud infrastructure.
  • Configure and support tools such as SolarWinds, Datadog, Azure Monitor, Grafana and Splunk.
  • Build dashboards covering infrastructure health, application performance, logs, events and service availability.
  • Improve monitoring effectiveness by tuning alerts, reducing unnecessary notifications and identifying service-impacting events.
  • Support anomaly detection, event correlation and proactive identification of potential outages.
  • Establish operational health checks, performance baselines and capacity monitoring.

Infrastructure & Cloud Operations

  • Support enterprise Windows Server, Linux, Azure and Kubernetes environments.
  • Monitor system availability, performance, capacity and overall platform health.
  • Support high-availability and disaster recovery activities.
  • Work with infrastructure and application teams to troubleshoot complex production issues.
  • Support configuration management and Infrastructure-as-Code practices.

Automation & AIOps

  • Develop automation using PowerShell, Bash, Python, Ansible and/or Terraform.
  • Automate monitoring, health checks, reporting, configuration and remediation activities.
  • Explore AI-assisted operations for log analysis, alert correlation, incident triage and operational reporting.
  • Maintain reusable scripts, automation workflows and technical runbooks.

Incident & Problem Management

  • Respond to critical infrastructure and application incidents in a 24/7 operational environment.
  • Perform root cause analysis and recommend corrective actions.
  • Support major incident reviews and identify opportunities to prevent recurring issues.
  • Maintain escalation procedures, recovery documentation and operational runbooks.

Security & IT Service Management

  • Support patching, vulnerability remediation, system hardening and access controls.
  • Follow established incident, problem, change and configuration management processes.
  • Support compliance requirements including PCI-DSS and ISO/IEC 27001.
  • Maintain accurate infrastructure, service and configuration information within CMDB and related repositories.
  • Work closely with infrastructure, applications, cybersecurity, vendors and managed service providers.


Requirements

  • Minimum 8 years of experience in Systems Administration, Infrastructure Operations, SysOps, SRE or a related technical environment.
  • Strong hands-on experience with Windows Server and Linux.
  • Experience supporting Microsoft Azure, including IaaS, PaaS and cloud security.
  • Experience with Kubernetes and containerized environments.
  • Hands-on knowledge of monitoring/observability platforms such as SolarWinds, Datadog, Azure Monitor, Grafana or Splunk.
  • Scripting or automation experience using PowerShell, Bash or Python.
  • Exposure to Terraform, Ansible, CI/CD or Infrastructure as Code would be advantageous.
  • Strong production troubleshooting, incident management and root cause analysis capabilities.
  • Understanding of ITIL processes and enterprise change management.
  • Knowledge of security and compliance requirements such as PCI-DSS and ISO/IEC 27001.
  • Experience using AI/LLM tools for operational support, log analysis or incident investigation would be an advantage.
  • Relevant certifications such as Microsoft Azure Administrator, AWS SysOps Administrator or CompTIA Security+ are advantageous.
  • Comfortable supporting a business-critical environment where evening, weekend or on-call support may occasionally be required.


Please send your detailed resume in MS Word format to ************* with

  • Education Level
  • Working experiences
  • Each employment background
  • Reason for leaving each employment
  • Last drawn salary
  • Expected salary
  • Date of availability

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多