jobs in U3 INFOTECH PTE. LTD.

U3 INFOTECH PTE. LTD. Hiring! Full Time Observability Engineer in Islandwide (Singapore), Earn up to SGD 7,000 - Ricebowl

Islandwide (Singapore)

Share
Save

Working Location

  • Islandwide (Singapore) Singapore

Job Description

Responsibilities

Key Responsibilities:

Observability Strategy and Governance

  • Define and own Enterprise Observability Architecture aligned with operational resilience mandates (MAS TRM, DORA, APRA CPS 230).
  • Deploy and optimize observability platforms (Datadog, Dynatrace, Splunk) for full-stack visibility across infra, application, network, and user experience.
  • Establish governance standards for telemetry data (metrics, logs, traces), ensuring consistency, retention compliance, and security controls.
  • Integrate observability platforms with incident management, ITSM, and AIOps systems for predictive alerting and anomaly detection.

Reliability Engineering and Automation

  • Implement SRE frameworks for infrastructure and business-critical applications.
  • Automate runbooks, alerts, self-healing actions, and auto-remediation workflows via Python, Ansible, and Terraform.
  • Partner with Application, Infrastructure, and Cyber teams to codify operational reliability into the delivery lifecycle.
  • Conduct resilience testing, chaos engineering, and capacity validation.
  • Develop error budget policies and reliability scorecards for key production services.

Cloud Observability and Platform Engineering

  • Architect and manage observability for Cloud-native workloads in AWS and Azure.
  • Integrate cloud observability into landing zones and CI/CD pipelines for continuous compliance.
  • Implement IaC models using Terraform and Ansible for consistent, auditable provisioning.
  • Collaborate with Cloud, DevOps, and Security teams on real-time telemetry aligned to audit requirements.

Operational Excellence and Stakeholder Management

  • Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.
  • Deliver executive dashboards highlighting availability, reliability KPIs, and operational risk indicators.
  • Act as technical advisor to senior management during major incidents, post-incident reviews, and audits.

Skillset Requirements:

  • At least 5 years of experience in Infrastructure, Cloud, or Site Reliability Engineering (SRE) related roles, with minimum 3 years in an SRE SME capacity, ideally within financial institutions or regulated environments.
  • Hands-on expertise with Observability Platforms: Datadog, Dynatrace, Splunk, ELK.
  • Hands-on expertise with Automation/IaC: Terraform, Ansible, Python, CI/CD tools.
  • Hands-on expertise with Cloud Platforms: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).
  • Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.
  • Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
  • A good team player with excellent written and verbal communication skills, able to coordinate across diverse stakeholders.
  • Certification in at least one of the following required:
  • Datadog Certified Observability Professional / Dynatrace Certified Associate
  • Terraform/Ansible/Python Certified Expert
  • Certification in the following will be advantageous:
  • AWS Certified DevOps Engineer / Azure DevOps Expert
  • SRE Foundation/Practitioner (DevOps Institute)
  • ITIL v4 Managing Professional

    U3 Privacy Notice: 
    Please refer to U3’s Privacy Notice for Job Applicants/Seekers at ************* When you apply, you voluntarily consent to the collection, use and disclosure of your personal data for recruitment/employment and related purposes.

    Important Information

    Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

    Learn More