jobs in Confidential Jobs

全职 Site Reliability Engineer 工作, 薪水, Confidential Jobs Federal Territory 公司招聘中 - Ricebowl

Site Reliability Engineer

Confidential Jobs

KL City, Federal Territory

分享
保存

工作地点

  • Jalan Sultan Mizan Zainal Abidin, Kompleks Kerajaan Kuala Lumpur Federal Territory Malaysia

职位描述

岗位职责

About the Role


As a Site Reliability Engineer, you will play a key role in owning the reliability, availability, and performance of production systems at scale. You will work closely with development, security, and compliance teams to design and enforce SLOs/error budgets, drive incident response, and build the automation that keeps the infrastructure resilient and audit-ready. This position is central to leading a modern SRE practice.


Key Responsibilities

  • Own reliability for production systems at scale — defining and tracking SLIs/SLOs, error budgets, and reliability roadmaps
  • Administer and harden Linux/Unix systems, and build automation and tooling in Python, Go, or similar languages to reduce manual toil
  • Manage infrastructure as code using Terraform, Ansible, and operate containerized workloads with Kubernetes and Docker
  • Build and maintain CI/CD pipelines (GitHub actions, Argo CD, Kargo) to enable safe, frequent, self-service deployments
  • Design and operate infrastructure across AWS, GCP, Azure or AliCloud (multi-cloud is a plus), with solid grounding in networking fundamentals — DNS, load balancing, TCP/IP, CDNs
  • Build and maintain observability — metrics, logs, and traces via Prometheus, Grafana, Datadog, or OpenTelemetry — to catch issues before they become incidents
  • Drive incident response and on-call rotations, including leading postmortems and blameless root-cause analysis, with clear, calm communication during active incidents
  • Run resilience testing to validate failover, redundancy, and recovery assumptions
  • Own capacity planning, performance tuning, and database reliability (replication, backups, failover) for critical systems
  • Analyze incidents to identify systemic issues, and lead initiatives that measurably reduce incident frequency and MTTR
  • Design and implement change management processes that satisfy audit and compliance requirements without slowing down engineering velocity
  • Apply and enforce security baseline standards across infrastructure and deployment pipelines
  • Lead or significantly contribute to the transformation from legacy operational practices to modern SRE workflows (automation-first, self-service, reduced toil)
  • Act as a technical lead on cross-functional reliability initiatives
  • Partner with development, security, and compliance stakeholders to embed reliability and audit-readiness into the SDLC
  • Champion SRE best practices — blameless postmortems, toil reduction, capacity planning, and progressive rollouts


Required Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Systems Engineering, or a related field
  • 3-5+ years of hands-on experience in SRE, DevOps, or infrastructure engineering roles
  • Demonstrated, hands-on track record designing and owning reliability for production systems at scale
  • Solid grounding in SRE principles (SLOs, error budgets, toil, blameless postmortems)
  • Experience building or operating change management processes aligned with audit requirements
  • Familiarity applying security baseline standards such as CIS benchmarks and NIST


重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多