jobs in Ambition

全职 VP, Site Reliability Engineer 工作, 薪水, Ambition 公司招聘中 - Ricebowl

VP, Site Reliability Engineer

Ambition

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

We are looking for an experienced and hands-on, Vice President of Site Reliability Engineering to lead the reliability, availability, and resilience strategy across critical platforms and services. This individual will establish SRE as a core engineering discipline, driving automation, observability, incident management, and reliability-by-design practices across the organization.


Key Responsibilities


  • Define and execute the enterprise SRE strategy, embedding reliability engineering across products, platforms, and services.
  • Own and govern SLIs, SLOs, and error budgets, ensuring reliability decisions are data-driven.
  • Drive improvements in service availability, resilience, recoverability, and performance.
  • Lead the strategy for observability, automation, self-healing capabilities, and resilience engineering platforms.
  • Oversee incident management, major incident response, post-incident reviews, and chaos engineering initiatives.
  • Drive operational excellence through automation, toil reduction, and platform standardization.
  • Partner with Engineering, Infrastructure, Security, and Product teams to embed reliability requirements early in the development lifecycle.
  • Build, mentor, and lead a high-performing team of SRE leaders and engineers.
  • Influence technical decisions across teams and stakeholders, driving reliability outcomes in a matrixed environment.


Required Qualifications


  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
  • At least 12 years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Infrastructure, DevOps, or related disciplines.
  • Proven experience building, scaling, or transforming SRE functions within complex enterprise environments.
  • Strong expertise in:
  • SRE principles and automation-first operations
  • Observability and monitoring frameworks
  • Resilience engineering and incident management
  • Capacity planning, performance optimization, and disaster recovery
  • Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budget management
  • Solid software engineering background with experience in Java and developing or supporting large-scale distributed systems.
  • Experience designing and implementing observability, automation, self-healing, or reliability platforms.
  • Strong stakeholder management skills with the ability to influence engineering teams and senior leaders without direct authority.


重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多