We are seeking a Site Reliability Engineer to join our clients team in Singapore. In this role, you will be dedicated to supporting our Microsegmentation Policy Engine project. You will ensure the maximum uptime, reliability, scalability, and operational support of the policy engine and its related synchronization processes. You will play a crucial role in establishing enterprise-grade observability, refining production support workflows, and onboarding applications and infrastructure objects onto the platform.
Key ResponsibilitiesObservability & System Monitoring
- Design, build, and maintain robust operational monitoring, structured logging, performance metrics, distributed tracing, alerting, and reporting systems for the policy engine.
- Integrate core infrastructure and synchronization tasks with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
- Develop real-time dashboards to provide visibility into policy engine reliability, data synchronization status, and overall system health.
Reliability, Support & Incident Management
- Ensure high availability and seamless scalability of the microsegmentation policy engine and synchronization pipelines.
- Define, maintain, and refine operational runbooks, escalation paths, standard operating procedures (SOPs), and production support models.
- Drive incident response, triage, root-cause analysis (RCA), and continuous remediation to improve platform resilience.
Onboarding & Operations
- Support the onboarding of applications, infrastructure objects, and policy data into the microsegmentation platform.
- Perform data quality reviews and coordinate closely with client stakeholders and engineering teams for issue remediation.
- Create comprehensive technical documentation, deployment notes, and knowledge-transfer materials for cross-functional teams.
Technical Skills & QualificationsMust-Have / Required Skills
- Experience: 5+ years of hands-on experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
- Systems & Scripting: Strong knowledge of Linux administration and core scripting languages (Python strongly preferred for automation, Shell scripting).
- Containerization & Orchestration: Strong, hands-on experience with Docker and Kubernetes.
- CI/CD & Version Control: Proficient in Git and building/maintaining automated CI/CD pipelines.
- Observability & Telemetry: Experience configuring observability platforms (logging, metrics, tracing, alerting, dashboarding).
- Operations & Integration: Experience integrating platforms with enterprise monitoring, SIEM, and Incident Response workflows.
- Process & Documentation: Proven experience creating and executing runbooks, operational procedures, and support models.
Desired / Good-to-Have Skills
- Domain Knowledge: Prior experience working with policy engines, security infrastructure, or microsegmentation architectures is highly desirable.
- Stakeholder Management: Strong communication skills to coordinate data remediation and technical onboarding directly with business and technical stakeholders.
Resource & Location Requirements
- Role Title: Senior Site Reliability Engineer (SRE)
- Region/Location: Singapore
- Minimum Experience: 5+ years of relevant industry experience
Pay: $5,000.00 - $6,000.00 per month
Experience:
- Site Reliability Engineer: 5 years (Required)
Work Location: In person