Analyze production issues, identify root causes, and implement long-term reliability improvements through automation, monitoring, and architectural enhancements.
Work collaboratively with other team members, provide technical leadership and guidance to a team of up to 10 SRE engineers, driving engineering excellence, reliability, and operational best practices.
Organize an efficient handover through high quality documentation and training.
...
Monitoring & Logging: Implement comprehensive system monitoring, alerting, and logging solutions to proactively detect and resolve performance bottlenecks or downtime.
Security & Compliance: Enforce infrastructure security best practices, access controls, and regular backup strategies across all deployment workflows.
Experience: 2+ years of hands-on experience in a DevOps, SRE, or Systems Engineering role.
...
Support RBC communication surveillance and related upstream / downstream applications through incident triage, root-cause analysis and service restoration.
Maintain production stability through monitoring, alert analysis, capacity awareness, runbook execution and risk escalation.
Work with CTB, RTB, vendor and cross-functional technology teams to support production fixes, enhancements, transition readiness and release/change activities.
...
Security & Governance: Implement enterprise security controls on-site, including IAM policy management, network security groups, encryption standards, and local regulatory compliance rules.
Operations & Troubleshooting: Perform real-time technical troubleshooting, root-cause analysis (RCA), and operational support directly within the client environment.
Client Engagement: Communicate technical progress, document architecture setups, and provide technical advisory to local stakeholder teams.
...