Role Summary
Our Product Development team is looking for a Site Reliability Engineer to help build reliable, scalable, and high-performing products. As we transform into a product-led organization, you will play a key role in improving system reliability, operational efficiency, and cost effectiveness through automation, monitoring, and data-driven insights.
You will collaborate with technical experts, contribute to our future technology strategy, and help deliver greater value to our clients. The ideal candidate is a engineer with strong experience designing and implementing solutions on AWS, together with a passion for continuous improvement, and operational excellence.
What You'll Be Doin
- Delivery of high-quality infrastructure as code solutions and CI/CD pipelines
- Implementation of monitoring solutions for client-facing products and internal data pipelines, including intelligent alarming for quicker incident detection and resolution
- Track and manage infrastructure & application vulnerabilities
- You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
- Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability.
- You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders.
- Development of and adherence to policies and procedures to ensure all work is executed to a high standard
- You will be reporting to Software Engineering Manager
Qualifications
- 5 years+ of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or a related role.
- Proven experience managing AWS workloads, particularly ECS and EKS clusters.
- Strong understanding of high availability, disaster recovery, platform and performance monitoring.
- Hands-on experience with DevOps and Infrastructure-as-Code tools, including Terraform and Jenkins.
- 3+ years of experience using data and monitoring insights to improve AWS performance, reliability, and cost efficiency.
- Complex technical matters conveyed clearly and drive alignment across technical and product teams.
- Strong mindset with the initiative to innovate, automate processes, and continuously improve platform performance.