Role Description The Site Reliability Engineer will be responsible for ensuring the reliability, availability, and performance of production systems by monitoring infrastructure, analyzing incidents, and implementing long-term improvements. Daily tasks include troubleshooting operational issues, collaborating with software development teams to design and implement scalable solutions, and automating deployment and maintenance processes. The role involves managing system administration activities, optimizing resource usage, and maintaining observability through logging, metrics, and alerting. The individual will also contribute to incident response, root cause analysis, and documentation of operational practices and runbooks.
Qualifications
- Candidates should possess strong Site Reliability Engineering skills, including experience with reliability best practices, monitoring, and incident management.
- Candidates should possess troubleshooting skills for diagnosing complex system, network, and application issues in production environments.
- Candidates should possess software development skills, with proficiency in one or more programming languages and experience in building automation and tooling.
- Candidates should possess system administration skills, including managing Linux/UNIX systems, configuration, security, and performance tuning.
- Candidates should possess infrastructure skills related to cloud platforms, containers, CI/CD pipelines, and infrastructure-as-code tools.
- Relevant experience with observability tools (metrics, logging, tracing) and high-availability architectures is beneficial.
- A bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience, is preferred.
- Strong communication, collaboration, and documentation abilities, with a proactive approach to continuous improvement, are expected.
We are looking for folks who can join on short notice. Kindly DM me if you are interested, or if you have a referral.