• Incident Detection & First-Level Troubleshooting: Detect and investigate incidents, perform initial troubleshooting and escalate unresolved issues to the
appropriate support teams.
• Ticket & Incident Management: Log, update, track and close incidents and service requests in accordance with agreed SLAs and operational procedures.
...
We are seeking a highly skilled Site Reliability Engineer (SRE) with 8+years of experience. The ideal candidate will possess a deep understanding of SRE principles and practices, ensuring the reliability, availability, and performance of our systems. You will work closely with development and operations teams to implement best practices in system reliability and automation.
Responsibilities:
Design, implement, and maintain scalable and reliable systems and services.
...