You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability.
You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders.
...
You participate in a 24/7 on-call rotation and drive improvements using SRE practices.
You actively participate in toil elimination, observability and monitoring improvements, knowledge management, error budget compliance, deployment designs and testing.
Bachelor’s degree and/or equivalent experience in Information Technology, Computer Science or Business Management.
...
Regularly deploy product updates as required to keep the platform vulnerability-free.
Work with open-source technologies, CI/CD, SCM tools as necessary, and source control such as Bitbucket, implement organization containers (e.g., Docker and Kubernetes). Stay current with industry trends and propose new ways for the business to improve.
Take accountability in considering business and regulatory compliance risks and take appropriate steps to mitigate the risks.
...
You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability.
You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders.
...
You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability.
You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders.
...
You participate in a 24/7 on-call rotation and drive improvements using SRE practices.
You actively participate in toil elimination, observability and monitoring improvements, knowledge management, error budget compliance, deployment designs and testing.
Bachelor’s degree and/or equivalent experience in Information Technology, Computer Science or Business Management.
...
Provide L2/L3 production support, including troubleshooting complex platform, infrastructure, application and networking issues.
Participate in a 24/7 on-call rotation and respond to critical production alerts and incidents.
Lead or support major incident management, including troubleshooting, vendor coordination, immediate remediation, root cause analysis and long-term corrective actions.
...