jobs in St. Jude Medical

St. Jude Medical Hiring! Full Time Remote Care Operations (Site Reliability Engineer) in Pulau Pinang - Ricebowl

Remote Care Operations (Site Reliability Engineer)

Undisclosed

George Town, Pulau Pinang

Share
Save

Working Location

  • Penang George Town Pulau Pinang Malaysia

Job Description

Responsibilities

We are seeking a proactive and reliability-focused engineer to strengthen the operational excellence of Remote Care platforms. This role is responsible for ensuring system stability, performance, and availability by delivering operational support, incident management, and continuous service improvements.

Key Responsibilities

1. Operational Support

  • Support production operations and issue investigations.

  • Participate in troubleshooting and cross-functional escalations.

  • Ensure operational stability and timely issue resolution.

2. Incident Management

  • Lead and coordinate incident response activities for production systems and services.

  • Perform incident triage, impact assessment, and root cause analysis (RCA).

  • Drive timely resolution of critical issues and ensure effective stakeholder communication throughout the incident lifecycle.

  • Track recurring incidents, identify systemic issues, and implement preventive measures to reduce operational risk.

  • Maintain incident documentation, post-mortem reports, and continuous improvement actions.

3. System Performance Monitoring and Tuning

  • Monitor system health, performance metrics, and service availability across Remote Care platforms.

  • Analyze trends and proactively identify potential performance bottlenecks or reliability risks.

  • Collaborate with engineering teams to optimize application, database, and infrastructure performance.

  • Develop and enhance monitoring dashboards, alerting mechanisms, and operational KPIs.

  • Recommend and implement performance tuning initiatives to improve system stability, scalability, and customer experience.

4. Quality

  • Assist with monitoring and addressing system performance and quality issues.

  • Identify trends of system health, engage in data review and provide recommendations for improvement.

  • Provide updates to stakeholders with regards to incident reports and performance metrics.


We Are Looking For

Required

  • Bachelor’s degree in Software Engineering, Systems Engineering, IT, Computer Science, or related field.

  • Minimum 3 years of experience working in site reliability, software engineering, computer engineering or a related technical discipline.

  • Strong analytical, problem-solving, and data interpretation skills.

  • Excellent communication skills with the demonstrated ability to work effectively in cross-functional teams and translate technical complexity for non-technical stakeholders.

  • Experience in software quality, system operations, or reliability engineering.

  • Ability to work with complex systems and large datasets.

  • Strong communication and stakeholder management skills.

  • Proficiency in at least one systems or automation programming language (e.g., Python, PowerShell) for building tooling, automation, and operational system.

Preferred

  • Experience with metrics visualization and reporting tools.

  • Familiarity with automation frameworks and CI/CD pipelines.

  • Knowledge of cloud environments and distributed systems.

  • Experience in regulated industries such as medical devices or healthcare.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More