600+ Reliability Jobs - September 2026 - High Salaries

Showing 671 jobs results for "reliability"
Never miss any updates for Reliability jobs

Singapore

  • Advise teams on suitable failure analysis techniques and ensure all potential failure mechanisms are evaluated
  • Collaborate with cross‑functional teams to support new package failure analysis readiness
  • Develop and enhance failure analysis capabilities, techniques, and lab processes for better efficiency and detection capability ...
Posted
3 days ago

Singapore

  • • Prepare and maintain:
  • - RAM Reports
  • - ILS Analysis Reports ...
Posted
3 days ago

Singapore

  • Manage and coordinate vendor activities to ensure operational tasks, incidents, service requests, and technical deliverables are completed effectively and within expected timelines.
  • Review and assess technical changes, service requests (SRs), implementation plans, and recovery approaches to ensure they are practical, supportable, and aligned with operational, security, and architectural requirements.
  • Support technical and security governance activities, including change review, risk awareness, compliance alignment, and operational readiness checks. ...
Posted
18 days ago

KL City

  • A key objective of this role is to drive AIOps-enabled operations and self-healing automation to reduce alert noise, improve Mean Time To Detect (MTTD), reduce Mean Time To Resolve (MTTR), and improve overall platform and application reliability.
  • Roles and Responsibilities:
  • Observability Architecture & Design ...
Posted
4 days ago

KL City

  • Plan and execute roadmap for strategic infrastructure improvement incorporating initiatives that align with the company goals.
  • Bachelor's or above degree in Computer Science, Software Engineering, or related field
  • Strong interest in operations and technical risk management ...
Posted
5 days ago

Singapore

  • Integrate Splunk with incident management, ITSM, and AIOps systems to enable predictive alerting and anomaly detection.
  • Act as the SIEM/Splunk subject matter expert (SME) for architecture reviews, platform upgrades, and performance tuning.
  • Implement and champion SRE frameworks and reliability practices for mission-critical systems. ...
Posted
6 days ago

Singapore

  • § Responsible to ensure documentation of problem-solving process, including the steps taken, findings, and resolutions. Share this information with relevant stakeholders, support teams, and knowledge repositories to facilitate learning and prevent similar issues in the future.
  • § Managing complex systems due to interdependencies between components, where changes or failures in one area can have cascading effects on other parts of the system.
  • § Minimum 5 years of operating on cloud environments.
Posted
a month ago

Singapore

  • Collaboration and Engagement: Collaboration with stakeholders is essential for prioritizing and addressing production issues effectively. The SRE and Production Support Lead should engage stakeholders in incident response efforts, problem-solving activities, and decision-making processes to ensure alignment and buy-in
  • Support Optimization: Streamlining processes, leveraging automation, fostering collaboration, and continuously improving operational efficiency
  • Leading and managing a team of SREs and production support engineers is a key aspect of the role. The primary focus is to ensure the reliability and stability of the organization's systems and applications ...
Posted
a month ago

Singapore

  • Customization: Customize job labelling process for different FA groups.
  • Classification: Catergorize failure mode through AIML classification.
  • Interface Design: Develop user interface to facilitate Man-machine interaction for above items. ...
Posted
19 days ago

Singapore

  • Evaluate emerging technologies, startups, and foundation model solutions relevant to industrial AI, predictive maintenance, and condition monitoring, assessing their technical capabilities.
  • Lead and manage vibration-based predictive maintenance projects as the technical focal point, overseeing the full project lifecycle from initiation and deployment to business value realization, in close collaboration with nominated subcontractors and system integrators.
  • Bachelor’s/Master’s/PhD degree in Computer Science, Electrical Engineering, Mechanical Engineering, Mathematics, Statistics, Physics, or a related field from a reputable institution. ...
Posted
7 days ago

Singapore

  • Identify service risks, improvement opportunities and drive actions to enhance service performance.
  • Coordinate and produce service reports, review and provide business input to post incident reports, monitor and report service usage trends, and lead capacity forecasting and planning in line with the agreed models.
  • Provide governance for RTP service changes with internal stakeholders and the customer, safeguarding service stability and resilience by ensuring the appropriate service change process is followed and all service changes are tested, coordinated and executed properly. ...
Posted
19 days ago

Singapore

  • Collaborate with the team to address the long-term requirements of CD systems, optimising computing resource utilisation, and enhancing the consistency and efficiency of system functions.
  • Collaborate with the team to conceive, design, and implement cloud-based CI/CD pipelines and related functions, ensuring seamless support for diverse application development teams and facilitating smooth development and release processes.
  • Develop tools and features to monitor cloud environments, enhancing the visibility and understanding of system operations within the cloud infrastructure. ...
Posted
8 days ago

TSC OFFSHORE PTE. LTD.

Outram

  • Collaborate closely with the Operations team, acting as the commercial focal point to support project execution and service delivery.
  • Oversee chartering processes, including technical, operations, contracts, and commercial.
  • Resolve operational challenges and emergency issues as they arise, maintaining continuity of service. ...
Posted
a month ago

Singapore

  • Bachelor’s degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 3+ years experience working in ML infra (PyTorch, Vertex AI, Sagemaker, etc.)
  • Experience building, scaling, and optimizing enterprise-grade Machine Learning systems ...
Posted
9 days ago

Singapore

  • Act as the Control-M SME for architecture reviews, capacity planning, and performance tuning of the scheduling estate.
  • Define and enforce job-definition, naming, and governance standards—including retention, audit logging, and change control for the batch environment.
  • Champion SRE frameworks and reliability practices for mission-critical batch workloads, defining SLAs, error budgets, and service-health indicators for scheduled jobs. ...
Posted
9 days ago

Singapore

  • Excellent programming abilities with a strong command of data structures and fundamental algorithms. For traditional coding roles, proficiency in C/C++ is required; for intelligent coding roles, proficiency in Python is required. Candidates are required to use these languages to implement complex algorithms and build iterative models. Candidates should also have a strong engineering mindset with the ability to balance performance and cost;
  • Ability to effectively communicate and collaborate with team members, such as algorithm engineers, data analysts, and product managers, to explore new technologies and drive innovation in generative recommendation and search systems.
Posted
9 days ago

Singapore

  • Efficiency: Designing and implementing software platforms and monitoring frameworks for efficient, automated, and intelligent service-oriented architecture (SOA) governance.
  • Cost: There are millions of CPUs. We should build delivery standards, and monitor and budget systems to optimize the cost of the company.
  • Compliance: Designing and setting up new IDC; designing and implementing data protection plan to meet the standard requirement. ...
Posted
9 days ago

Malaysia

  • Take ownership of certain tools and equipment.
  • Bachelor in Electronics/ Electrical/ Microelectronics / Material Engineering, Material Science, Physics or Chemistry.
  • Minimum 8-10 years’ experience in silicon level failure analysis on IC failure. ...
Posted
9 days ago

KL City

  •  Understand the game architecture, analyze, evaluate and respond to potential risks, such as: Hidden troubles , performance bottlenecks,
  •  Responsible for daily communication and coordination between various teams
  • Job Requirements & Qualifications: ...
Posted
9 days ago

Singapore

  • Experiment with prompt engineering, fine-tuning, or orchestration of AI workflows
  • Cloud Operations Optimization
  • Analyze existing cloud infrastructure processes and identify opportunities for automation and optimization. ...
Posted
9 days ago

Singapore

  • Stability and security: Build comprehensive K8s cluster monitoring, alerting, logging, and distributed tracing systems; define operations runbooks, change processes, and incident response plans; strengthen cluster security controls, disable high-risk permissions, harden container runtime environments, and ensure infrastructure and business data security.
  • Automated operations and DevOps: Develop operations automation scripts using Shell/Python; integrate Jenkins, GitLab CI, and ArgoCD to build automated release, inspection, and backup systems; implement Infrastructure as Code (IaC) principles to improve efficiency and reduce human error.
  • Incident management and post-mortem optimization: Lead online incident response, conduct root cause analysis, produce post-mortem reports, and continuously optimize cluster architecture, resource allocation, monitoring strategy, and long-term stability assurance mechanisms. ...
Posted
9 days ago

Singapore

  • Exposure to state-of-the-art IT and datacenter technologies and large-scale various fleets.
  • Server Operations & Infrastructure Support: Assist in the deployment, monitoring, and maintenance of large-scale server fleets across our global datacenters.
  • Lifecycle Management: Support the full lifecycle of servers, from system design, deployment, operation, troubleshooting, and decommissioning. ...
Posted
9 days ago

Singapore

  • Collaborate with multi-disciplinary skilled professionals to address business needs and develop new capabilities
  • Degree in Information Technology, Computer Science, Computer Engineering, Electrical Engineering, or related discipline
  • Proficiency in one or more of the related Hyper-Converged Infrastructure, Cloud Network/Security/Platform Services technologies will be preferred (e.g. Nutanix, HyperV, SAN, Object and File Storage, IAM, Configuration and Compliance Management, etc.) ...
Posted
9 days ago

Singapore

  • Write and maintain automated tests and CI/CD-related configurations to improve development efficiency and code quality.
  • Support the engineering deployment of AI/LLM capabilities based on business needs, such as integrating LLM APIs, building basic RAG pipelines, and embedding tool/agent capabilities into existing systems.
  • Help build and maintain datasets and benchmarks for evaluation/regression, track online performance, and assist with debugging and fixing issues. ...
Posted
9 days ago

Bukit Panjang

  • Ensure the teams have the necessary qualifications and authorization to carry out daily tasks and any assigned tasks.
  • Attend necessary skills courses to be able to carry out maintenance and assigned tasks.
  • Responsible for the safety, welfare and discipline of the teams. ...
Posted
9 days ago

Singapore

Posted
9 days ago

Singapore

Posted
9 days ago

Singapore

Posted
9 days ago

KL City

  • Accelerate onboarding and enable successful adoption of platform capabilities.
  • Promote Compliance-as-Code and Policy-as-Code principles.
  • Embed security, compliance, and resilience controls into delivery pipelines. ...
Posted
9 days ago