800+ Reliability Jobs - October 2026 - High Salaries

Showing 850 jobs results for "reliability"
Never miss any updates for Reliability jobs
  • Monitor and analyze equipment reliability, availability, and maintenance performance, and track and report reliability KPIs such as MTBF, MTTR, PM compliance, and equipment availability.
  • Lead troubleshooting and Root Cause Analysis (RCA) for equipment failures and recurring issues.
  • Develop and optimize Preventive Maintenance (PM), Condition-Based Maintenance (CBM), and Predictive Maintenance (PdM) programs. ...
Posted
4 days ago

KL City

  • Drive the development and execution of cutting-edge equipment strategies and predictive maintenance programs to maximize reliability and minimize downtime.
  • Lead advanced reliability investigations using sophisticated techniques like Fault Tree Analysis and Root-Cause Failure Analysis to swiftly resolve critical issues.
  • Engineer and optimize annual maintenance plans, leveraging preventive and predictive tasks to ensure peak equipment performance. ...
Posted
4 days ago

KL City

  • • Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.
  • • Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.
  • • Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams. ...
Posted
3 days ago

Singapore

Posted
3 days ago

Singapore

Posted
4 days ago

Singapore

  • Analyze equipment performance, failure history, and plant data (including predictive monitoring tools) to detect degradation trends and prevent failures
  • Support development and optimization of maintenance strategies, including preventive, predictive, and reliability-centered maintenance (RCM)
  • Track and report reliability KPIs (e.g., MTBF, MTTR, availability, OEE) and support continuous improvement initiatives ...
Posted
4 days ago

Malacca City

  • Perform and coordinate internal package analysis, utilizing failure analysis (FA) tools and techniques to evaluate stressed samples and generate comprehensive technical reports.
  • Support reliability assessments by providing technical recommendations and expert consultation to project teams.
  • Coordinate and execute second-level reliability evaluations to ensure product robustness and qualification readiness. ...
Posted
5 days ago

Malacca City

  • Maintenance / repair / modification / optimization / buy-off Automatic Test Equipment, Load / DUT boards, adapter & socket, handlers and temperature conditioning system, and resolve all the issues with low & medium complexity / quantity
  • Review, update, and documentation (work instruction, specification, OJTI..)
  • Monitor and report on test activities and provide relevant information of machine down / malfunction, early warning on delays and bottlenecks to the Head of Test Engineering ...
Posted
5 days ago

Singapore

  • Monitor power, cooling, environmental, and IT infrastructure health across multiple sites.
  • Perform routine infrastructure health checks, preventive maintenance, and operational readiness reviews.
  • Manage rack space allocation, power capacity planning, and infrastructure utilization. ...
Posted
7 days ago

Hong Kong

  • Coordinate incident management, root cause analysis, and issue resolution.
  • Participate in after-hours production support and critical incident response when required.
  • Design and enhance platform observability, monitoring, and alerting capabilities. ...
Posted
7 days ago

Singapore

Posted
7 days ago

Malaysia

  • Ensure design verification and product qualification testing prior to project Gate 5.
  • Perform Process Capability Studies to support Final Tool Approval (FTAR) using CTAR methodology.
  • Establish and maintain quality controls and documentation for production environments. ...
Posted
8 days ago

Singapore

Posted
22 days ago

KL City

  • Job Description:
  • • You understand how Docker and Kubernetes work end to end.
  • • You are an engineer with an interest in service reliability, automation, monitoring, scalability, and high-availability systems. ...
Posted
17 days ago

Singapore

  • 5+ years of Python development
  • ·Strong proficiency in Python programming language
  • ·Hands-on experience with Kubernetes for deploying, managing, and scaling containerized workloads ...
Posted
23 days ago

Singapore

  • Define and manage monitoring, alerting, SLIs, and SLOs for critical production services.
  • Automate operational processes, enhance runbooks, and minimise manual support effort.
  • Diagnose and resolve complex issues across applications, infrastructure, and cloud environments. ...
Posted
10 days ago

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
23 days ago

Singapore

  • Develop and maintain CI/CD pipelines supporting deployment and release activities.
  • Drive observability, monitoring and alerting best practices across the platform.
  • Support incident investigation, root cause analysis and platform troubleshooting. ...
Posted
10 days ago

Singapore

  • Monitor Condition Monitoring program, analyse and propose recommendations for predictive maintenance.
  • Establish programs and initiatives to maximize plant availability and equipment reliability.
  • Degree in Mechanical Engineering ...
Posted
23 days ago

Singapore Refining Company Private Limited (SRC)

Singapore

  • Monitor and review implementation of Preventive Maintenance program to achieve high plant availability.
  • Monitor Condition Monitoring program, analyse and propose recommendations for predictive maintenance.
  • Establish programs and initiatives to maximize plant availability and equipment reliability. ...
Posted
23 days ago

Singapore

  • Design and implement solutions that are secure and compliant by collaborating with dedicated security teams, conducting regular audits, and integrating advanced vulnerability scanning tools.
  • Identify and resolve performance bottlenecks and operational issues, define and track KPIs (e.g., MTTR, system uptime, cost efficiency), and drive ongoing optimisation efforts.
  • Act as a technical advisor for tenants, guiding them on containerization, and best practices for cloud-native deployments, and participating in strategic initiatives to enhance platform scalability and performance. ...
Posted
11 days ago
  • Conduct root cause analysis reviews with site teams and key OEM vendors and develop a corrective action program.
  • Provide systems reliability and maintainability feedback to the Design Engineering teams for future design considerations.
  • Work with Design Engineering and Construction teams to ensure the reliability and maintainability of new and modified installations. ...
Posted
11 days ago

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
23 days ago

KL City

  • • Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.
  • • Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.
  • • Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams. ...
Posted
12 days ago

KL City

  • Incident Management & Root Cause Analysis
  • Participate as a Subject Matter Advisor during production incidents and outages.
  • Provide insights backed by system monitoring, code review, and database analysis. ...
Posted
18 days ago

Singapore

Posted
19 days ago

Singapore

  • Release and change management: build and maintain CI/CD pipelines supporting canary releases and fast rollback; enforce change review and checklists; advance infrastructure-as-code
  • Middleware and database operations: own backup, recovery, tuning, and capacity assessment for regional databases and middleware; work with developers on slow queries and performance issues
  • Security and access: manage regional IAM accounts and access under company policy; support rollout of WAF, gateway, and other security capabilities; enforce least privilege and audit trails for data access ...
Posted
24 days ago

River Valley

Posted
12 days ago
  • Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
  • Gauge the effectiveness and efficiency of existing systems and infrastructure; implement strategies for improving or further leveraging these systems within a geoscience workflow.
  • Collaborate with network and security staff to ensure smooth, secure and reliable operation of application software and systems. ...
Posted
24 days ago