700+ Reliability Jobs - October 2026 - High Salaries

Showing 709 jobs results for "reliability"
Never miss any updates for Reliability jobs
  • Analyze quality metrics, statistical data, and performance trends to drive evidence-based decision-making
  • Lead cross-functional teams in process improvement initiatives using methodologies such as Six Sigma and Lean
  • Establish and maintain quality standards, specifications, and acceptance criteria aligned with industry best practices ...
Posted
13 days ago
  • Regularly deploy product updates as required to keep the platform vulnerability-free.
  • Work with open-source technologies, CI/CD, SCM tools as necessary, and source control such as Bitbucket, implement organization containers (e.g., Docker and Kubernetes). Stay current with industry trends and propose new ways for the business to improve.
  • Take accountability in considering business and regulatory compliance risks and take appropriate steps to mitigate the risks. ...
Posted
24 days ago

Beijing

  • Supporting different teams to make long-term infrastructure decisions, provide suggestions for infrastructure optimizations.
  • Drive operational excellence practice and help improve reliability, stability and scalability challenges with engineering teams.
  • Lead junior engineers to complete projects and help them grow. ...
Posted
24 days ago

Singapore

  • ·      Own production reliability for AI-based media pipelines
  • ·      Read, debug, and patch application code when issues surface — not just restart services or check dashboards
  • ·      Step into pipeline engineering work- Ad /promo replacement, watermarking, inpainting — when core developers are stretched ...
Posted
24 days ago

Singapore

  • Support cloud security operations, including cloud security alert management and compliance auditing.
  • 3+ years of DevOps or SRE experience; experience with AIOps or observability platform development is a plus.
  • Proficient in Python; familiar with at least one of Go or Java. Full-stack capability (React/Vue frontend + backend API) is a plus. ...
Posted
24 days ago

Centre For Strategic Infocomm Technologies (CSIT)

Singapore

  • Provide network consultation to project teams for system deployment, working with various stakeholders to resolve technical problems and optimise network performance
  • Background in Computer Science, Computer/Electrical Engineering, Information Technology or related fields
  • 2-5 years of relevant working experience in network design, implementation and/or operations preferred ...
Posted
a month ago

Singapore

  • Implement SRE practices including monitoring, observability, alerting, incident management, and automation.
  • Support Platform Engineering initiatives to improve deployment, scalability, resilience, and operational efficiency.
  • Automate repetitive operational processes and improve system reliability. ...
Posted
24 days ago
  • We are seeking a highly skilled Site Reliability Engineer (SRE) with 10 -15 years of experience to join our dynamic team in Petaling Jaya. The ideal candidate will possess a deep understanding of SRE principles and practices, ensuring the reliability, availability, and performance of our systems. You will work closely with development and operations teams to implement best practices in system reliability and automation.
  • Responsibilities:
  • Design, implement, and maintain scalable and reliable systems and services. ...
Posted
24 days ago

KL City

  • Analyze production issues, identify root causes, and implement long-term reliability improvements through automation, monitoring, and architectural enhancements.
  • Work collaboratively with other team members, provide technical leadership and guidance to a team of up to 10 SRE engineers, driving engineering excellence, reliability, and operational best practices.
  • Organize an efficient handover through high quality documentation and training. ...
Posted
24 days ago

Singapore

  • Build and maintain production tooling that supports deployment, orchestration, monitoring, and system diagnostics
  • Define and maintain observability, SLI/SLOs, and performance metrics in partnership with product owners
  • Leverage metrics and capacity planning to ensure scalability and uptime ...
Posted
a month ago

Singapore

  • Experience programming in Python or another modern language such as Java, Go, C#, or Scala.
  • Strong SQL skills with experience working with large-scale datasets and data analysis tools (e.g. Pandas, Snowflake, Elasticsearch/Kibana).
  • Exposure to orchestration and data workflow technologies such as Airflow, Dagster, Argo, Spark, or Hive. ...
Posted
24 days ago

KL City

  • Conduct in‑depth analysis of system deficiencies, pinpoint system bottlenecks and optimization opportunities, and formulate actionable solutions to enhance system stability and enable cost‑effective, highly‑available system operations.
  • Perform 7×24‑hour On‑call duties to respond, track and resolve online incidents in a timely manner for continuous business stability.
  • Design and build automated platforms and services to improve O&M and delivery efficiency and reduce repetitive manual work. ...
Posted
24 days ago

Hong Kong

Posted
24 days ago

Singapore

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
  • Abide by Mastercard’s security policies and practices;
  • Ensure the confidentiality and integrity of the information being accessed; ...
Posted
20 days ago

KL City

  • Drive the development and execution of cutting-edge equipment strategies and predictive maintenance programs to maximize reliability and minimize downtime.
  • Lead advanced reliability investigations using sophisticated techniques like Fault Tree Analysis and Root-Cause Failure Analysis to swiftly resolve critical issues.
  • Engineer and optimize annual maintenance plans, leveraging preventive and predictive tasks to ensure peak equipment performance. ...
Posted
20 days ago

Singapore

  • Leverage firm-wide metrics to improve scalability and system performance
  • Collaborate across the technology organization to analyze and troubleshoot complex system problems
  • Work closely with Risk Management and Operational Trading Support teams to coordinate changes and manage incidents ...
Posted
15 days ago

Malaysia

  • Logistics shipment handling IN/Out from laboratory to ensure in time delivery and fulfilment cycle time expectation
  • Reliability engineering support responsibilities may be assigned on an as needed basis
  • Participation in RCCA efforts and assist engineer in reliability activities ...
Posted
4 days ago

Singapore

  • Design and develop automations for our workflow.
  • Capacity and Resource management.
  • Responsible for the full-chain stress test to enhance the performance and remove redundancy of applications. ...
Posted
5 days ago

Singapore

  • Enhance observability across metrics, logs, and distributed tracing
  • Partner with engineering teams to improve scalability, reliability, and security
  • Support incident response and continuous platform improvement initiatives ...
Posted
5 days ago

Malaysia

  • Participate in 24/7 on-call rotations, including scheduled shifts and holidays. Conduct root-cause analysis and lead blameless post-mortems to prevent recurrence.
  • Implement monitoring tools (SLIs/SLOs/SLAs), automated alerting, and metrics to track system health and performance.
  • Implement and maintain security best practices and ensure systems meet regulatory requirements. ...
Posted
5 days ago

Singapore

  • Insurance Coverage
  • Entitled to Yearly Bonus & Performance Bonus
  • Focused on maintaining high availability and stability of application systems, ensuring compliance with privacy and data protection laws, managing changes and releases, and driving operational automation and infrastructure optimization. ...
Posted
6 days ago

KL City

  • Strong experience in site reliability engineering, infrastructure engineering or a similar role.
  • Strong knowledge on network and protocols, network security and cloud networking
  • Proven strong record of cloud cost optimisation ...
Posted
20 days ago

Singapore

  • Write well-tested, documented, and maintainable code with proper versioning, release processes, and code review practices
  • Architect, develop, and refactor Ansible roles and playbooks across a large-scale inventory spanning 30+ datacenters, 80+ group variable files, and 40+ roles
  • Design reusable, composable Ansible role patterns that scale cleanly as the DC footprint grows -- new DCs should be deployable with minimal variable additions ...
Posted
8 days ago

Singapore

Posted
21 days ago

Seri Manjung

  • Experience: At least 10 years of experience in technical leadership roles within maintenance, reliability, or process plant technical management.
  • Analytical Skills: Strong analytical and problem-solving skills, with the ability to identify and evaluate options, assess risks, and make informed decisions aligned with organizational objectives.
  • Communication Skills: Excellent communication and interpersonal skills, with the ability to collaborate effectively with cross-functional teams, stakeholders, and senior management. ...
Posted
a month ago

Singapore

  • Drive observability, monitoring, and SLO discipline
  • Support release engineering, including data migration planning and verification
  • Partner with platform, infra, and research teams to ensure reliability and scalability ...
Posted
16 days ago

Singapore

  • Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch
  • Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity
  • Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts ...
Posted
10 days ago

Singapore

  • Collaborating closely with traders, developers, and infrastructure engineers
  • 1+ years of experience in technical support or engineering in a trading or finance environment
  • Strong Linux systems experience (debugging, scripting, tuning) ...
Posted
11 days ago