73 Reliability Jobs in Federal Territory - September 2026 - High Salaries

Showing 73 jobs results for "reliability" in Federal Territory
Never miss any updates for Reliability jobs in Federal Territory

KL City

  • Incident Management: Respond to and manage incidents to minimize downtime and resolve issues quickly, including on-call support.
  • System Performance: Measure, analyze, and tune system performance to ensure efficiency and stability.
  • Infrastructure Management: Provision and manage cloud infrastructure, sometimes using Infrastructure as Code (IaC), and assist in platform management and capacity planning. ...
Posted
19 hours ago

KL City

  • • Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.
  • • Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.
  • • Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams. ...
Posted
4 days ago

KL City

  • Database Design & Recommendations: Review database schema and query design, and propose recommendations to improve indexing, query efficiency, and overall database performance.
  • Upgrades & Migrations: Plan and execute database version upgrades, patching, and migrations (including cross-engine or cross-environment migrations) with minimal downtime and clear rollback plans.
  • Backup & Restoration: Own backup strategy, retention, and restoration testing across managed database services, ensuring recovery objectives (RPO/RTO) are consistently met. ...
Posted
9 days ago

KL City

  • Incident Management: Respond to and manage incidents to minimize downtime and resolve issues quickly, including on-call support.
  • System Performance: Measure, analyze, and tune system performance to ensure efficiency and stability.
  • Infrastructure Management: Provision and manage cloud infrastructure, sometimes using Infrastructure as Code (IaC), and assist in platform management and capacity planning. ...
Posted
14 days ago

KL City

  • • Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.
  • • Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.
  • • Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams. ...
Posted
3 days ago
  • We are seeking a highly skilled Site Reliability Engineer (SRE) with 8+years of experience. The ideal candidate will possess a deep understanding of SRE principles and practices, ensuring the reliability, availability, and performance of our systems. You will work closely with development and operations teams to implement best practices in system reliability and automation.
  • Responsibilities:
  • Design, implement, and maintain scalable and reliable systems and services. ...
Posted
3 days ago

KL City

  • Partner with engineering teams to design reliable and scalable architectures for new and existing applications and services.
  • Improve observability across applications and infrastructure through effective monitoring, alerting, logging, tracing, and automation.
  • Participate in the SRE on-call rotation, lead response to critical production incidents, perform root cause analysis, and drive corrective and preventive improvements. ...
Posted
21 hours ago

KL City

Posted
a day ago

KL City

  • • Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.
  • • Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.
  • • Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams. ...
Posted
a day ago

KL City

  • Driving continuous improvement initiatives to strengthen reliability, resilience, and performance across the portfolio.
  • Bachelor’s degree in mechanical/ electrical engineering, Building Services Engineering or related discipline.
  • Candidates with SCEM, GMAP, CDCS or other related professional certificates are preferred. ...
Posted
4 days ago

KL City

  •  Understand the game architecture, analyze, evaluate and respond to potential risks, such as: Hidden troubles , performance bottlenecks,
  •  Responsible for daily communication and coordination between various teams
  • Job Requirements & Qualifications: ...
Posted
17 days ago
  • We are seeking a highly skilled Site Reliability Engineer (SRE) with 10 -15 years of experience to join our dynamic team in Petaling Jaya. The ideal candidate will possess a deep understanding of SRE principles and practices, ensuring the reliability, availability, and performance of our systems. You will work closely with development and operations teams to implement best practices in system reliability and automation.
  • Responsibilities:
  • Design, implement, and maintain scalable and reliable systems and services. ...
Posted
9 days ago

KL City

  • Analyze production issues, identify root causes, and implement long-term reliability improvements through automation, monitoring, and architectural enhancements.
  • Work collaboratively with other team members, provide technical leadership and guidance to a team of up to 10 SRE engineers, driving engineering excellence, reliability, and operational best practices.
  • Organize an efficient handover through high quality documentation and training. ...
Posted
9 days ago

KL City

  • Conduct in‑depth analysis of system deficiencies, pinpoint system bottlenecks and optimization opportunities, and formulate actionable solutions to enhance system stability and enable cost‑effective, highly‑available system operations.
  • Perform 7×24‑hour On‑call duties to respond, track and resolve online incidents in a timely manner for continuous business stability.
  • Design and build automated platforms and services to improve O&M and delivery efficiency and reduce repetitive manual work. ...
Posted
9 days ago

KL City

  • Lead incident response activities, coordinate cross-functional resolution efforts, and drive continuous improvement through post-incident reviews and preventive actions.
  • Design, implement, and optimize cloud infrastructure configurations across Alicloud and AWS that support scalability, security, and operational stability.
  • Collaborate with development and product teams to improve application reliability, deployment processes, and service performance throughout the software lifecycle. ...
Posted
20 days ago

KL City

  • Help consolidate migration feedback and support issue resolution.
  • Collect and organize project or operational data from agreed sources.
  • Assist in tracking project progress, migration metrics, and operational KPIs ...
Posted
11 hours ago

KL City

  • Drive the development and execution of cutting-edge equipment strategies and predictive maintenance programs to maximize reliability and minimize downtime.
  • Lead advanced reliability investigations using sophisticated techniques like Fault Tree Analysis and Root-Cause Failure Analysis to swiftly resolve critical issues.
  • Engineer and optimize annual maintenance plans, leveraging preventive and predictive tasks to ensure peak equipment performance. ...
Posted
5 days ago

KL City

  • Strong experience in site reliability engineering, infrastructure engineering or a similar role.
  • Strong knowledge on network and protocols, network security and cloud networking
  • Proven strong record of cloud cost optimisation ...
Posted
5 days ago

KL City

  • Build and enhance our observability platform, enabling real-time monitoring of our golden signals (uptime, latency, saturation, error rate)
  • Develop automation solutions for incident response, disaster recovery, and business continuity
  • Drive our DevSecOps platform to enable safe, rapid deployments through CI/CD, GitOps, and self-service capabilities ...
Posted
16 days ago

KL City

  • Analyze production issues, identify root causes, and implement long-term reliability improvements through automation, monitoring, and architectural enhancements.
  • Work collaboratively with other team members and provide guidance to more junior team members.
  • Organize an efficient handover through high quality documentation and training. ...
Posted
16 days ago

KL City

  • Incident Management: Respond to and manage incidents to minimize downtime and resolve issues quickly, including on-call support.
  • System Performance: Measure, analyze, and tune system performance to ensure efficiency and stability.
  • Infrastructure Management: Provision and manage cloud infrastructure, sometimes using Infrastructure as Code (IaC), and assist in platform management and capacity planning. ...
Posted
23 days ago

KL City

  • Build and enhance CI/CD pipelines using GitLab to enable secure, reliable, and automated software delivery.
  • Develop automation solutions to eliminate repetitive operational tasks using scripting and APIs.
  • Manage and optimize Cloudflare services including DNS, WAF, CDN, Load Balancing, Zero Trust, and security controls. ...
Posted
18 days ago

KL City

  • Incident Management & Root Cause Analysis
  • Participate as a Subject Matter Advisor during production incidents and outages.
  • Provide insights backed by system monitoring, code review, and database analysis. ...
Posted
21 days ago

KL City

  • Review team members' deliverables by evaluating adherence to best practices and standards to ensure high-quality software delivery and guide juniors.
  • General Responsibilities: All employees are required to ensure adherence to the compliance of company policies, industry regulations and legal requirements. All employees are expected to assist with tasks, projects, and other duties related to the role, as and when deemed necessary.
  • 2+ years of experience in site reliability engineering or software engineering ...
Posted
16 days ago

KL City

  • Experience: 3+ years to 7 years
  • Site Reliability Engineer (SRE) is an IT professional who applies software engineering methods to IT operations to keep systems reliable, fast, and available.
  • An SRE typically: ...
Posted
18 days ago

KL City

  • Help consolidate migration feedback and support issue resolution.
  • Collect and organize project or operational data from agreed sources.
  • Assist in tracking project progress, migration metrics, and operational KPIs ...
Posted
20 days ago

KL City

  • Experience working within large-scale enterprise environments
  • Bachelor's Degree in Computer Science, Engineering, Information Technology, or equivalent practical experience
  • Contribute to the enhancement and continuous improvement of the Managed Patching Service (MPS) ...
Posted
5 days ago

KL City

  • Manage error budgets aligned with business-critical retail operations
  • Ensure high availability of transaction processing systems (payments, receipts, inventory sync)
  • Design systems resilient to network instability in retail stores ...
Posted
25 days ago

KL City

  • Assist with validation activities and issue tracking.
  • Help consolidate migration feedback and support issue resolution.
  • Collect and organize project or operational data from agreed sources. ...
Posted
21 days ago

KL City

Posted
24 days ago