60 Reliability Jobs in Selangor - September 2026 - High Salaries

Showing 60 jobs results for "reliability" in Selangor
Never miss any updates for Reliability jobs in Selangor
  • Regularly deploy product updates as required to keep the platform vulnerability-free.
  • Work with open-source technologies, CI/CD, SCM tools as necessary, and source control such as Bitbucket, implement organization containers (e.g., Docker and Kubernetes). Stay current with industry trends and propose new ways for the business to improve.
  • Take accountability in considering business and regulatory compliance risks and take appropriate steps to mitigate the risks. ...
Posted
13 days ago
  • Additional Information:
Posted
20 days ago
Posted
20 days ago
  • You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
  • Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability.
  • You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders. ...
Posted
20 days ago
  • You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
  • Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability.
  • You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders. ...
Posted
20 days ago
  • Optimise alerting frameworks to improve signal quality and reduce noise
  • Ensure telemetry pipelines are scalable, reliable and cost-efficient
  • Partner with SRE teams to improve visibility into error budgets and performance trends ...
Posted
18 hours ago
  • Monitor and analyze equipment reliability, availability, and maintenance performance, and track and report reliability KPIs such as MTBF, MTTR, PM compliance, and equipment availability.
  • Lead troubleshooting and Root Cause Analysis (RCA) for equipment failures and recurring issues.
  • Develop and optimize Preventive Maintenance (PM), Condition-Based Maintenance (CBM), and Predictive Maintenance (PdM) programs. ...
Posted
14 hours ago
  • You participate in a 24/7 on-call rotation and drive improvements using SRE practices.
  • You actively participate in toil elimination, observability and monitoring improvements, knowledge management, error budget compliance, deployment designs and testing.
  • Bachelor’s degree and/or equivalent experience in Information Technology, Computer Science or Business Management. ...
Posted
13 days ago
  • Improve CI/CD pipeline reliability and deployment safeguards
  • Integrate observability tooling programmatically to improve detection and recovery
  • Standardise automation frameworks across platforms ...
Posted
16 hours ago
  • Conduct daily health checks and monitor VCF infrastructure metrics via vROPS, vCenter, Dynatrace to ensure optimal workload performance and timely issue resolution.
  • Analyze VCF components and perform NVA security remediation to maintain compliance across vSphere, NSX, vSAN, and other VCF elements.
  • Maintains awareness of industry trends on regulatory MAS (SG) and BNC (MY) compliance, emerging threats and technologies to understand the risk and better safeguard the company. Experience with HPSA scanning tools is a plus. ...
Posted
a month ago
  • Conduct root cause analysis reviews with site teams and key OEM vendors and develop a corrective action program.
  • Provide systems reliability and maintainability feedback to the Design Engineering teams for future design considerations.
  • Work with Design Engineering and Construction teams to ensure the reliability and maintainability of new and modified installations. ...
Posted
18 days ago
  • Provide L2/L3 production support, including troubleshooting complex platform, infrastructure, application and networking issues.
  • Participate in a 24/7 on-call rotation and respond to critical production alerts and incidents.
  • Lead or support major incident management, including troubleshooting, vendor coordination, immediate remediation, root cause analysis and long-term corrective actions. ...
Posted
13 days ago
  • Conduct root cause analysis reviews with site teams and key OEM vendors and develop a corrective action program.
  • Provide systems reliability and maintainability feedback to the Design Engineering teams for future design considerations.
  • Work with Design Engineering and Construction teams to ensure the reliability and maintainability of new and modified installations. ...
Posted
19 days ago
  • Optimise alerting frameworks to improve signal quality and reduce noise
  • Ensure telemetry pipelines are scalable, reliable and cost-efficient
  • Partner with SRE teams to improve visibility into error budgets and performance trends ...
Posted
21 days ago
  • Key Responsibilities:
  • • Maintaining the stability, reliability, and efficiency of GEL’s internal container platform and its supporting infrastructure. Responsible for resource provisioning and management, responding to platform and application outages, capacity planning, monitoring, and driving reliability enhancements.
  • • You will continuously evaluate platform’s technical architecture to ensure it scales effectively with evolving application demands, including proactively identifying and resolving reliability issues, analyzing product dependencies, pinpointing performance bottlenecks, and implementing optimization strategies to enhance platform availability and cost efficiency. ...
Posted
a month ago
  • Key Responsibilities:
  • • Maintaining the stability, reliability, and efficiency of GEL’s internal container platform and its supporting infrastructure. Responsible for resource provisioning and management, responding to platform and application outage and monitoring.
  • • This includes proactively identifying and resolving reliability issues, analysing product dependencies, pinpointing performance bottlenecks, and implementing optimization strategies to enhance platform availability and cost efficiency. ...
Posted
a month ago
  • Improve CI/CD pipeline reliability and deployment safeguards
  • Integrate observability tooling programmatically to improve detection and recovery
  • Standardise automation frameworks across platforms ...
Posted
21 days ago

Linergy Power Sdn Bhd

  • Hands-on experience in DevOps / AIOps, skilled in Shell & Python scripting.
  • Solid knowledge of public cloud & private cloud, with cloud project delivery experience.
  • Experience with APM & BCM tools (AppDynamic, Dynatrace); able to troubleshoot applications and optimize performance. Familiar with Zabbix, Grafana, Prometheus monitoring stack. ...
Posted
a day ago
  • Drive technical architecture and engineering decisions across APIs, databases, integrations, infrastructure and customer-facing systems.
  • Set engineering standards around code quality, system design, testing, observability, security and production reliability.
  • Stay close to the technical work through architecture reviews, code reviews, debugging and critical engineering decisions. ...
Posted
5 days ago

Virtuosity Solutions Sdn. Bhd.

  • · Fresh graduates with great attitude and personal qualities are encouraged to apply.
  • · Internship for Final Year Students will be considered for exceptional cases.
  • Basic to intermediate understanding of Software Engineering. ...
Posted
21 days ago
  • Integration Services – MWS/IS support and end-to-end message flow troubleshooting.
  • PostgreSQL – Database administration, query tuning and performance monitoring.
  • Cross-functional Collaboration – Supporting Development, QA and Platform teams to maintain operational excellence. ...
Posted
a month ago

Restoran Mango Chutney

  • Build a reputation for consistent, well-executed dishes that open pathways to head chef or multi-outlet roles.
  • Ready to cook up something meaningful? Join our small, busy kitchen and bring flavour to every plate, working with us at Restoran Mango Chutney where we serve approachable Malaysian and fusion dishes that keep our regulars smiling.
  • As a Builder for the kitchen, you will help design menus, set kitchen standards, and scale our food operations so quality stays high as we grow our footprint. ...
Posted
21 days ago
Posted
13 days ago
Posted
21 days ago

Bandar Puteri 12

  • Mentor engineers through code reviews, pair programming, and documentation to raise team standards and reduce defects. Set coding conventions, run knowledge-sharing sessions, and help junior engineers take ownership of operational tasks.
  • Build a visible portfolio of production systems by driving deployments, monitoring, and structured postmortems that show operational thinking. Own service-level metrics, alerting, and incident follow-up, and present outcomes to product and client stakeholders.
  • Job Summary ...
Posted
5 days ago

Restoran Mango Chutney

  • Build a reputation for consistent, well-executed dishes that open pathways to head chef or multi-outlet roles.
  • Ready to cook up something meaningful? Join our small, busy kitchen and bring flavour to every plate, working with us at Restoran Mango Chutney where we serve approachable Malaysian and fusion dishes that keep our regulars smiling.
  • As a Builder for the kitchen, you will help design menus, set kitchen standards, and scale our food operations so quality stays high as we grow our footprint. ...
Posted
a month ago
Posted
a month ago

MARVEL CAPITAL HOLDINGS LIMITED

  • Strengthen your process skills by leading automation and ERP improvements that reduce close time and improve accuracy.
  • Ready to keep financial operations accurate, timely and useful for decision makers? By working with us at MARVEL CAPITAL HOLDINGS LIMITED, you will help deliver clear financial reports and steady month-end closes that our clients and teams rely on.
  • As the Senior Account backbone for our finance function, you will prepare and review financial statements and lead month-end and year-end closings. You will oversee reconciliations, accruals and journal entries, support budgeting, forecasting and variance analysis, and ensure tax and audit compliance while liaising with external auditors and tax agents. You will mentor junior colleagues and drive process improvements and automation to make the close faster and more reliable. ...
Posted
a month ago
  • Full Stack EngineeringFrontend: React, Next.js, TypeScript, modern component architectures, state management, real-time and streaming AI interfaces, agent activity and execution interfaces, data visualization.Backend: Node.js, TypeScript, Python, REST APIs, GraphQL, WebSockets and streaming, event-driven architectures, background workers, job queues, distributed systems, authentication and authorization.
  • Distributed SystemsDesign systems that reliably execute thousands or millions of AI and data-processing tasks. Kubernetes, Docker, Cloud Run and serverless, message queues, Redis, Kafka or equivalent, distributed job processing, concurrency management, rate limiting, retries, idempotency, fault tolerance, observability. You know how to build systems that stay reliable when agents fail, APIs time out, models hallucinate, or downstream services go away.
  • Data & Learning InfrastructureBuild the infrastructure agents need to learn from historical executions. PostgreSQL, BigQuery or equivalent data warehouses, ClickHouse or analytical databases, vector databases, embeddings, retrieval systems, event logs, feature stores, analytics pipelines, data ingestion. ...
Posted
19 days ago