200+ Site Reliability Engineering Jobs in Malaysia | Job Vacancies | August 2026 | Ricebowl

显示298个工作的结果 "site reliability engineering"
不要错过任何 Site Reliability Engineering 的新工作机会
支持聊天
Undisclosed

George Town, Pulau Pinang

  • Conduct risk assessments and manage wafer incoming issues at suppliers and on-hold lots with assembly issues on the test floor.
  • Generate and evaluate supplier quality reports, establish baselines, and identify opportunities for continuous improvement.
  • Coordinate with FA engineers to identify assembly failures, determine root causes, and implement corrective actions with assemblers. ...
Posted
a month ago
支持聊天
Undisclosed

Bandar Kuala Lumpur, WP Kuala Lumpur

靠近火车站
  • • Cloud Resources & CDN Management: manage multi-cloud environments (Alibaba, Tencent, Huawei, AWS) with cost and capacity control; lead CDN configuration for Web traffic security, acceleration, and HA
  • • Domain & Certificate Systems: manage all DNS resolution and filings; build a unified SSL certificate issuance, deployment, monitoring, and auto-renewal mechanism to prevent expiry incidents
  • • Cloud Security: lead server and business security-protection strategies (DDoS, brute-force, WAF, intrusion detection); coordinate security-incident response ...
Linux Shell
+3
Posted
a month ago
支持聊天
SGD3,600 - SGD4,000 每月

Singapore, Singapore

  • • Those with Diploma in Electrical Engineering are also welcome to apply
  • • Possess 2-3 years related hands-on working experiences in electrical switchboards servicing/testing would be an advantage
  • • Good communication, interpersonal and time management skills. ...
Electrical Maintenance Electrical Engineering
+1
Posted
2 days ago
Undisclosed

Bandar Kuala Lumpur, WP Kuala Lumpur

靠近火车站
  • Drive infrastructure automation using Terraform, CI/CD, and GitOps practices.
  • Improve observability across teams by standardizing monitoring, tracing, and alerting.
  • Collaborate on incident response and postmortems to reduce MTTR and build resilience. ...

最后机会申请此工作。

Posted
a month ago
Undisclosed

Bandar Kuala Lumpur, WP Kuala Lumpur

靠近火车站
  • Drive infrastructure automationusingTerraform, CI/CD, and GitOps practices
  • Improve observability across teams by standardizing monitoring, tracing, and alerting
  • Collaborate on incident response and postmortems to reduce MTTR and build resilience ...

最后机会申请此工作。

Posted
a month ago
Undisclosed
靠近火车站
  • Participate actively in site meetings, factory inspections, and acceptance testing to resolve design and construction challenges while ensuring technical compliance and quality assurance.
  • Oversee the preparation and submission of Operation and Maintenance Manuals, guaranteeing user-friendly and thorough documentation for project handover.
  • Manage construction site supervision workflows, including the review and approval of shop drawings, materials submissions, As-Built drawings, and technical queries to ensure conformity with approved designs and regulation. ...
Electrical Design Project Management
+4

最后机会申请此工作。

Posted
2 months ago
Undisclosed
  • Support the verification of site levels, alignments, and construction layouts under the guidance of the Land Surveyor. •
  • Assist with survey activities across multiple work fronts to ensure timely completion of project survey requirements. •
  • Maintain basic survey records, field notes, and activity logs to support project documentation. • ...

最后机会申请此工作。

Posted
a month ago
Undisclosed

Singapore

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
  • Abide by Mastercard’s security policies and practices;
  • Ensure the confidentiality and integrity of the information being accessed; ...
Posted
9 days ago
SGD9,500 - SGD9,500 每月

Singapore

  • Leverage automation and AI technologies to enhance proactive issue detection, enable self-healing capabilities, reducing Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM).
  • Develop testing and validation plans for new environment builds, disaster recovery exercises and post-maintenance activities to certify environment readiness before customer traffic is routed to it.
  • Champion continuous learning, development, and knowledge sharing across networking and other infrastructure disciplines to strengthen multi-disciplinary SRE team capabilities. Lead training initiatives for team members and Product and Development on networking aspects of the platforms. ...
Posted
15 days ago
Undisclosed

Singapore

  • Manage and optimise observability data pipelines, ensuring scalability, performance, and data fidelity.
  • Develop internal observability tooling, automation utilities, and integrations using Python.
  • Build and maintain RESTful API integrations connecting observability platforms with internal systems and third-party services. ...
Posted
a day ago
Undisclosed

Singapore

  • Working experience with OpenTelemetry (OTel) — instrumentation, collectors, exporters, and signal types (logs, metrics, traces)
  • Experience building or consuming RESTful APIs
  • Familiarity with AWS services, particularly AWS EventBridge for event-driven architecture and integration workflows ...
Posted
6 days ago
Undisclosed

Singapore

  • Manage and optimise observability data pipelines, ensuring scalability, performance, and data fidelity.
  • Develop internal observability tooling, automation utilities, and integrations using Python.
  • Build and maintain RESTful API integrations connecting observability platforms with internal systems and third-party services. ...
Posted
6 days ago
Undisclosed

KL City

  • Organise collaboration within technology teams by creating efficient communication to enhance collaboration and achieve shared goals.
  • Assist on identifying risk in technology infrastructure by conducting proactive risk assessments and develop contingency plan to ensure regulatory and security compliance for technology infrastructure deliverables
  • Execute strategic vision for technology infrastructure improvement by executing the roadmap of initiatives which align with company goals to ensure continuous improvement ...
Posted
13 days ago
SGD13,000 - SGD13,000 每月

Singapore

  • Strong experience with AWS (EC2, EKS, VPC, Route 53, RDS, Aurora, S3, EFS, ALB/NLB) and cloud infrastructure design.
  • Hands-on expertise in Kubernetes, Docker, Helm, Kustomize, Terraform, and CI/CD tools such as Jenkins, GitLab CI/CD, and ArgoCD.
  • Experience with monitoring, observability, and reliability engineering using Prometheus, Grafana, CloudWatch, and SigNoz. ...
Posted
14 days ago
Undisclosed

Hong Kong

  • Develop automation solutions to reduce manual effort, improve reliability, and optimise resource utilisation
  • Leverage data, AI and anomaly detection to proactively identify and resolve system issues
  • Collaborate with cross-functional teams and technical partners to evaluate tools, run PoCs, and implement new capabilities ...
Posted
a month ago
Undisclosed

KL City

  • Lead every Sev1/2 Incident, run the bridge, write RCA within 48H, enforce blameless post-mortems the same week, and ship permanent automated fixes so the same outage never happens twice.
  • Review team members' code scripts by evaluating adherence to better code quality standards to ensure high-quality software delivery.
  • Evolve product Observability. This includes metrics (Prometheus/Tempo), Logs (Loki/Cloudwatch), Traces (Tempo/OpenTelemetry) and proactively updates on the design, and implementation. ...
Posted
22 days ago
Undisclosed

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
2 days ago
Undisclosed

Singapore

  • Support incident management, troubleshooting, root-cause analysis, and continuous improvement.
  • Integrate monitoring and alerting systems with SIEM and Incident Response workflows.
  • Improve automation, deployment processes, and operational efficiency. ...
Posted
3 days ago
Undisclosed

Singapore

  • Experience integrating systems with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
  • Experience defining and implementing runbooks, operational procedures, escalation paths, and production support models.
  • Role – Site Reliability Engineer (SRE) ...
Posted
2 days ago
Undisclosed

KL City

  • Travel allowance
  • 21 vacation days
  • 13th salary ...
Posted
2 days ago
Undisclosed

Singapore

  • Experience integrating systems with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
  • Experience defining and implementing runbooks, operational procedures, escalation paths, and production support models.
  • Technical Skills Required for SRE role
Posted
a day ago
Undisclosed

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
a day ago
Undisclosed
  • · Design containerization, Kubernetes and service mesh implementation plans.
  • Full stack monitoring and logging system settings
  • · Design/develop a full-stack monitoring system for good monitoring of development and production scenarios within the entire cloud resource. ...
Posted
13 hours ago
Undisclosed

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
3 days ago
Undisclosed

Singapore

  • Build and maintain production tooling that supports deployment, orchestration, monitoring, and system diagnostics
  • Define and maintain observability, SLI/SLOs, and performance metrics in partnership with product owners
  • Leverage metrics and capacity planning to ensure scalability and uptime ...
Posted
5 days ago
Undisclosed

KL City

  • Support deployment activities and configuration changes in accordance with established operational procedures.
  • Assist in managing and maintaining Kubernetes environments, including monitoring workloads and investigating cluster-related issues.
  • Utilize Linux commands and Shell scripts to automate repetitive tasks, collect operational data, and improve efficiency. ...
Posted
6 days ago
Undisclosed

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
6 days ago
Undisclosed

Singapore

  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK ...
Posted
7 days ago

TalentVibe Business Consultancy

SGD5,000 - SGD6,000 每月

Singapore

  • Ensure high availability and seamless scalability of the microsegmentation policy engine and synchronization pipelines.
  • Define, maintain, and refine operational runbooks, escalation paths, standard operating procedures (SOPs), and production support models.
  • Drive incident response, triage, root-cause analysis (RCA), and continuous remediation to improve platform resilience. ...
Posted
8 days ago
Undisclosed

Singapore

  • Implement and manage monitoring, logging, metrics, tracing and alerting
  • Support incident management and production troubleshooting
  • Develop and maintain operational runbooks and support procedures ...
Posted
8 days ago