88 Kubernetes Jobs - September 2026 - High Salaries

Showing 88 jobs results for "kubernetes"
Never miss any updates for Kubernetes jobs

Singapore

  • Collaborate with developers, infra, and cloud teams to ensure reliable delivery
  • Hands-on experience in DevOps, CI/CD, or Kubernetes administration
  • Strong scripting skills (Python, Bash, Groovy) ...
Posted
4 days ago

Singapore

  • Amazon EKS, Azure AKS, Google Kubernetes Engine (cloud)
  • Deploy and configure Kubernetes/OpenShift clusters
  • Implement control plane and worker node setup ...
Posted
7 days ago

Singapore

Posted
12 days ago

Singapore

  • User access management
  • Pod/application restarts
  • Log analysis and basic issue diagnosis ...
Posted
20 days ago

Singapore

  • Collaborate with developers, infra, and cloud teams to ensure reliable delivery
  • Hands-on experience in DevOps, CI/CD, or Kubernetes administration
  • Strong scripting skills (Python, Bash, Groovy) ...
Posted
18 days ago

Singapore

  • Build and automate scalable container platforms using Infrastructure as Code (IaC), GitOps, reusable templates, and standardized Dev/UAT/Prod environments.
  • Implement and manage cluster components, including control planes, worker nodes, networking, storage, monitoring, logging, and observability tools such as Prometheus, Grafana, and ELK.
  • Ensure platform reliability through performance tuning, alerting, enterprise monitoring integration, troubleshooting, and 24/7 operational support when required. ...
Posted
10 days ago

Singapore

  • Amazon EKS, Azure AKS, Google Kubernetes Engine (cloud)
  • Deploy and configure Kubernetes/OpenShift clusters
  • Implement control plane and worker node setup ...
Posted
19 days ago

Singapore

  • Build and automate scalable container platforms using Infrastructure as Code (IaC), GitOps, reusable templates, and standardized Dev/UAT/Prod environments.
  • Implement and manage cluster components, including control planes, worker nodes, networking, storage, monitoring, logging, and observability tools such as Prometheus, Grafana, and ELK.
  • Ensure platform reliability through performance tuning, alerting, enterprise monitoring integration, troubleshooting, and 24/7 operational support when required. ...
Posted
12 days ago

Singapore

Posted
12 days ago

Singapore

  • Implement GitOps workflows for cluster and platform configuration.
  • Establish tenant onboarding, access control, quotas, policies, and workload-isolation standards.
  • Monitor platform health and troubleshoot incidents across Kubernetes, Linux, networking, storage, and infrastructure. ...
Posted
12 days ago

Singapore

  • Support application deployments, releases and upgrades and investigate deployment failures.
  • Analyse application and system logs, monitoring alerts and performance metrics to identify root causes.
  • Work with Microsoft Azure environments, with exposure to Azure Kubernetes Service (AKS) highly advantageous. ...
Posted
12 days ago
  • Experience integrating Kubernetes with CI/CD and GitOps processes.
  • Knowledge of monitoring, observability, performance tuning, and troubleshooting.
  • Experience with Rancher for Kubernetes management. ...
Posted
12 days ago

Singapore

  • Stability and security: Build comprehensive K8s cluster monitoring, alerting, logging, and distributed tracing systems; define operations runbooks, change processes, and incident response plans; strengthen cluster security controls, disable high-risk permissions, harden container runtime environments, and ensure infrastructure and business data security.
  • Automated operations and DevOps: Develop operations automation scripts using Shell/Python; integrate Jenkins, GitLab CI, and ArgoCD to build automated release, inspection, and backup systems; implement Infrastructure as Code (IaC) principles to improve efficiency and reduce human error.
  • Incident management and post-mortem optimization: Lead online incident response, conduct root cause analysis, produce post-mortem reports, and continuously optimize cluster architecture, resource allocation, monitoring strategy, and long-term stability assurance mechanisms. ...
Posted
12 days ago

Singapore

  • Develop and maintain documentation for solution configurations, deployment guides, and knowledge-sharing resources.
  • Monitor and optimize infrastructure performance, ensuring high availability and resource efficiency for AI model execution.
  • Act as a subject matter expert, providing guidance on GPU acceleration technologies and their integration in AI workflows. ...
Posted
19 days ago

Singapore

Posted
20 days ago

Singapore

Posted
20 days ago

Singapore

  • Knowledge of computer hardware components and ability to restore faulty servers to working condition.
  • Experience in coding or scripting languages such as bash, windows batch script, powershell, perl, python, etc.- Ability to install, use and configure various Linux Operating Systems, such as Redhat, Ubuntu, CentOS.- Experience in system installation planning, execution, routine
Posted
20 days ago

Singapore

Posted
8 days ago

Singapore

  • Perform capacity planning, performance tuning, backup, recovery, and restoration activities.
  • Kubernetes
  • Docker ...
Posted
8 days ago

Singapore

Posted
22 days ago

Singapore

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • 10+ years of experience in DevOps, Platform Engineering, Infrastructure Engineering, Cloud Engineering, Site Reliability Engineering (SRE), or related technical roles.
  • Strong experience supporting and operating both on-premise and cloud-based environments, including hybrid infrastructure architectures. ...
Posted
12 days ago

Singapore

  • Build topology-aware placement mechanisms that account for GPU locality, NVLink/NVSwitch domains, node topology, NUMA affinity, NIC placement, RDMA paths, network-fabric topology, storage locality, and fault-domain boundaries.
  • Build AI-factory resource-aware scheduling mechanisms that account for cluster capacity, node health, fabric condition, storage performance, GPU availability, maintenance windows, software compatibility, capacity reservations, power limits, thermal conditions, and operational constraints.
  • Implement workload-control capabilities including admission control, queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, co-scheduling, backfilling, workload aging, retry policies, checkpoint-aware scheduling, deferred execution, and failure recovery. ...
Posted
12 days ago

Singapore

  • Build topology-aware placement mechanisms that account for GPU locality, NVLink/NVSwitch domains, node topology, NUMA affinity, NIC placement, RDMA paths, network-fabric topology, storage locality, and fault-domain boundaries.
  • Build AI-factory resource-aware scheduling mechanisms that account for cluster capacity, node health, fabric condition, storage performance, GPU availability, maintenance windows, software compatibility, capacity reservations, power limits, thermal conditions, and operational constraints.
  • Implement workload-control capabilities including admission control, queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, co-scheduling, backfilling, workload aging, retry policies, checkpoint-aware scheduling, deferred execution, and failure recovery. ...
Posted
12 days ago

Singapore

  • Build topology-aware placement mechanisms that account for GPU locality, NVLink/NVSwitch domains, node topology, NUMA affinity, NIC placement, RDMA paths, network-fabric topology, storage locality, and fault-domain boundaries.
  • Build AI-factory resource-aware scheduling mechanisms that account for cluster capacity, node health, fabric condition, storage performance, GPU availability, maintenance windows, software compatibility, capacity reservations, power limits, thermal conditions, and operational constraints.
  • Implement workload-control capabilities including admission control, queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, co-scheduling, backfilling, workload aging, retry policies, checkpoint-aware scheduling, deferred execution, and failure recovery. ...
Posted
12 days ago

Singapore

  • Support CI/CD pipelines and deployment processes across development, testing, and production environments.
  • Perform production incident management, root cause analysis (RCA), and service reliability improvements.
  • Work closely with development, infrastructure, and security teams to ensure platform stability and availability. ...
Posted
12 days ago

Singapore

  • Collaborate with security and networking teams to define and enforce policies across containerized platforms.
  • Build and maintain observability (monitoring, logging, alerting) and operational dashboards.
  • Work closely with development and DevOps teams to provide reliable platform capabilities. ...
Posted
20 days ago

Singapore

  • Investigate customer technical issues and provide timely resolutions
  • Prepare implementation plans, technical documentation, reports and deployment-related materials
  • Conduct product demonstrations, technical discussions and customer training where required ...
Posted
4 days ago

Singapore

Posted
20 days ago

Singapore

  • Analyse logs, identify root causes and reproduce technical issues.
  • Work with engineering teams to resolve product-related issues.
  • Provide clear technical guidance and updates to customers. ...
Posted
24 days ago

Singapore

  • Amazon EKS, Azure AKS, Google Kubernetes Engine (cloud)
  • Deploy and configure Kubernetes/OpenShift clusters
  • Implement control plane and worker node setup ...
Posted
a month ago