About Us
The Senior Cloud Engineer has the deepest technical authority in a major area of Traveloka's cloud platform, owning reliability, architecture, and cost decisions that shape how the wider engineering organization builds and operates. You will drive large-scale infrastructure programs end-to-end, from technical roadmap to execution, and partner closely with the Cloud Platform Engineering team and engineering leaders across the organization to raise the bar on reliability, automation, and operational excellence.
What You'll Do
- Define and drive a roadmap to migrate from non-K8s-based infrastructure to a K8s-based one across public cloud platforms (AWS, GCP, or AliCloud).
- Maintain multi-tenant Kubernetes clusters across public cloud platforms and regions.
- Maintain cloud infrastructure using Terraform.
- Build platform-grade automation and self-service tools to simplify application deployment, incorporating AI-assisted workflows where they genuinely accelerate delivery without compromising verification rigor.
- Own technical decisions (e.g., reliability, scalability, security, and cost efficiency) for a major platform domain, and drive continuous improvement in IaC, observability, and cost efficiency.
- Lead cross-cutting incident responses as a senior technical commander, and owned systemic fixes through to completion.
- Identify systemic weaknesses across the platform (e.g., architectural, operational, or cost-related) and lead the cross-functional programs that fix them for good.
- Communicate technical trade-offs and risk clearly to both technical and non-technical stakeholders to support fast, informed decisions.
Core Capabilities
- At least 8 years in Cloud Engineering, with 5 years deeply focused on production Kubernetes preferred.
- Mastery of Helm, CNI plugins (Cilium/Calico), Service Meshes (Istio), and K8s Operator development.
- Deep knowledge and hands-on expertise of cloud-native services (e.g., AWS EKS, AliCloud ACK, etc.) and multi-region deployment.
- Strong grounding in distributed systems failure modes, capacity planning, and operational architecture.
- Strong scripting and automation skills in Go or Python.
- Extensive Terraform / IaC experience, with a track record of setting standards other teams operate by.
- Practical experience applying AI-assisted tooling to engineering workflows (e.g., automation, triage, incident response), with the judgment to know where it accelerates work versus where it introduces risk.
- Proven ability to lead major incident responses for high-traffic, consumer-facing systems.
What Makes You Stand Out
- You've architected and implemented a migration to a K8s-based infrastructure serving an internet-scale of traffic.
- A Certified Kubernetes Administrator (CKA) or a Certified Kubernetes Application Developer (CKAD) certification is highly desirable.
- Active involvement in Cloud/K8s/DevOps communities, conferences, or open-source contributions.