jobs in CoreWeave

CoreWeave Hiring! Full Time Production Engineer – Team Lead in - Ricebowl

Production Engineer – Team Lead

CoreWeave

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

CoreWeave is The Essential Cloud for AI. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at *************

What You'll Do

CoreWeave’s Production Engineering organisation sits at the heart of our cloud infrastructure stability, availability, and resilience. We build, maintain, and safeguard the operational platform that powers massive multi-tenant GPU workloads for the world's leading AI labs and enterprises.

About The Role

As a Production Engineer – Team Lead, you will serve as a senior generalist providing strategic technical direction, operational continuity, and incident response leadership across our cloud platform. You will act as a primary Incident Commander during major service disruptions, coordinating cross-functional engineering teams, guiding root cause analysis (RCA) efforts, and refining post-incident review (PIR) frameworks. In this role, you will define and track Service Level Objectives (SLOs), implement automation to reduce MTTD/MTTR, and drive deep-dive investigations across complex distributed systems. Additionally, you will mentor Production Engineers (I/II), build scalable incident runbooks, and foster an engineering culture centered on continuous improvement and platform reliability.

Who You Are

  • 4+ years of professional experience in production engineering, cloud operations, Site Reliability Engineering (SRE), or major incident management.
  • Deep technical expertise with cloud platform architectures, containerisation, and Kubernetes-based infrastructure.
  • Strong familiarity with incident management frameworks, ITIL methodologies, and SRE best practices.
  • Proficiency with observability ecosystems, alerting tools (Prometheus, Grafana), and telemetry principles.
  • Hands-on scripting, automation, and configuration management experience (Python, Bash, Terraform).
  • Proven experience leading incident command under high-pressure scenarios and communicating complex technical status to both technical and executive stakeholders.
  • Track record of mentoring technical engineers and up-skilling team capabilities in operational best practices.

Preferred

  • Direct experience operating in an Incident Commander role for major, high-priority service restorations at scale.
  • Advanced knowledge of Kubernetes internals, distributed systems architecture, and container runtimes.
  • Hands-on experience developing self-healing infrastructure, runbook automation, and progressive change management frameworks.

Wondering if you're a good fit?

We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams—even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk.

  • You love to step up during critical production incidents, lead technical teams through ambiguity, and drive swift, blameless recovery.
  • You're curious about dissecting complex distributed failures, finding "unknown unknowns" across system boundaries, and automating toil away.
  • You're an expert in setting reliability standards, establishing SLO frameworks, and mentoring SREs to build resilient cloud platforms.

Why CoreWeave?

About

At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the organisation's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

The base salary range for this role is SGD 196,000 to SGD 262,000 annually. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More