Production Engineering Lead (.NET)
Location
Kuala Lumpur / Hybrid
Position Summary
We are seeking an experienced Production Engineering Lead to own the operational stability, reliability, and performance of our mission-critical digital platforms and enterprise applications built on the Microsoft technology stack.
This role sits at the intersection of Software Engineering, Site Reliability Engineering (SRE), DevOps, and Application Production Support. The successful candidate will lead production engineering practices across a portfolio of .NET applications, driving operational excellence, incident reduction, observability, automation, and release governance.
The ideal candidate combines strong production support leadership with deep technical knowledge of C#, ASP.NET Core, Azure, APIs, and distributed architectures, enabling effective partnership with development teams and rapid resolution of complex production issues.
Key Responsibilities
Production Stability & Reliability
- Own the availability, reliability, and operational health of business-critical applications.
- Lead major incident management and coordinate cross-functional teams during Sev-1 and Sev-2 incidents.
- Drive root cause analysis (RCA) and ensure preventive actions are implemented.
- Establish proactive monitoring and alerting to improve service availability.
- Monitor platform performance and identify opportunities for optimization.
.NET Platform Ownership
- Serve as the production engineering lead for applications built using:
- C#
- ASP.NET Core
- REST APIs
- Microservices
- Azure-hosted services
- Partner closely with engineering teams during troubleshooting and production investigations.
- Review application architecture and identify operational risks.
- Support platform modernization and cloud adoption initiatives.
DevOps & Release Governance
- Collaborate with development teams on deployment planning and production readiness reviews.
- Govern release processes, rollback strategies, and change controls.
- Drive CI/CD best practices through Azure DevOps and GitHub workflows.
- Promote Infrastructure-as-Code and deployment automation.
Site Reliability Engineering (SRE)
- Establish SRE practices and reliability metrics.
- Define and monitor:
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Operational KPIs
- Improve:
- Availability
- Incident response
- Mean Time to Recovery (MTTR)
- Change success rate
Observability & Monitoring
- Build and maintain observability standards across applications.
- Utilize tools such as:
- Azure Monitor
- Application Insights
- Grafana
- Splunk
- Dynatrace
- Drive proactive issue detection through dashboards and automated alerting.
Leadership & Stakeholder Management
- Lead and mentor Production Support / Production Engineering teams.
- Act as the escalation point for business-critical incidents.
- Collaborate with:
- Software Engineering teams
- Product Owners
- Architects
- Infrastructure teams
- Security teams
- Communicate operational risks and service performance to senior stakeholders.
Required Experience
Must Have
- 10+ years of IT experience.
- 3+ years in Production Engineering, SRE, Application Support Lead, or Production Support Lead roles.
- Strong hands-on experience supporting applications built using:
- C#
- ASP.NET Core
- Web APIs
- Experience working within Azure cloud environments.
- Experience managing critical production incidents and driving RCA.
- Experience supporting distributed systems and microservices architectures.
- Strong understanding of SDLC, deployment processes, and release management.
- Experience with SQL Server and application performance troubleshooting.
- Strong stakeholder management and communication skills.
Technical Skills
Application Technologies
- C#
- ASP.NET Core
- REST APIs
- Microservices
- Authentication / Authorization (OAuth, OIDC, JWT)
Cloud & DevOps
- Microsoft Azure
- Azure DevOps
- GitHub Actions
- CI/CD Pipelines
- Docker
- Kubernetes (preferred)
Monitoring & Observability
- Azure Monitor
- Application Insights
- Grafana
- Dynatrace
- Splunk
Databases
- SQL Server
- PostgreSQL (preferred)
Preferred Qualifications
- Experience within Insurance, Banking, or Financial Services.
- Experience supporting customer-facing digital platforms or portals.
- Microsoft Azure certifications (AZ-104, AZ-204, AZ-305).
- Exposure to Site Reliability Engineering practices.
- Experience implementing production automation and operational tooling.