Why PayNet / Why Now
- Lead the operations of Malaysia's national payment infrastructure, ensuring the availability and stability of systems relied upon by millions every day.
- Build and mature an Application Support function that enables highly resilient, secure, and always-on payment services.
- Drive operational excellence across critical payment platforms through automation, governance, and continuous improvement.
- Influence technology strategy and operational standards that directly support Malaysia's growing digital economy.
TL; DR
- Own the end-to-end Application Support function for PayNet's business-critical platforms.
- Lead teams responsible for production support, incident management, problem management, change management, and service reliability.
- Partner closely with Engineering, Infrastructure, Cyber Security, Product, and Business teams to deliver stable, secure, and high-performing production services.
- Drive operational excellence through automation, governance, service improvements, and people leadership.
Why This Role Matters
- Ensure PayNet's mission-critical applications remain highly available, resilient, and performant to support Malaysia's payment ecosystem.
- Lead the response and recovery of major production incidents while driving permanent corrective actions to reduce operational risk.
- Establish operational governance, service management processes, and support standards that improve reliability and customer confidence.
- Develop a high-performing Application Support organisation with strong technical capability, ownership, and continuous improvement culture.
- Act as the operational bridge between Technology and Business, ensuring production services support evolving business priorities and regulatory expectations.
What You Will Actually Do
- Lead Application Support Operations
Own the day-to-day operation of all production applications, ensuring service availability, system health, operational stability, and compliance with agreed service levels.
- Drive Incident, Problem & Major Incident Management
Lead the organisation's response to production incidents, establish clear escalation paths, conduct root cause analysis, and ensure long-term preventive actions are implemented.
- Govern Production Change & Release Management
Oversee production deployments, infrastructure changes, configuration management, and release activities to minimise operational risk while enabling timely delivery.
- Build Operational Excellence
Define and improve support processes, monitoring capabilities, automation initiatives, operational KPIs, dashboards, documentation, and service governance frameworks.
- Lead People & Cross-Functional Collaboration
Develop and mentor Application Support leaders and engineers while partnering with Development, Infrastructure, Security, Architecture, Vendors, and Business stakeholders to continuously improve production services.
Example of This Role in Practice
- Lead a Severity 1 production incident affecting national payment services, coordinating multiple technology teams until full service restoration while providing timely executive and stakeholder communications.
- Identify recurring production issues through trend analysis and establish permanent fixes that significantly reduce incident volumes and improve platform stability.
- Introduce monitoring automation, operational dashboards, and predictive alerting to improve service visibility and reduce manual operational effort.
- Lead production readiness reviews for new platform launches, ensuring operational documentation, monitoring, support procedures, disaster recovery, and service acceptance criteria are fully completed.
- Coach Application Support Managers and Engineers to strengthen technical capability, improve ownership, and foster a culture of continuous operational improvement.
Required
What Will Help You Succeed
- Leadership in Enterprise Application Support
Proven experience leading Application Support or Production Support teams responsible for business-critical enterprise platforms operating under stringent availability and service level requirements.
- Incident, Problem & Service Management
Deep expertise in ITIL Service Management practices including Incident, Problem, Change, Release, Knowledge, and Major Incident Management with demonstrated experience improving operational maturity.
Strong understanding of enterprise application architecture, APIs, middleware, databases, networking, infrastructure, cloud platforms, and system integrations, enabling effective leadership during complex production issues.
- Stakeholder & Vendor Management
Ability to engage effectively with senior business stakeholders, technology leaders, regulators, vendors, and external partners while managing expectations during both operational and strategic initiatives.
- Operational Excellence & Continuous Improvement
Demonstrated ability to define KPIs, drive automation, improve operational processes, optimise support organisations, and build high-performing teams focused on reliability, efficiency, and customer outcomes.
Good to Have
- Financial Services / Payments Industry Experience
Experience supporting payment systems, financial platforms, or other highly regulated, mission-critical environments with demanding availability requirements.
- Cloud & Modern Platform Operations
Experience supporting cloud-native applications, container platforms, Kubernetes, microservices, DevOps environments, and Site Reliability Engineering (SRE) practices.
- Monitoring & Observability
Hands-on exposure to enterprise monitoring, logging, and observability platforms such as Grafana, Prometheus, ELK, Splunk, AppDynamics, Dynatrace, or similar technologies.
- Security, Audit & Regulatory Compliance
Understanding of operational controls, cybersecurity practices, disaster recovery, business continuity, audit requirements, and regulatory compliance within enterprise technology environments.
- Automation & Digital Operations
Experience driving automation through scripting, orchestration, self-healing capabilities, AI-assisted operations (AIOps), or operational tooling to improve service reliability and reduce manual intervention.