Work with level 1 support in application knowledge transfer in establishing standard first level recovery process/system health check monitoring/scheduled activities as well as common application user queries.
Collaborate with internal stakeholders like application, business, infrastructure, security, and operations to continuously improve the stability of applications.
Develop support operational documents to summarize common support issues.
...
Act as Technical Product Manager in operations, ensuring all technical components (servers, middleware components, application modules, interfaces, job chains, schedulers, connectivity paths) remain secure, compliant and lifecycle‑current.
Monitor, track and plan technology lifecycle events (EoL/EoS, patch cycles, hardware refresh, OS upgrades, middleware version changes) together with infrastructure and platform teams, ensuring risks are identified early and scheduled into IBSol governance cycles.
Execute Incident, Problem, Change and Release Management according to IBSol service standards, including detailed analysis, task coordination, root‑cause identification and technical approvals.
...
Work with Service Center to ensure average turn around time (TAT) is within agreed KPI thresholds.
Create a weekly report to send to the Info Systems Project Management team to assist them in reviewing incidents to determine what events have occurred that may be operational losses.
Work with Service Center to ensure average turn around time (TAT) is within agreed KPI thresholds.
...
Operational Governance: Maintaining operational governance of AI solutions by managing production stability, performance thresholds, and compliance requirements.
Stakeholder Collaboration: Collaborating with application operations, platform operations, engineering, architecture, and business stakeholders to design, deploy, and improve agent-based solutions. Compute Fundamentals: Demonstrates a deep understanding of compute concepts, including virtualization, containerization, operating systems, and system administration. Cloud Technologies: Experience with cloud platforms like GCP, AWS, and Azure, including their AI/ML services. Generative AI Experience with LLM and Generative AI and Google Cloud Products and services (e.g Vertex AI, Dialogflow, Gemini) ML Development: Experience with Machine Learning model development and deployment. AIML Frameworks: Experience with frameworks for deep learning (e.g. PyTorch, Tensorflow, Jax, Ray, etc.), AI accelerators (e.g. TPUs, GPUs), model architectures (e.g. encoders, decoders, transformers), and using machine learning APIs. Must have Associate Cloud Engineer (ACE) Certification Malaysia Software Engineering Professional PETALING JAYA, MY (0088) IBM Malaysia Sdn. Bhd.
Operational Governance: Maintaining operational governance of AI solutions by managing production stability, performance thresholds, and compliance requirements.
Stakeholder Collaboration: Collaborating with application operations, platform operations, engineering, architecture, and business stakeholders to design, deploy, and improve agent-based solutions.
Compute Fundamentals: Demonstrates a deep understanding of compute concepts, including virtualization, containerization, operating systems, and system administration.
...