- Petaling Petaling Selangor Malaysia
工作地点
职位描述
岗位职责
Project Implementation and Solution Design:
• Conduct technical discovery sessions and workshops to understand customer requirements, existing infrastructure, operating model, security requirements, and service expectations.
• Lead OpenShift, Ansible Automation Platform and related Opensource project implementations, including scope review, Pre-Implementation Validation (PIV), solution design, deployment planning, and acceptance criteria.
• Design scalable OpenShift architectures covering cluster topology, compute, networking, ingress, load balancing, storage, image registry, identity integration, monitoring, logging, high availability, cluster operators and disaster recovery.
• Design and implement AAP solutions including automation controller, private automation hub, execution environments, inventories, projects, credentials, job templates, workflows, schedules, and role-based access controls.
• Prepare high-level and low-level technical design documents, implementation runbooks, deployment checklists, text plans, handover documentation and capturing technical issues for future reference.
• Act as Project Technical Lead when required, contributing technical scope definition, implementation approach, assumptions, dependencies, risks, and man-day estimate for pre-sales and project teams.
• Ensure project completion within agreed timelines, coordinating closely with
project managers.
Technical Support and Incident Management:
• Provide advanced technical support for OpenShift and Ansible Automation Platform environments, including remote and onsite troubleshooting for production, UAT, and disaster recovery environments.
• Diagnose and resolve complex issues involving Kubernetes workloads, cluster operators, networking, DNS, ingress, certificates, storage, authentication, automation execution, and integration points.
• Lead incident investigation, root cause analysis, corrective actions, and technical problem management activities for high-severity customer issues.
• Plan and execute platform lifecycle activities such as OpenShift upgrades, operator upgrades, security remediation, AAP upgrades, backup validation, and health checks in accordance with approved change procedures.
• Ensure incident updates, service records, and technical findings are documented accurately and in a timely manner to support contractual Service Level Agreements (SLAs).
Platform Engineering, Automation and Integration:
• Develop, test, maintain, and support Ansible playbooks, roles, collections, inventories, execution environments, and automation workflows using secure, reusable, and version-controller practices.
• Integrate AAP with enterprise platforms and services such as Git repositories, CI/CD tools, ITSM systems, VMWare, Linux, cloud platforms, APIs, databases, monitoring tools, and identity services.
• Support container application onboarding by providing guidance on image build practices, container registries, namespaces, RBAC, resource management, persistent storage, network policies, and day-2 operations.
• Apply infrastructure-as-code and DevOps practices using Git, YAML, jinja2, Ansible, Terraform or equivalent tools, while maintaining appropriate peer review, testing, documentation, and change control.
• Support security and operational resilience through least-privilege access, secrets management, certificate lifecycle management, auditability, backup and recovery validation, capacity monitoring, and performance optimization.
Technical Leadership and Knowledge Sharing:
• Conduct technical workshops, solution walkthroughs, demonstrations, and knowledge-transfer sessions for customers and internal technical teams.
• Provide technical guidance and escalation support to Level 1 and Level 2 engineers, promoting consistent support practices and faster issue resolution.
• Work closely with cross-functional teams, technology partners, and Red Hat support where required to ensure timely resolution of complex technical issues.
• Stay current with Red Hat product updates, Kubernetes ecosystem developments, automation practices, and relevant industry standards; recommend practical improvements to services and delivery methods.
• Contribute reusable templates, automation content, reference architectures, troubleshooting guides, and lessons learned to the team knowledge base.
Qualifications: Degree in Information Technology, Computer Science, Engineering, or
an equivalent related field.
Relevant Industry Experience: Minimum of 5 years of related working experience in an IT environment, or technical consulting. Candidates with 5 or more years of customer facing OpenShift, Kubernetes, or automation delivery experience are preferred.
Technical Expertise:
• Hands-on experience with Red Hat OpenShift and Kubernetes, including cluster administration, core control plane concepts, worker nodes, namespaces/projects, operators, application deployment, resource management, and troubleshooting.
• Strong knowledge of container technologies and runtime concepts, including Docker or Podman, CRI-O, container images, image registries, image security, and container build practices.
• Hands-on experience with Ansible and Ansible Automation Platform, including playbooks, roles, collections, inventories, credentials, projects, job templates, workflows, execution environments, automation hub, and API or webhook integration.
• Good Linux administration skills, particularly Red Hat Enterprise Linux or compatible distributions, including system services, package management, filesystem and storage administration, performance analysis, security hardening, and troubleshooting.
• Sound networking knowledge, including TCP/IP, DNS, NTP, VLANs, load balancing, ingress, routing, proxy configuration, firewall concepts, certificates/TLS, and network troubleshooting.
• Experience with OpenShift infrastructure integrations such as storage (CSI, ODF/Ceph, NFS, SAN/NAS), virtualization (VMware vSphere, KVM, or equivalent), identity services (LDAP, Active Directory, SSO), and enterprise backup solutions.
• Familiarity with observability and logging technologies such as Prometheus, Grafana, Alert manager, Loki, Elasticsearch, Open Telemetry, or equivalent tools.
• Knowledge of high availability, multi-tenancy, backup and recovery, disaster recovery, capacity planning, security governance, and operational readiness for enterprise platform environments.
• Experience with source control, CI/CD, and GitOps tools such as GitLab, Bitbucket, GitHub, Jenkins, Argo CD, Tekton, or equivalent tools is an advantage.
• Knowledge of cloud platforms such as AWS, Microsoft Azure, Google Cloud, or Red Hat OpenShift cloud services is an advantage.
Preferred Experience:
• Experience delivering OpenShift and AAP solution in regulated, financial services, enterprise, or multi-team environments.
• Exposure to service mesh technologies such as Red Hat OpenShift Service Mesh, lstio, or related traffic-management solutions.
• Experience with infrastructure provisioning or configuration tools such as Terraform, Red Hat Satellite, VMware automation, REST APIs, Bash, or Python scripting.
• Experience working directly with customers to facilitate workshops, present technical designs, manage technical risks, and support project acceptance.
Certifications:
Preferably with one or more certifications below:
• Red Hat Certified System Administrator (RHCSA)
• Red Hat Certified System Engineer (RHCE)
• Red Hat Certified Specialist in OpenShift or OpenShift-related certification
• Red Hat Certified Specialist in Ansible Automation Platform or automation-related certification
• Red Hat Certified Specialist in High Availability Clustering
• Red Hat Certified Specialist in Deployment and Systems Management
• Relevant Kubernetes, infrastructure-as-code, or DevOps certifications
Additional Skills: Strong communication and problem-solving skills, with the ability to explain complex technical concepts clearly to both technical and non-technical stakeholders. work independently. Proactive with a positive attitude.
重要安全守则
申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。