We are Lenovo. We do what we say. We own what we do. We WOW our customers.
Lenovo is a US$83 billion revenue global technology powerhouse, ranked #196 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit *************, and read about the latest news via our StoryHub.
Description and Requirements
Key Responsibilities
Deploy, manage, monitor, and troubleshoot Kubernetes (K8s) clusters and containerized applications in production environments.
Support and optimize CI/CD pipelines for software deployment and release management.
Develop automation scripts and tools to eliminate manual operational tasks and improve efficiency.
Manage day-to-day operational incidents, troubleshooting, and root cause analysis (RCA).
Monitor application and infrastructure performance, identify bottlenecks, and implement optimization plans.
Build and maintain observability solutions including logging, metrics collection, monitoring, and APM platforms.
Manage and maintain databases, including performance tuning, backup/recovery, and fault resolution.
Collaborate with development teams to improve deployment processes, system reliability, and operational excellence.
Support cloud-native infrastructure and AI/ML-related services.
Administer Linux-based environments and web platforms while ensuring high availability and security.
Implement infrastructure automation using modern DevOps tools and Infrastructure-as-Code principles.
Work with networking and security teams to maintain reliable and secure services.
Continuously evaluate emerging technologies and recommend improvements to systems and processes.
Preferred Qualifications
Experience supporting AI/ML platforms or AI-related infrastructure.
Experience building end-to-end observability solutions.
Strong understanding of GitOps methodologies.
Knowledge of Infrastructure as Code (Terraform, Ansible).
Experience in enterprise application operations and production support.
Familiarity with modern microservices and cloud-native environments.
Desired Competencies
Strong troubleshooting and analytical skills.
Excellent problem-solving capabilities.
Ability to work independently in a fast-paced environment.
Strong communication and collaboration skills.
Self-motivated with a passion for learning new technologies.
Good time management and prioritization skills.
Customer-focused mindset with a strong sense of ownership.
Education & Experience
Bachelor's degree in computer science, Information Technology, Engineering, or related discipline.
3+ years of experience in DevOps, Site Reliability Engineering (SRE), Cloud Operations, Platform Engineering, or IT Operations.
Hands-on experience supporting production applications and cloud-native infrastructure.