Job Summary
Gather and analyze system and application metrics to optimize performance and reliability. Collaborate with cross-functional teams to enhance services, automate workflows, and deliver projects on time while balancing feature development and operational stability.
Responsibilities
Analyze operating system and application metrics to identify performance issues and support fault diagnosis
Collaborate with cross-functional teams to test workflows and implement best practices that improve service quality
Participate in system design consulting, platform management, and capacity planning using Broadcom VMware knowledge where applicable
Automate processes and implement system uplifts to develop sustainable systems and services
Define and meet service-level objectives to balance feature development speed with system reliability
Proactively identify problems, performance bottlenecks, and improvement opportunities through automation and workflow enhancements
Coordinate with delivery teams to ensure timely project and service delivery while maintaining operational focus
Apply software development lifecycle (SDLC) understanding and development knowledge to support system improvements
Required Competencies and Certifications
Bachelor Degree in Computer Science or related fields
6–8 years or more of experience in IT infrastructure operations or related fields
Preferred Competencies and Qualifications
Proficiency in programming using structured and object-oriented programming languages
Experience with infrastructure and related technologies
Understanding of continuous integration and continuous delivery (CI/CD) processes
Excellent written and verbal communication skills
Team-oriented with a strong customer service mindset
Ability to work independently and manage multiple tasks simultaneously
Strong analytical skills and attention to detail
Technical Knowledge
Operating Systems: RedHat Linux, Windows, CentOS, Ubuntu
Container Orchestration & Runtime: Docker, Kubernetes
Monitoring & Performance: Observability tools