Collaborate with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
Troubleshoot and resolve issues across production and non-production environments, providing root cause analysis and designing solutions to prevent future occurrences
Determine areas of inefficient resource utilization and implement changes to ensure both platform stability and cost efficiency
...
Industry: Enterprise IT Infrastructure, Managed Services, Banking, Data Centers, Cloud & Infrastructure Services
Environment: Large enterprise or multinational organizations supporting mission-critical infrastructure, Microsoft and Linux platforms, virtualization, storage, and identity services.
Proven hands-on experience with one or more programming languages (Java & SpringBoot preferred), scripting language and CICD/DevOps tooling (GitLab, Terraform, EKS) using modern technologies.
Ability to take ownership, collaborate closely with cross-functional teams and communicate clearly in an agile environment.
Solid understanding of cloud security, resilience, and operational best practices.
...
Analyze energy losses, downtime events, inverter trips, curtailment, shading, equipment degradation, and grid-related issues.
Work closely with relevant stakeholders to conduct root cause analysis (RCA) for underperforming assets and develop corrective and preventive action plans.
Prepare daily, weekly, monthly, quarterly, and annual asset performance reports for management and stakeholders.
...
In addition, you are leading the software engineering direction by establishing engineering standards, promoting DevSecOps best practices, providing technical coaching and contributing to the AWS cloud software engineering community.
Bachelor’s degree in Computer Science, Information Systems Management or in similar fields – Master/PhD is a bonus (CGPA > 3.0 or similar grading is a must)
8+ years of experience in software engineering with a track record of ownership in cross-functional teams, excellent communication skills, and a strong focus on continuous learning, technical curiosity, mentoring and driving high-impact engineering outcomes.
...
Collaborate with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems
Troubleshoot and resolve issues across production and non-production environments, providing root cause analysis and designing solutions to prevent future occurrences
Determine areas of inefficient resource utilization and implement changes to ensure both platform stability and cost efficiency
...
Observability Implementation: You will support the observability culture by directly assisting developers to onboard their applications onto OpenTelemetry (Otel) for distributed tracing and metric collection. You will deploy and maintain the Grafana LGTM stack (Loki, Tempo, Mimir) and Grafana Alloy to implement designated alerting and notification strategies.
On-Prem Platform Stability: You will maintain and optimize our established on-premise infrastructure to ensure high availability and stability. This includes supporting our Docker Swarm cluster and RHEL-based internal VMs across segmented networks (Dev, Staging, Prod) according to standard operating and audit procedures.
Infrastructure as Code (IaC) & Configuration Management: You will write, maintain, and version-control clear IaC scripts (such as Ansible, Terraform) to consistently provision AWS infrastructure and maintain configuration baselines across both cloud and on-premise environments.
...
Proven hands-on experience with one or more programming languages (Java & SpringBoot preferred), scripting language and CICD/DevOps tooling (GitLab, Terraform, EKS) using modern technologies.
Ability to take ownership, collaborate closely with cross-functional teams and communicate clearly in an agile environment.
Solid understanding of cloud security, resilience, and operational best practices.
...