Build, manage, and enhance observability platforms, including logging, metrics collection, monitoring, alerting, and performance analysis.
Collaborate closely with software development teams to improve CI/CD pipelines, deployment automation, and software delivery processes.
Perform daily operational health checks, respond to production incidents, conduct root cause analysis (RCA), and implement preventive improvements.
...
Observability Implementation: You will support the observability culture by directly assisting developers to onboard their applications onto OpenTelemetry (Otel) for distributed tracing and metric collection. You will deploy and maintain the Grafana LGTM stack (Loki, Tempo, Mimir) and Grafana Alloy to implement designated alerting and notification strategies.
On-Prem Platform Stability: You will maintain and optimize our established on-premise infrastructure to ensure high availability and stability. This includes supporting our Docker Swarm cluster and RHEL-based internal VMs across segmented networks (Dev, Staging, Prod) according to standard operating and audit procedures.
Infrastructure as Code (IaC) & Configuration Management: You will write, maintain, and version-control clear IaC scripts (such as Ansible, Terraform) to consistently provision AWS infrastructure and maintain configuration baselines across both cloud and on-premise environments.
...