Observability Implementation: You will support the observability culture by directly assisting developers to onboard their applications onto OpenTelemetry (Otel) for distributed tracing and metric collection. You will deploy and maintain the Grafana LGTM stack (Loki, Tempo, Mimir) and Grafana Alloy to implement designated alerting and notification strategies.
On-Prem Platform Stability: You will maintain and optimize our established on-premise infrastructure to ensure high availability and stability. This includes supporting our Docker Swarm cluster and RHEL-based internal VMs across segmented networks (Dev, Staging, Prod) according to standard operating and audit procedures.
Infrastructure as Code (IaC) & Configuration Management: You will write, maintain, and version-control clear IaC scripts (such as Ansible, Terraform) to consistently provision AWS infrastructure and maintain configuration baselines across both cloud and on-premise environments.
...
Define and document standard operating procedures and runbooks for using automation and tools (how to request changes, how pipelines work, how to roll back)
Support incident and problem management by providing tooling, diagnostics scripts, and data exports to accelerate root cause analysis and post incident reviews
Strong practical Python and shell scripting skills to build automation, small tools, and integrations without existing frameworks
...