Responsibilities include monitoring system performance, resolving technical issues, coordinating with different teams for problem resolution, creating preventive measures(optional), and maintaining documentation related to system configuration, process, and service records.
· Monitor and analyse the current state of various product runtime environment (production and non-production) to ensure optimum system performance, and work out data-based strategy for continuous improvement. Work with application teams, solution architects, security consultants, and other teams to implement improvement plans.
Work closely with business users to setup AWS Quick environment to support AI use cases.
...
Evaluate, select, and validate servers and key hardware components, including CPUs, GPUs, memory, storage, network interface cards (NICs), and power systems.
Develop and maintain Bill of Materials (BOM), sizing models, capacity plans, and hardware configuration standards.
Perform hardware benchmarking, performance testing, compatibility validation, and capacity planning to ensure optimal infrastructure performance.
...
High-Severity Leadership: Excellent interpersonal ability to manage P1/P2 crises and deliver technical root-cause analyses directly to enterprise customers and accelerate L1 L2 teams
Ideal: Systems & Code Debugging skills, Strong Python coding/debugging skills, API troubleshooting, and distributed systems log analysis (Cloud Logging/Monitoring).
Project & Client Management – Take ownership of assigned technical projects and client engagements from planning and implementation through testing, deployment, and handover. Work directly with clients to understand requirements, provide technical recommendations, and communicate project or support updates.
Technical Collaboration & Continuous Improvement – Work closely with Sales, Pre-Sales, Product, vendors, and internal technical teams on solution design, proof-of-concepts, escalations, and technical issues. Research emerging technologies and identify opportunities to improve security, performance, and operational efficiency.
Documentation, Mentoring & Knowledge Sharing – Maintain technical documentation, implementation guides, configuration records, and support reports. Provide technical guidance to junior engineers and contribute to knowledge sharing, best practices, and team capability development.
...
Mentor engineers through code reviews, pair programming, and documentation to raise team standards and reduce defects. Set coding conventions, run knowledge-sharing sessions, and help junior engineers take ownership of operational tasks.
Build a visible portfolio of production systems by driving deployments, monitoring, and structured postmortems that show operational thinking. Own service-level metrics, alerting, and incident follow-up, and present outcomes to product and client stakeholders.