Evaluate and pilot Generative AI and agentic services (Bedrock, AgentCore, Claude) for internal developer tooling.
Maintain the AI agent skills and runbooks used for automated incident investigation and root-cause analysis.
Implement and maintain monitoring, logging and alerting solutions to ensure the health and performance of cloud environments and applications running in the cloud following Site Reliability best practices.
...
Manage and optimize cloud infrastructure (AWS, Alicloud) for performance, cost, and reliability
Develop Devops platform like online load test, change management system
Leverage LLMs or AI frameworks (OpenAI, Dify, Agno, LangChain) to enhance automation in infrastructure operations, including intelligent alert triage, RCA (Root Cause Analysis), and chat-based operations (ChatOps)
...
We are seeking a DevOps Engineer (OpenStack). You will design, automate and run a secure, scalable OpenStack cloud, partnering with teams to integrate Kubernetes and improve reliability. You will troubleshoot complex platform issues, strengthen monitoring and security controls and help introduce AIOps capabilities to speed up operations.
We have wide experience in designing, developing, managing, delivering, testing and commissioning of large-scale systems system & solutions. Our pool of highly-trained IT Engineers and Managers helps to identify and implement the best and most appropriate IOT technology, to enable you to introduce timely and innovative products and services. In addition to the large-scale IOT systems system & solutions, we are also much engaged in digitalisations including customised mobile apps on IOS and Android platform with integrated support to different 3rd party systems.
We are committed to deliver top quality services by conforming to important ISO/IEC quality standards backed by our established best practices and effective service management systems and processes.
We are seeking a DevOps Engineer (OpenStack). You will design, automate and run a secure, scalable OpenStack cloud, partnering with teams to integrate Kubernetes and improve reliability. You will troubleshoot complex platform issues, strengthen monitoring and security controls and help introduce AIOps capabilities to speed up operations.
Real-world Feedback Loops: Leverage multi-channel user feedback and real-world task data as primary research signals; design experiments and datasets to continuously improve agent and retrieval performance in production scenarios
2-8+ Year hands-on experience with LLM, RAG and AI agent systems in production