You will design and implement the platform's identity, access and security model across AWS IAM, Kubernetes RBAC and service identities, working with the storage access and security teams.
You will build the observability, alerting, capacity planning, and incident tooling for the platform. You will contribute to the SRE practice, which includes SLOs, runbooks, on-call, and post-incident reviews. You will reduce toil and MTTR.
You will own compute cost efficiency: instance and storage strategy, spot and right-sizing, bin-packing, idle reclamation, and cost attribution back to tenants.
...
Preparing material lists needed for service activity and ensuring that all materials, parts, and equipment are available and of appropriate quality for service activities.
Planning and executing work on a first-time right approach with customers, and ensuring the job is done on time and as per quality standards.
Identifying improvement needs and potential solutions for them in the ways of working.
...
Preparing material lists needed for service activity and ensuring that all materials, parts, and equipment are available and of appropriate quality for service activities.
Planning and executing work on a first-time right approach with customers, and ensuring the job is done on time and as per quality standards.
Identifying improvement needs and potential solutions for them in the ways of working.
...
Job SummaryWe are looking for a hands-on AI Engineer to design, develop, integrate, and deploy Generative AI, AI chatbot, and Agentic AI solutions. The role will focus on building practical AI applications and integrating LLM capabilities with existing applications, APIs, databases, and enterprise systems.Key ResponsibilitiesDesign and develop AI chatbots, copilots, and LLM-powered applications.Build Agentic AI solutions with reasoning, tool/function calling, multi-step task execution, and workflow automation.Implement RAG, embeddings, vector search, prompt engineering, and conversational memory.Work with commercial and open-source LLMs and select appropriate models based on business and technical requirements.Integrate AI solutions with REST APIs, databases, CRM/ERP systems, and enterprise applications.Develop AI integration services and APIs for existing systems and business workflows.Evaluate and optimize AI solutions for accuracy, reliability, latency, scalability, security, and cost.Collaborate with software engineers, architects, product teams, and business stakeholders.Keep up to date with developments in Generative AI, LLMs, Agentic AI, and AI automation.RequirementsBachelor's degree in Computer Science, AI/ML, Software Engineering, or a related field.3+ years of experience in AI/ML engineering, software engineering, or a related role.Hands-on experience developing LLM applications, AI chatbots, or Generative AI solutions.Practical experience with AI Agents, Agentic AI, tool/function calling, or AI workflow automation.Experience with open-source LLMs such as Llama, Qwen, Mistral, Gemma, or equivalent.Strong Python programming and software engineering skills.Experience with REST APIs, system integration, databases, and cloud/on-premise environments.Hands-on experience with RAG, vector databases, embeddings, and prompt engineering.Experience with frameworks such as LangChain, LangGraph, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, or equivalent.Experience with LLM platforms such as OpenAI, Azure OpenAI, Anthropic, Gemini, and/or open-source models.Familiarity with Docker, Git, CI/CD, and production deployment.Preferred SkillsExperience serving open-source LLMs using vLLM, Hugging Face, Ollama, or equivalent.Experience with MCP, multi-agent systems, LLM evaluation/observability, fine-tuning, LoRA/QLoRA, or model optimization.Experience with AWS, Azure, or Google Cloud.Knowledge of AI security, data privacy, access control, and responsible AI.
Conduct rigorous testing and provide ongoing maintenance for applications to uphold performance and reliability standards.
Cultivate deep insights into clients' businesses and industries, leveraging this understanding to identify and capitalize on new opportunities for innovation and improvement.
Foster strong client relationships through clear and effective communication, collaborating closely with both clients and superiors to achieve project success.
...
Deep expertise in modern Java (17+) including concurrency, multithreading, reactive programming, build tools (Gradle/Maven lifecycle), and design principles (SOLID), with experience in performance tuning, JVM optimization and troubleshooting.
Proven hands-on experience with Spring Boot ecosystem (Security, WebFlux, Data/MyBatis), strong understanding of Spring internals (beans, AOP, request handling, serialization), and API design using REST and OpenAPI.
Strong experience designing and operating microservices on AWS using ECS, DynamoDB and Aurora, with solid knowledge of event-driven systems and services such as SQS, SNS, EventBus and Lambdas, combined with CI/CD and infrastructure-as-code practices.
...
Proven hands-on experience with Spring Boot ecosystem (Security, WebFlux, Data/MyBatis), strong understanding of Spring internals (beans, AOP, request handling, serialization), and API design using REST and OpenAPI.
Strong experience designing and operating microservices on AWS using ECS, DynamoDB and Aurora, with solid knowledge of event-driven systems and services such as SQS, SNS, EventBus and Lambdas, combined with CI/CD and infrastructure-as-code practices.
Strong experience in modern testing approaches (e.g., RestAssured, TestContainers, WireMock, Localstack) and a solid understanding of cloud security, resilience, and operational excellence in enterprise environments.
...
Act as Technical Product Manager in operations, ensuring all technical components (servers, middleware components, application modules, interfaces, job chains, schedulers, connectivity paths) remain secure, compliant and lifecycle‑current.
Monitor, track and plan technology lifecycle events (EoL/EoS, patch cycles, hardware refresh, OS upgrades, middleware version changes) together with infrastructure and platform teams, ensuring risks are identified early and scheduled into IBSol governance cycles.
Execute Incident, Problem, Change and Release Management according to IBSol service standards, including detailed analysis, task coordination, root‑cause identification and technical approvals.
...