Senior AI DevOps Engineer
AI-Native Infrastructure & Agentic Swarm Oversight
Contract · Remote · Commitment to Results
Apply via ReturnOnTalent
About ReturnOnTalent
ReturnOnTalent connects high-performing professionals with cutting-edge roles in transformative organizations. Our client for this position is embarking on a landmark AI-project, leveraging models like Claude, DeepSeek, and MiniMax to rewrite and optimize their complex multi-language codebase (Go, Rust, Python). Integral to this endeavor is a secure, production-grade infrastructure led by a visionary Senior AI DevOps Engineer.
Job Responsibilities
As the Senior AI DevOps Engineer, you'll be pivotal in architecting, securing, and managing the infrastructure central to this AI initiative. Your role spans across the following focus areas:
1. DevSecOps & Infrastructure Foundation
- Mitigate critical vulnerabilities (e.g., exposed S3 buckets, leaked credentials).
- Implement secure, centralized secrets management using tools like HashiCorp Vault, AWS Secrets Manager, or SOPS.
- Develop infrastructure as code (IaC) frameworks with Terraform or Pulumi to enable reproducible, AI-integrated deployment.
- Design and manage container orchestration platforms (Kubernetes, AWS ECS, or GCP Cloud Run) focusing on Go, Rust, and Python workloads.
- Set up API Gateways (Kong, Traefik, Nginx) for seamless traffic redirection from legacy to microservices architecture.
- Implement observability platforms (OpenTelemetry, Prometheus/Grafana, or Datadog) for real-time monitoring of applications and AI workflows.
2. AI Vibecoding & Agentic Orchestration
- Configure AI-integrated IDEs (Cursor, GitHub Copilot) with standards enforcing secure and architecture-aligned code generation.
- Build sophisticated CI/CD pipelines (GitHub Actions) tailored to integrate secure AI-generated pull requests (SAST/DAST tools: SonarQube, Trivy).
- Operationalize a human-in-the-loop (HITL) system: AI proposes → CI/CD tests → Human reviews → Deploy.
- Manage isolated environments for safe execution of AI Agents’ code and testing.
- Optimize LLM API usage, controlling token costs and maintaining throughput via rate limits.
Ideal Talent
The successful candidate will excel in high-pressure, innovation-driven contexts, with the following strengths:
DevOps & Cloud Expertise
- 4–5+ years as a Senior DevOps or SRE Engineer in SaaS or multi-tenant architectures.
- Advanced skills in AWS or GCP, with networking fundamentals (VPC, Subnets, Security Groups, WAF).
- In-depth experience with Terraform or Pulumi for IaC, as well as Docker, Kubernetes, and Helm.
- Proficiency in designing multi-stage CI/CD pipelines that incorporate rigorous test suites.
- Familiarity with Go, Rust, Python, and modern TypeScript/Angular frontends.
AI Vibecoding & Mindset
- Highly competent user of AI coding tools (Cursor, GitHub Copilot, Claude, ChatGPT) with a collaborative AI-human workflow approach.
- Strong skills in prompt crafting and engineering for optimal AI-generated code outcomes.
- Familiarity with AI orchestration tools (LangChain, AutoGen, CrewAI) or a willingness to dive into them.
- Operate with a zero-trust security philosophy, ensuring AI-generated outputs meet reliable and secure deployment standards.
Communication & Working Style
- Fluent in English, with exceptional verbal and written communication skills.
- Self-starter with the confidence to build and execute without constant direction.
- Results-driven approach with a focus on delivering infrastructure solutions efficiently and sustainably.
Day-to-Day Activities
- Secure and optimize the infrastructure enabling AI-driven workflows and development.
- Collaborate seamlessly with AI agents, validating their solutions via a structured HITL CI/CD pipeline.
- Monitor, assess, and guide AI performance across LLM models and generated pipelines.
- Document processes meticulously to maintain clarity and collaboration in a fully remote, asynchronous environment.
- Lead infrastructure decision-making to enable scalability for an AI-native platform evolution.
Benefits
- Work Remotely: Flexibility to work from anywhere while actively contributing to a cutting-edge AI project.
- Innovative Environment: Be a key player in AI-native development leveraging the latest technologies and methodologies.
- High-Impact Role: Take ownership of a mission-critical infrastructure platform for a transformative organization.
- Competitive Compensation: Receive an attractive daily rate based on your expertise and performance.
How to Apply
If you’re excited about redefining how DevOps and AI converge, we invite you to apply. Please share the following:
- A brief summary of your DevOps/SRE experience relevant to this role.
- Your current AI tooling workflow and how you collaborate with AI agents.
- Examples of infrastructure you’ve built from scratch, including any noteworthy security remediation projects.
- Your daily rate expectations.