Position Summary
As an AI Platform Engineer, you will be a key technical driver in designing, operating, securing, and scaling Central AI Platform built on hybrid cloud infrastructure, Red Hat OpenShift, and RE:AI IaaS. This role bridges enterprise-grade networking, site reliability engineering (SRE), cloud platform operations, and DevSecOps. You will ensure seamless multicloud connectivity, robust security postures, high availability, and operational excellence for critical enterprise AI/ML workloads.
Key Responsibilities
1. Platform Infrastructure & Operations
- Oversee continuous availability, performance optimization, and observability of the Azure AI and Red Hat OpenShift-powered RE:AI cloud platform environments.
- Manage end-to-end incident response, root cause analysis (RCA), disaster recovery (DR) strategies, and SRE practices to guarantee platform resilience.
- Drive platform automation through Infrastructure-as-Code (IaC) and script workflows to support model hosting, GPU workloads, and data pipeline scalability.
- Monitor system health, bandwidth utilization, latency, and throughput while proactively leading capacity planning and scaling initiatives.
2. Networking & Hybrid Connectivity Architecture
- Design, build, and maintain secure hybrid connectivity connecting RE:AI IaaS, public cloud environments (Azure/AWS/GCP), on-premises data centers, and external partner systems.
- Configure and manage routing protocols (BGP), network segmentation, peering, VPNs, DirectConnect/ExpressRoute, and API gateway architectures.
- Coordinate physical connectivity needs, including cross-connects, colocation facilities, and data center interconnects.
3. Security, Compliance & DevSecOps
- Enforce Zero Trust principles, firewall policies, network security zoning, encryption in transit, and privileged access controls.
- Oversee cybersecurity operations including vulnerability management, threat detection, endpoint detection, and SIEM integration.
- Ensure platform compliance with cybersecurity policies and industry standards (ISO 27001, CIS, NIST).
- Partner with DevSecOps and developer teams to embed security and automated observability into AI/ML lifecycle pipelines.
Skills & Qualifications
- Bachelor’s degree in Computer Engineering, Computer Science, Network Engineering, Information Security, or a related technical discipline.
- Certifications (Advantageous):
- Networking/Cloud: CCNP/CCIE, AWS/Azure/GCP Networking or Solutions Architect Specialty
- Security/Operations: CISSP, CISM, or Red Hat OpenShift certifications.
Required Technical Experience
- Experience: 5+ years of experience in cloud infrastructure administration, hybrid networking, platform engineering, or SRE environments.
- Cloud & Container Infrastructure: Deep expertise in Azure platform operations and hands-on experience with Red Hat OpenShift clusters, Kubernetes, or AKS environments.
- Networking & Security: Strong foundation in TCP/IP, BGP, SD-WAN, cloud connectivity, firewall management, and network segmentation.
- Automation & Observability: Proficiency in IaC (Terraform, Bicep, ARM) and scripting (Python, PowerShell, Bash), along with experience in monitoring suites (Azure Monitor, Log Analytics, Application Insights, or OpenShift observability).
- AI/ML Familiarity: Experience supporting or operating AI/ML infrastructure needs (e.g., GPU VMs, inference endpoints, data pipelines, model hosting).
Core Competencies
- Systems Thinking: Strong architectural mindset capable of balancing performance, security, availability, and cost.
- Crisis Management: Excellent problem-solving skills and calm execution during high-pressure incidents.
- Collaboration: Clear communication style with the ability to bridge technical requirements across multi-disciplinary teams.