Incident Management: Handle and manage network incidents, including recording incidents, escalating as needed, and ensuring timely resolution. Collaborate with cross-functional teams, such as network engineers, system administrators, and vendors, to resolve complex issues.
Performance Monitoring and Optimization: Monitor network performance metrics and identify areas for improvement. Recommend and implement network optimization strategies to enhance performance, reliability, and scalability.
Collaboration and Communication: Collaborate with other IT teams, including network engineers, system administrators, and help desk personnel, to address network-related issues. Communicate effectively with stakeholders, providing timely updates on network incidents and resolutions.
...
Incident Management: Handle and manage network incidents, including recording incidents, escalating as needed, and ensuring timely resolution. Collaborate with cross-functional teams, such as network engineers, system administrators, and vendors, to resolve complex issues.
Performance Monitoring and Optimization: Monitor network performance metrics and identify areas for improvement. Recommend and implement network optimization strategies to enhance performance, reliability, and scalability.
Collaboration and Communication: Collaborate with other IT teams, including network engineers, system administrators, and help desk personnel, to address network-related issues. Communicate effectively with stakeholders, providing timely updates on network incidents and resolutions.
...
Responsible for the design, implementation, operation and continuous improvement of network and security infrastructure supporting large-scale AI/GPU computing platforms, data centres and cloud services.The role combines strong data-center networking and security fundamentals with high-performance AI networking technologies, including NVIDIA InfiniBand, Spectrum Ethernet/Spectrum-X and RoCEv2. The candidate must demonstrate strong networking fundamentals, hands-on troubleshooting capability and the ability to rapidly learn and develop expertise in AI/GPU networking.The role will also provide technical leadership during complex incidents, infrastructure deployments and network expansion projects, while mentoring other engineers and working closely with Operations, Systems, Platform, Security, Data Centre teams and external technology partners. The role requires participation in on-call support as needed.
...