Anlyvis develops Edge AIoT and AI-powered video analysis solutions that help cities and enterprises improve safety, security and operations. This role converts multimodal and vision models into efficient, reliable deployments beside live cameras.
Key Responsibilities
Deploy VLM and computer vision models on NVIDIA edge platforms.
Optimise latency, throughput, GPU utilisation, VRAM use and power consumption.
Apply FP16, FP8, INT8 or INT4 quantisation where appropriate.
Convert and optimise models using ONNX, TensorRT, TensorRT-LLM or vLLM.
Build RTSP, GStreamer and DeepStream pipelines with event-triggered VLM inference.
Benchmark, profile, debug and containerise production AI services.
Requirements
Degree in computer science, electronic engineering, AI or a related discipline.
Strong Linux and Python skills with NVIDIA GPU and CUDA experience.
Hands-on experience with TensorRT, DeepStream, GStreamer, ONNX, PyTorch or Docker.
Understanding of GPU memory, inference latency and throughput trade-offs.
Experience integrating and processing RTSP or other live video streams.
Preferred Qualifications
NVIDIA Jetson Orin and edge LLM or VLM deployment experience.
CUDA or C++, TensorRT plugin development, and Nsight profiling.
Kubernetes, remote-device operations or edge-fleet management experience.
What Success Looks Like
Datacentre-class models operate efficiently on edge GPUs.
Multiple camera streams run reliably without exhausting GPU memory.
Continuous perception invokes expensive VLM reasoning only when it adds value.
Local safety alerts remain available when cloud connectivity is interrupted.
How to Apply
Submit your CV, relevant publications or project portfolio, and a short note explaining your interest in Anlyvis and this role.
Full-time