- Singapore
Working Location
Job Description
Responsibilities
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to join our Engineering and Technology team. You will establish how we measure, validate and communicate the health of GPU infrastructure used for customer and internal workloads. You will define trusted health signals and service-readiness criteria, and turn them into reusable dashboards, alerts, queries, diagnostic checks and operational guidance. Your work will help commissioning, infrastructure and operations teams bring capacity online safely, identify degradation early and recover from failures quickly. You will also make knowledge self-service by publishing clear reference implementations, runbooks and AI-ready operational knowledge that other teams can use and extend.
Define GPU and host health criteria for customer and internal workloads, and the service-readiness gates that repair and capacity workflows depend on. Cover what stops or corrupts AI jobs: GPUs falling off the bus, XID events, ECC and memory faults, NVLink/NVSwitch degradation, thermal and power capping, NCCL and collective failures, silent data corruption, and stragglers running below fleet baseline.
Publish golden dashboards, alerts, PromQL/LogQL queries and health checks that other teams adopt and extend. Your output is the standard and the examples. The alerts must be specific, low-noise, and with a clear next action.
DCGM checks, NCCL and bandwidth tests, stress and burn-in, and validation jobs that confirm a server matches expected performance. Others must be able to run them without you.
Separate a bad GPU from a cooling, host, network or power-limit problem using host and BMC/management-interface telemetry together, including when the same signature appears across many servers. A clean management view means nothing if the host is throwing faults.
Rising correctable ECC counts, NVLink retries, thermal slowdown, XID patterns, wrong results with no error. Keep the knowledge current: what each signal means, what to do next on the machine, and where the operation team must make the final decision.
Work with commissioning on bring-up and acceptance, operations on break-fix, infrastructure on what healthy hardware looks like, and the telemetry owners on making your signals production-grade. Run technical sessions with customer teams on the telemetry they need. Join GPU and host incidents, including debugging on the server and convert every finding into a reusable check, alert or runbook. Give engineering leadership a straight read on fleet health and risk.
Highly Desirable Experiences
Location & Reporting
Employment Basis
Full-time
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.