jobs in Autonomai Recruitment

Autonomai Recruitment Hiring! Full Time Production Engineer in - Ricebowl

Production Engineer

Autonomai Recruitment

Undisclosed

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

C++ Site Reliability Engineer

Location: Singapore

Industry: Electronic Markets


About the Firm

We are working with a technology-led trading firm that builds and operates highly automated, low-latency systems across global financial markets.


The firm operates more like a world-class technology company than a traditional financial institution: engineers own complex systems end to end, work closely with quantitative researchers and traders, and are expected to solve difficult reliability, performance, and scale challenges with a high degree of autonomy.


This is an opportunity for an experienced engineer from a leading technology organisation to apply their systems, production engineering, and distributed infrastructure expertise in an environment where microseconds, resilience, and engineering quality matter.


The Role

As a C++ Site Reliability Engineer, you will work at the intersection of software engineering, infrastructure, and production reliability.

You will design, build, and operate the platforms that support latency-sensitive trading applications and critical real-time services. The role involves developing high-performance C++ systems while also improving deployment, observability, automation, incident response, and operational resilience.


You will be expected to go beyond traditional operations. Rather than simply maintaining existing infrastructure, you will improve the engineering systems, tooling, and architecture that allow trading platforms to operate reliably at very high throughput and low latency.


Responsibilities

  • Design, develop, and maintain high-performance C++ services and systems used in production trading environments.
  • Improve the reliability, scalability, and operational performance of latency-sensitive applications.
  • Build tooling for service orchestration, deployment, configuration management, health checking, and automated remediation.
  • Investigate complex production issues across application, operating system, network, and infrastructure layers.
  • Develop monitoring, alerting, logging, tracing, and service-level indicators for critical systems.
  • Improve release engineering and deployment processes, including safe rollout, rollback, and validation mechanisms.
  • Automate operational processes using C++, Python, Go, or similar languages.


Required Experience

  • Strong professional experience developing production software in modern C++.
  • Experience operating or supporting large-scale, highly available, or performance-sensitive production systems.
  • Strong understanding of Linux internals, including processes, scheduling, memory, filesystems, networking, and system performance.
  • Experience debugging complex issues using tools such as gdb, perf, eBPF, flame graphs, core dumps, packet captures, or equivalent technologies.
  • Demonstrable experience with automation, infrastructure tooling, deployment systems, or production engineering.
  • Solid understanding of distributed systems, service reliability, failure modes, and fault-tolerant design.


Desirable Experience

Experience in one or more of the following areas would be advantageous:

  • Ultra-low-latency or real-time systems.
  • Kubernetes, container platforms, or large-scale service orchestration.
  • Bare-metal fleet management and data centre infrastructure.
  • Time synchronisation, precision timing, PTP, or hardware timestamping.
  • Python, Go, Rust, or shell scripting.
  • Experience at a major technology company operating internet-scale, distributed, or mission-critical platforms.


Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More