jobs in Autonomai Recruitment

全职 Production Engineer 工作, 薪水, Autonomai Recruitment 公司招聘中 - Ricebowl

Production Engineer

Autonomai Recruitment

Undisclosed

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

C++ Site Reliability Engineer

Location: Singapore

Industry: Electronic Markets


About the Firm

We are working with a technology-led trading firm that builds and operates highly automated, low-latency systems across global financial markets.


The firm operates more like a world-class technology company than a traditional financial institution: engineers own complex systems end to end, work closely with quantitative researchers and traders, and are expected to solve difficult reliability, performance, and scale challenges with a high degree of autonomy.


This is an opportunity for an experienced engineer from a leading technology organisation to apply their systems, production engineering, and distributed infrastructure expertise in an environment where microseconds, resilience, and engineering quality matter.


The Role

As a C++ Site Reliability Engineer, you will work at the intersection of software engineering, infrastructure, and production reliability.

You will design, build, and operate the platforms that support latency-sensitive trading applications and critical real-time services. The role involves developing high-performance C++ systems while also improving deployment, observability, automation, incident response, and operational resilience.


You will be expected to go beyond traditional operations. Rather than simply maintaining existing infrastructure, you will improve the engineering systems, tooling, and architecture that allow trading platforms to operate reliably at very high throughput and low latency.


Responsibilities

  • Design, develop, and maintain high-performance C++ services and systems used in production trading environments.
  • Improve the reliability, scalability, and operational performance of latency-sensitive applications.
  • Build tooling for service orchestration, deployment, configuration management, health checking, and automated remediation.
  • Investigate complex production issues across application, operating system, network, and infrastructure layers.
  • Develop monitoring, alerting, logging, tracing, and service-level indicators for critical systems.
  • Improve release engineering and deployment processes, including safe rollout, rollback, and validation mechanisms.
  • Automate operational processes using C++, Python, Go, or similar languages.


Required Experience

  • Strong professional experience developing production software in modern C++.
  • Experience operating or supporting large-scale, highly available, or performance-sensitive production systems.
  • Strong understanding of Linux internals, including processes, scheduling, memory, filesystems, networking, and system performance.
  • Experience debugging complex issues using tools such as gdb, perf, eBPF, flame graphs, core dumps, packet captures, or equivalent technologies.
  • Demonstrable experience with automation, infrastructure tooling, deployment systems, or production engineering.
  • Solid understanding of distributed systems, service reliability, failure modes, and fault-tolerant design.


Desirable Experience

Experience in one or more of the following areas would be advantageous:

  • Ultra-low-latency or real-time systems.
  • Kubernetes, container platforms, or large-scale service orchestration.
  • Bare-metal fleet management and data centre infrastructure.
  • Time synchronisation, precision timing, PTP, or hardware timestamping.
  • Python, Go, Rust, or shell scripting.
  • Experience at a major technology company operating internet-scale, distributed, or mission-critical platforms.


重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多