Master Complex Problem Solving: Collaborate with cross-functional teams to diagnose complex network incidents, performing deep root-cause analysis to prevent recurrence
Maximize System Reliability: Define and maintain Service Level Objectives (SLOs) through advanced telemetry and streaming analytics to ensure "five-nines" availability
Influence Strategic Roadmaps: Translate technical roadmaps into business value for stakeholders and steer vendor product evaluations to align with long-term strategy
...
POC & demonstrations by lead and validate complex Proof of Concepts (POC), technical demonstrations, and site surveys to ensure solutions align with customer requirements and business objectives.
Vendor compliance, manage the team’s vendor certification roadmap to ensure the organization maintains necessary partner statuses and meets all technical requirement thresholds.
Business growth by collaborating with product team to design marketing programs and involves in meeting and quarterly sales quotas like QBR through technical leadership.
...
Ensure asset lifecycle aligns with security and audit requirements.
Manage secure disposal of IT assets in line with company policy and regulatory requirements and proper data sanitization / destruction before disposal.
...
· Collaborate with teams (e.g. Website, Communications) to solve technical challenges such as Wordpress management, integrations (Adaptis-ipay88) and system optimisation.
· Manage and optimise back-end systems (e.g. DynaMail, Alaya, SiteGIant, CiviCRM, Vodia, Automate, SiteGround, etc.) and system integrations (e.g. Adaptis-iPay88) to ensure reliability and scalability.
...
Full Stack EngineeringFrontend: React, Next.js, TypeScript, modern component architectures, state management, real-time and streaming AI interfaces, agent activity and execution interfaces, data visualization.Backend: Node.js, TypeScript, Python, REST APIs, GraphQL, WebSockets and streaming, event-driven architectures, background workers, job queues, distributed systems, authentication and authorization.
Distributed SystemsDesign systems that reliably execute thousands or millions of AI and data-processing tasks. Kubernetes, Docker, Cloud Run and serverless, message queues, Redis, Kafka or equivalent, distributed job processing, concurrency management, rate limiting, retries, idempotency, fault tolerance, observability. You know how to build systems that stay reliable when agents fail, APIs time out, models hallucinate, or downstream services go away.
Data & Learning InfrastructureBuild the infrastructure agents need to learn from historical executions. PostgreSQL, BigQuery or equivalent data warehouses, ClickHouse or analytical databases, vector databases, embeddings, retrieval systems, event logs, feature stores, analytics pipelines, data ingestion.
...