Quality Assurance & Testing: Execute user acceptance testing (UAT) for new automations and integrations, document test cases, verify edge-case handling, and monitor post-deployment error rates to ensure production stability.
System Operations & Backlog Management: Triage, investigate, and prioritize internal support tickets, bugs, and feature requests while ensuring projects meet delivery timelines and SLAs.
Impact Tracking & Analytics: Analyze and monitor operational dashboards tracking system adoption, process efficiency (hours saved, error reduction, SLA compliance), and automation health. Produce weekly reports, proactively flag regressions, and present data-driven recommendations in team huddles.
...
Operational Excellence: Ensure timely renewal and compliance of hardware maintenance contracts and licenses. Implement and regularly test disaster recovery plans for business continuity. Recommend tech enhancements to streamline operations and reduce costs. Manage physical and virtual servers for optimal performance
Cloud & Infrastructure Strategy: Architect and implement cloud-based solutions aligned with business goals, ensuring consistency and cost-efficiency. Collaborate across teams to drive a standardized, modernized cloud infrastructure roadmap with a unified operating model
Solution Delivery & Support: Develop and support IT infrastructure solutions with high availability, scalability, and performance as well as act as a Level 3/4 subject matter expert for server and cloud-related issues
...
Participate actively in site meetings, factory inspections, and acceptance testing to resolve design and construction challenges while ensuring technical compliance and quality assurance.
Oversee the preparation and submission of Operation and Maintenance Manuals, guaranteeing user-friendly and thorough documentation for project handover.
Manage construction site supervision workflows, including the review and approval of shop drawings, materials submissions, As-Built drawings, and technical queries to ensure conformity with approved designs and regulation.
...
Help align the design system with company goals by improving efficiency, consistency, and scalability across product teams and brands
Participate in conversations and workshops with stakeholders, users, and contributors to ensure system scalability and strong relationships
Explore and apply AI-assisted development workflows to help the team work more effectively, improve quality, strengthen documentation, and create a better developer experience
...
Help align the design system with company goals by improving efficiency, consistency, and scalability across product teams and brands
Participate in conversations and workshops with stakeholders, users, and contributors to ensure system scalability and strong relationships
Explore and apply AI-assisted development workflows to help the team work more effectively, improve quality, strengthen documentation, and create a better developer experience
...
Help align the design system with company goals by improving efficiency, consistency, and scalability across product teams and brands
Participate in conversations and workshops with stakeholders, users, and contributors to ensure system scalability and strong relationships
Explore and apply AI-assisted development workflows to help the team work more effectively, improve quality, strengthen documentation, and create a better developer experience
...
Full Stack EngineeringFrontend: React, Next.js, TypeScript, modern component architectures, state management, real-time and streaming AI interfaces, agent activity and execution interfaces, data visualization.Backend: Node.js, TypeScript, Python, REST APIs, GraphQL, WebSockets and streaming, event-driven architectures, background workers, job queues, distributed systems, authentication and authorization.
Distributed SystemsDesign systems that reliably execute thousands or millions of AI and data-processing tasks. Kubernetes, Docker, Cloud Run and serverless, message queues, Redis, Kafka or equivalent, distributed job processing, concurrency management, rate limiting, retries, idempotency, fault tolerance, observability. You know how to build systems that stay reliable when agents fail, APIs time out, models hallucinate, or downstream services go away.
Data & Learning InfrastructureBuild the infrastructure agents need to learn from historical executions. PostgreSQL, BigQuery or equivalent data warehouses, ClickHouse or analytical databases, vector databases, embeddings, retrieval systems, event logs, feature stores, analytics pipelines, data ingestion.
...