This role combines systems administration, scripting, and process improvement to reduce manual work, improve reliability, and support day-to-day IT operations. Working closely with the Site Reliability Engineer, you will help maintain core systems, drive meaningful automation, and contribute to a more efficient and stable IT environment for our engineering teams.
Roles & Responsibilities:
- 1Own LAN and site connectivity, including switching, VLAN design, inter-VLAN routing, wireless, firewall policy, site-to-site VPN (including China links), and network segmentation. Maintain accurate topology records and ensure changes are documented and reversible.
- Administer core infrastructure, including Windows and Linux servers, virtualization (Proxmox/PVE or equivalent), storage, backup and restore, and Microsoft Entra ID / Microsoft 365 identity and access management.
- Own endpoint and workstation operations for engineering systems, including CAD/SolidWorks-class PCs: imaging, drivers, patching, endpoint security, and asset hygiene. Provide Level 2 support for floor incidents and resolve issues across network, identity, and endpoint layers.
- Apply operational security controls, including patching, system hardening, least-privilege access, logging and alerting, and periodic access reviews. Security is treated as part of day-to-day operations rather than a separate function.
- Manage vendor engagement for circuits, hardware, and warranty support, remaining accountable for service outcomes and delivery quality.
- Automate repeatable operational work using PowerShell, and Microsoft Power Automate where it is the appropriate tool. All automations must include access control, error handling, auditability, and a defined rollback path.
- Execute user lifecycle operations, including onboarding, offboarding, access provisioning, and ad-hoc operational support for engineering teams.
- Produce and maintain operational documentation suitable for handover and incident response, including network diagrams, runbooks, change records, and restore procedures.
- Track operational performance (incident rate, restore time, provisioning time, patch compliance) and propose improvements. Escalate risk to the Site Reliability Engineer with clear impact and recommended action.
Requirements
- Demonstrated hands-on experience operating enterprise or mid-size LAN environments, including VLAN design, switching, firewall policy, VPN, and wireless. Documentation-only network experience is insufficient.
- Systems administration experience covering Windows Server and Linux, plus virtualization (Proxmox/PVE, or equivalent).
- Practical identity and access administration in Microsoft Entra ID / Microsoft 365, or Active Directory integrated with Entra, including group management and joiner/mover/leaver processes.
- Endpoint and workstation support experience, including engineering or CAD-class PCs, imaging, peripheral support, and troubleshooting of connectivity and license-server access issues.
- Proven backup and restore experience, and applied operational security practices (patching, hardening, access restriction, and logging).
- Ability to work independently in a small team, document changes, and take operational issues from physical connectivity through identity and endpoint resolution.
Nice to have:
- Microsoft Intune
- Asset management platforms
- Production PowerShell
- Dual-site or China/APAC VPN operations
- Bilingual English and Mandarin
- A diploma or degree in IT or a related field is preferred but not required where equivalent hands-on experience is demonstrated.