Senior Software Engineer, GPU Cluster Infrastructure
FAR.AI
Posted
Sep 14, 2026
Location
Remote
Type
Full-time
Mission
What you will drive
Build and operate large-scale GPU cluster infrastructure for AI safety research, managing Kubernetes fleets, storage, scheduling, and security.
- Operate Kubernetes GPU clusters day-to-day: node lifecycle, upgrades, capacity planning, and staged rollouts with safe rollback.
- Own batch scheduling, multi-tenancy, quotas, priorities, and fair-share allocation across research teams.
- Design and manage distributed storage for datasets, checkpoints, and backups; ensure multi-node training fault tolerance and NCCL debugging.
- Harden platform security: identity, access control, network policy, workload isolation, and sandboxing for autonomous agents.
- Requires 3+ years in production Linux systems or infrastructure engineering on GPU, HPC, or batch platforms; strong Kubernetes and infrastructure-as-code experience.
Profile
What makes you a great fit
Build and operate large-scale GPU cluster infrastructure for AI safety research, managing Kubernetes fleets, storage, scheduling, and security.
- Operate Kubernetes GPU clusters day-to-day: node lifecycle, upgrades, capacity planning, and staged rollouts with safe rollback.
- Own batch scheduling, multi-tenancy, quotas, priorities, and fair-share allocation across research teams.
- Design and manage distributed storage for datasets, checkpoints, and backups; ensure multi-node training fault tolerance and NCCL debugging.
- Harden platform security: identity, access control, network policy, workload isolation, and sandboxing for autonomous agents.
- Requires 3+ years in production Linux systems or infrastructure engineering on GPU, HPC, or batch platforms; strong Kubernetes and infrastructure-as-code experience.
About
Inside FAR.AI
FAR.AI is an AI safety research nonprofit focused on addressing risks from advanced AI systems.