Senior Software Engineer, GPU Cluster Infrastructure
FAR AI
Posted
Sep 14, 2026
Location
Remote
Type
Full-time
Compensation
$150000 - $275000
Mission
What you will drive
- In this role, you'll operate FAR.AI's Kubernetes GPU cluster infrastructure for AI research.
- Manage cluster fleet operations including node lifecycle, driver rollouts, and capacity planning.
- Own batch scheduling and multi-tenancy with queues, quotas, priorities, and fair-share allocation.
- Design distributed storage systems for datasets and checkpoints with performance and fault-tolerance.
- Harden platform security through access controls, network policies, and workload isolation.
Profile
What makes you a great fit
- In this role, you'll operate FAR.AI's Kubernetes GPU cluster infrastructure for AI research.
- Manage cluster fleet operations including node lifecycle, driver rollouts, and capacity planning.
- Own batch scheduling and multi-tenancy with queues, quotas, priorities, and fair-share allocation.
- Design distributed storage systems for datasets and checkpoints with performance and fault-tolerance.
- Harden platform security through access controls, network policies, and workload isolation.
About
Inside FAR AI
FAR AI aims to ensure AI systems are trustworthy and beneficial to society. They incubate and accelerate research agendas that are too resource-intensive for academia but not yet ready for commercialisation by industry.