AI Safety & Governance Full-time

Senior Software Engineer, GPU Cluster Infrastructure

FAR AI

Posted

Sep 14, 2026

Location

Remote

Type

Full-time

Compensation

$150000 - $275000

Mission

What you will drive

  • In this role, you'll operate FAR.AI's Kubernetes GPU cluster infrastructure for AI research.
  • Manage cluster fleet operations including node lifecycle, driver rollouts, and capacity planning.
  • Own batch scheduling and multi-tenancy with queues, quotas, priorities, and fair-share allocation.
  • Design distributed storage systems for datasets and checkpoints with performance and fault-tolerance.
  • Harden platform security through access controls, network policies, and workload isolation.

Profile

What makes you a great fit

  • In this role, you'll operate FAR.AI's Kubernetes GPU cluster infrastructure for AI research.
  • Manage cluster fleet operations including node lifecycle, driver rollouts, and capacity planning.
  • Own batch scheduling and multi-tenancy with queues, quotas, priorities, and fair-share allocation.
  • Design distributed storage systems for datasets and checkpoints with performance and fault-tolerance.
  • Harden platform security through access controls, network policies, and workload isolation.

About

Inside FAR AI

Visit site →

FAR AI aims to ensure AI systems are trustworthy and beneficial to society. They incubate and accelerate research agendas that are too resource-intensive for academia but not yet ready for commercialisation by industry.