AI Safety & Governance Full-time

Senior Software Engineer, GPU Cluster Infrastructure

FAR.AI

Posted

Sep 14, 2026

Location

Remote

Type

Full-time

Mission

What you will drive

Build and operate large-scale GPU cluster infrastructure for AI safety research, managing Kubernetes fleets, storage, scheduling, and security.
- Operate Kubernetes GPU clusters day-to-day: node lifecycle, upgrades, capacity planning, and staged rollouts with safe rollback.
- Own batch scheduling, multi-tenancy, quotas, priorities, and fair-share allocation across research teams.
- Design and manage distributed storage for datasets, checkpoints, and backups; ensure multi-node training fault tolerance and NCCL debugging.
- Harden platform security: identity, access control, network policy, workload isolation, and sandboxing for autonomous agents.
- Requires 3+ years in production Linux systems or infrastructure engineering on GPU, HPC, or batch platforms; strong Kubernetes and infrastructure-as-code experience.

Profile

What makes you a great fit

Build and operate large-scale GPU cluster infrastructure for AI safety research, managing Kubernetes fleets, storage, scheduling, and security.
- Operate Kubernetes GPU clusters day-to-day: node lifecycle, upgrades, capacity planning, and staged rollouts with safe rollback.
- Own batch scheduling, multi-tenancy, quotas, priorities, and fair-share allocation across research teams.
- Design and manage distributed storage for datasets, checkpoints, and backups; ensure multi-node training fault tolerance and NCCL debugging.
- Harden platform security: identity, access control, network policy, workload isolation, and sandboxing for autonomous agents.
- Requires 3+ years in production Linux systems or infrastructure engineering on GPU, HPC, or batch platforms; strong Kubernetes and infrastructure-as-code experience.

About

Inside FAR.AI

Visit site →

FAR.AI is an AI safety research nonprofit focused on addressing risks from advanced AI systems.