Application Guide

How to Apply for Software Engineer, ML Platform (ML Training)

at Zoox

๐Ÿข About Zoox

Zoox is reimagining urban mobility from the ground up, designing a fully autonomous, electric vehicle fleet for dense cities. Unlike many self-driving companies, Zoox is building a purpose-built vehicle without a steering wheel, prioritizing passenger experience and safety. Working here means contributing to a moonshot project that could transform transportation, with a culture that values deep technical innovation and cross-functional collaboration.

About This Role

As a Software Engineer on the ML Training team, you will own the core training infrastructure that powers all machine learning at Zooxโ€”from perception to planning. Your work will directly impact how quickly and reliably models are trained, validated, and deployed, enabling breakthroughs in autonomous driving. This role is critical because the ML platform must scale to handle petabytes of data and complex distributed training jobs while maintaining high uptime.

๐Ÿ’ก A Day in the Life

A typical day might start with a stand-up with your team to triage training pipeline issues, then dive into coding a new feature for the training framework, like adding support for gradient checkpointing. You might pair with a researcher to debug a model that's not converging due to data loading bottlenecks, then spend the afternoon monitoring cluster performance and optimizing resource allocation on AWS. The day ends with documenting architectural decisions and reviewing a colleague's pull request for a new validation pipeline.

๐ŸŽฏ Who Zoox Is Looking For

  • Has 2+ years building ML infrastructure, with hands-on experience optimizing distributed training using PyTorch, DeepSpeed, or JAX for large-scale models.
  • Is comfortable operating cloud-native infrastructure (AWS, Kubernetes, Docker) and can design fault-tolerant training pipelines that minimize cost and maximize throughput.
  • Understands the full ML lifecycleโ€”from data preprocessing and training to serving and monitoringโ€”and can debug performance bottlenecks in distributed systems.
  • Thrives in a collaborative environment, working closely with ML researchers to translate experimental code into robust, production-grade training workflows.

๐Ÿ“ Tips for Applying to Zoox

1

Highlight specific examples of scaling training workloads (e.g., multi-GPU, multi-node) using frameworks like PyTorch DDP or DeepSpeed, and quantify improvements (e.g., reduced training time by 40%).

2

Mention any experience with AWS services relevant to ML (SageMaker, EKS, S3, EFS) and how you optimized for cost or performance.

3

Show that you understand the unique challenges of autonomous vehicle dataโ€”like handling large, heterogeneous datasets and ensuring reproducibility.

4

Tailor your resume to emphasize platform engineering over model development; use keywords like 'distributed training', 'ML pipeline', 'infrastructure as code', and 'monitoring'.

5

Include a brief note in your cover letter about why Zoox's mission resonates with you and how your past work aligns with building reliable, scalable training systems.

โœ‰๏ธ What to Emphasize in Your Cover Letter

['Your experience building and maintaining ML training platforms at scale, with concrete metrics or outcomes.', 'Your comfort with cross-functional collaboration, especially translating researcher needs into robust infrastructure.', "Your familiarity with the specific tools listed (PyTorch, DeepSpeed, Ray, JAX) and how you've used them in production.", "A genuine interest in autonomous driving and Zoox's unique approach (purpose-built vehicle, zero-emission fleet)."]

Generate Cover Letter โ†’

๐Ÿ” Research Before Applying

To stand out, make sure you've researched:

  • โ†’ Read Zoox's engineering blog or tech talks about their ML infrastructure and autonomous driving stack.
  • โ†’ Understand Zoox's vehicle design and how it differs from other AV companies (e.g., no steering wheel, bidirectional driving).
  • โ†’ Familiarize yourself with Zoox's approach to safety and validation of ML models in safety-critical systems.
  • โ†’ Check recent news about Zoox's testing permits, partnerships, or funding to show you're up-to-date.

๐Ÿ’ฌ Prepare for These Interview Topics

Based on this role, you may be asked about:

1 Design a distributed training system for a large model (e.g., 100B parameters) on AWS, considering data parallelism, model parallelism, and fault tolerance.
2 Debug a scenario where training is slow despite high GPU utilization; how would you identify and fix bottlenecks?
3 How would you ensure reproducibility of training runs when using dynamic data pipelines and cloud spot instances?
4 Describe a time you had to balance researcher velocity with platform stability; what trade-offs did you make?
5 What metrics would you use to monitor the health of a training cluster, and how would you set up alerts?
Practice Interview Questions โ†’

โš ๏ธ Common Mistakes to Avoid

  • Focusing too much on model architecture or accuracy improvements rather than infrastructure and platform work.
  • Not demonstrating experience with cloud and distributed systemsโ€”this role is heavy on AWS and scaling.
  • Ignoring the safety-critical nature of autonomous driving; avoid suggesting untested or risky approaches.

๐Ÿ“… Application Timeline

This position is open until filled. However, we recommend applying as soon as possible as roles at mission-driven organizations tend to fill quickly.

Typical hiring timeline:

1

Application Review

1-2 weeks

2

Initial Screening

Phone call or written assessment

3

Interviews

1-2 rounds, usually virtual

โœ“

Offer

Congratulations!

Ready to Apply?

Good luck with your application to Zoox!