9d ago

avatar

Zoox

Staff Software Engineer, HPC

$240K - $360K

Foster City, CA

Senior (10+ years)

AI / ML

Enterprise (1000+)

[object Object], ,[object Object], ,[object Object], ,[object Object], ,[object Object],[object Object], ,[object Object], ,[object Object], ,[object Object], ,[object Object], ,[object Object], ,[object Object]

Questions about the Staff Software Engineer, HPC role at Zoox

What key technical skills define success for this HPC platform role?

Success in this Staff Software Engineer role hinges on deep expertise in high-performance computing (HPC) infrastructure and distributed systems. Candidates must demonstrate proficiency in orchestrating large-scale compute, storage, and scheduling demands using industry-standard tools like SLURM, Kubernetes, and Ray.io. The role requires architecting scalable, reliable platforms that enhance developer velocity across complex AI and simulation workflows. Furthermore, technical success involves translating intricate autonomy and software workload requirements into robust infrastructure strategies. Candidates should possess strong system design capabilities to manage high-growth compute environments, enabling seamless integration between data engineering and machine learning pipelines. Ultimately, mastery of distributed computing frameworks and infrastructure-as-code is essential to empower Zoox’s autonomous vehicle development at scale.

How do you balance high-performance compute needs with developer velocity?

To balance high-performance compute (HPC) needs with developer velocity at Zoox, I focus on building a self-service, scalable infrastructure that abstracts complexity. By leveraging modern orchestrators like Ray.io, SLURM, and Kubernetes, I create a seamless interface where developers can access massive compute resources without getting bogged down by underlying infrastructure hurdles. My strategy involves implementing robust automation, comprehensive monitoring, and optimized resource scheduling to ensure high reliability. By working closely with AI and autonomy teams to understand their specific workloads, I can provide purpose-built platforms that reduce latency and friction. Ultimately, the goal is to create a "paved road" environment where engineers spend less time managing compute and more time driving innovation.

Which emerging HPC trends are most critical for scaling AI infrastructure?

To scale AI infrastructure effectively, as highlighted by the Zoox HPC role, the most critical emerging trends center on hybrid orchestration and elastic scalability. Leveraging platforms like Ray.io alongside traditional schedulers like SLURM and Kubernetes is essential to manage the massive, bursty compute demands of AI training and simulation. Furthermore, infrastructure abstraction is vital; engineering teams require a platform that hides the underlying complexity of high-performance storage and compute, prioritizing developer velocity. As AI models grow, automating resource allocation and ensuring reliable, low-latency access to distributed GPU clusters are paramount. Ultimately, the ability to modernize legacy HPC architectures to support seamless, large-scale autonomous vehicle development workflows defines current industry leadership.

How does the HPC team prioritize competing demands from autonomy groups?

The HPC team at Zoox prioritizes competing demands by acting as a strategic partner to the Autonomy and Software teams. Rather than reacting to individual requests, the Staff Software Engineer is tasked with defining the platform strategy by proactively engaging with stakeholders to understand their specific workload requirements, such as model training, data engineering, and simulation. By translating these diverse technical needs into a robust, scalable infrastructure based on technologies like Ray.io, SLURM, and Kubernetes, the team ensures that the HPC platform remains reliable and efficient. This collaborative approach allows the HPC team to balance velocity with infrastructure stability, directly impacting the productivity of all engineering teams company-wide.

How does your HPC strategy specifically accelerate Zoox's simulation goals?

Our HPC strategy accelerates Zoox’s simulation goals by providing a highly scalable, modernized infrastructure that supports the intense compute demands of autonomous vehicle development. By optimizing the integration of Ray.io, SLURM, and Kubernetes, we eliminate performance bottlenecks, ensuring rapid execution of complex simulation workloads. This platform serves as the foundational backbone for our AI and perception models, directly improving developer velocity and iteration cycles. By translating evolving workload requirements into robust, reliable infrastructure, we empower engineering teams to run simulations at scale, faster and more efficiently. Ultimately, our strategy transforms raw compute power into actionable insights, enabling the rapid development and validation of our fully autonomous vehicle fleet.