2mo ago

avatar

Waymo

Senior Staff Machine Learning Engineer, LLM/VLM Model Architecture & Optimization

$298K - $368K

Mountain View, CA

Senior (10+ years)

AI / ML

Enterprise (1000+)

[object Object],[object Object],[object Object], ,[object Object],[object Object], ,[object Object],[object Object], ,[object Object],[object Object], ,[object Object],[object Object], ,[object Object],[object Object], ,[object Object],[object Object],[object Object]

Questions about the Senior Staff Machine Learning Engineer, LLM/VLM Model Architecture & Optimization role at Waymo

What key skills ensure success in large-scale LLM/VLM model roles?

Success in large-scale LLM/VLM roles hinges on seven years of experience in machine learning, with a focus on foundation models like LLMs and VLMs. Key technical skills include mastery of deep learning frameworks (PyTorch, JAX), expertise in low-latency on-device inference, and deep understanding of hardware acceleration. Professionals must design model architectures that align with hardware constraints and optimize performance for memory, power, and compute-limited environments. Soft skills like operating effectively under ambiguity and collaborating across research, software, and hardware teams are critical. Additionally, experience applying these models in safety-critical domains, such as autonomy or robotics, ensures the ability to handle complex real-world challenges[1][2].

Which frameworks and tools best support efficient on-device model optimization?

The frameworks and tools best supporting efficient on-device model optimization are TensorFlow Lite, PyTorch Mobile, NCNN, OpenVINO, and ONNX Runtime, which provide optimized implementations to reduce memory and power consumption [1]. LiteRT (with the TensorFlow Model Optimization Toolkit) specifically enables quantization, pruning, and clustering to shrink model size and accelerate inference [5]. For hardware-specific performance, NVIDIA TensorRT is highly recommended [10]. Additionally, emerging frameworks like Apple’s AXLearn (built on JAX and XLA) facilitate efficient on-device training and inference for foundation models [2]. These tools collectively address latency, memory, and compute constraints in edge environments [1][4].

What industry challenges most impact safety-critical ML applications today?

The most critical safety challenges impacting ML applications today are data quality, model generalizability, and explainability. Biases and erroneous data collection frequently undermine trustworthiness, while poor generalizability limits reliability in real-world scenarios like autonomy. Crucially, the inherent lack of interpretability in complex models hinders regulatory approval and user trust, making it difficult to verify safety. Furthermore, uncertainty quantification and security vulnerabilities pose significant risks, as models may behave unpredictably or be exploited. Ensuring continuous AI assurance through rigorous verification, validation, and runtime monitoring remains essential to mitigate residual risks in safety-critical domains. [1][2][3]

How does Waymo integrate LLM/VLM research into autonomous driving systems?

Waymo integrates LLM/VLM research into autonomous driving primarily through its End-to-End Multimodal Model (EMMA), which leverages Gemini, a multimodal large language model by Google. EMMA employs a unified transformer architecture to process raw sensor data—cameras, LiDAR, radar, and maps—alongside textual inputs, generating driving outputs like trajectories and object detections directly in a natural language space. This enables chain-of-thought reasoning for complex decision-making, improving planning performance by 6.7%. Additionally, Waymo uses a Driving VLM trained on its own data to handle rare, safety-critical scenarios (e.g., a burning vehicle ahead) by providing semantic cues that guide safer navigation, bridging linguistic understanding with physical control.

[1][6][7][4]

What growth strategies does Waymo prioritize for perception and model innovation?

Waymo prioritizes growth in perception and model innovation by developing the Waymo Foundation Model, which integrates advanced Large Language Models (LLMs) and Vision-Language Models (VLMs) with its proprietary driving experience and AV-specific AI. This architecture enables superior scene interpretation, driving plan generation, and agent trajectory prediction. The company focuses on continuous learning from its massive real-world dataset—over 100 million miles driven—to refine models for safety-critical domains. Additionally, Waymo enhances closed-loop simulation capabilities to generate realistic future states and road user behaviors, ensuring robust model generalization for domestic and international expansion while optimizing performance for on-device hardware constraints.

[5] [2] [4]