Institution profile

Zoox Inc.

Industry researchnorthamerica · us
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

May 13, 2026

This work addresses the high inference latency of Vision Transformers in multi-view 3D object detection, which stems from dense processing of 2D image tokens and 3D queries, and the inability of existing sparse methods to jointly optimize both. The authors propose a correlation-aligned sparsification framework that, for the first time, enables joint sparsification of 2D tokens and 3D queries within a ViT architecture. Their approach employs a 2D–3D cross-correlation head to dynamically assess relevance and co-select critical tokens and queries, complemented by a feature caching and reactivation mechanism to reuse filtered features. Evaluated on nuScenes and a newly introduced nuScenes-Relevance benchmark, the method achieves up to 3× speedup with only marginal accuracy degradation, substantially improving inference efficiency and enabling scalable, real-time 3D detection.

0 citationsRead paper

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

Apr 06, 2026

Existing diffusion language models struggle to effectively improve generation quality at test time through increased inference computation, primarily because conventional best-of-K sampling is constrained by repeatedly drawing from the same distribution. This work proposes Hierarchical Scaling Search ($S^3$), which introduces verifier-guided, trajectory-level search into diffusion language models for the first time. During the denoising process, $S^3$ dynamically expands multiple candidate trajectories and employs a lightweight, reference-free verifier to evaluate and selectively resample high-potential paths while preserving diversity. Notably, this approach approximates a reward-weighted distribution without modifying the base model or decoding schedule. Experiments demonstrate that $S^3$ significantly enhances performance on mathematical reasoning benchmarks such as MATH-500 and GSM8K when applied to LLaDA-8B-Instruct, validating the efficacy of test-time scaling.

0 citationsRead paper

Continuous-Utility Direct Preference Optimization

Jan 31, 2026

Traditional binary preference supervision struggles to capture fine-grained quality in reasoning processes, limiting the alignment efficacy of large language models on complex reasoning tasks. This work proposes CU-DPO, a novel framework that replaces binary labels with continuous utility scores to enable fine-grained alignment across diverse prompt-driven cognitive strategies. The approach employs a two-stage decoupled training procedure: first selecting strategies via a best-vs-all mechanism, then refining strategy execution through margin-stratified contrastive learning combined with entropy regularization. Theoretical analysis reveals a sample complexity improvement of Θ(K log K). Empirical results demonstrate that, across seven base models, strategy selection accuracy improves from 35–46% to 68–78%, with in-distribution mathematical reasoning performance gaining up to 6.6 points and strong generalization observed on out-of-distribution tasks.

0 citationsRead paper
Recent publications

Latest Papers

SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

May 13, 2026

This work addresses the high inference latency of Vision Transformers in multi-view 3D object detection, which stems from dense processing of 2D image tokens and 3D queries, and the inability of existing sparse methods to jointly optimize both. The authors propose a correlation-aligned sparsification framework that, for the first time, enables joint sparsification of 2D tokens and 3D queries within a ViT architecture. Their approach employs a 2D–3D cross-correlation head to dynamically assess relevance and co-select critical tokens and queries, complemented by a feature caching and reactivation mechanism to reuse filtered features. Evaluated on nuScenes and a newly introduced nuScenes-Relevance benchmark, the method achieves up to 3× speedup with only marginal accuracy degradation, substantially improving inference efficiency and enabling scalable, real-time 3D detection.

0 citationsRead paper

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

Apr 06, 2026

Existing diffusion language models struggle to effectively improve generation quality at test time through increased inference computation, primarily because conventional best-of-K sampling is constrained by repeatedly drawing from the same distribution. This work proposes Hierarchical Scaling Search ($S^3$), which introduces verifier-guided, trajectory-level search into diffusion language models for the first time. During the denoising process, $S^3$ dynamically expands multiple candidate trajectories and employs a lightweight, reference-free verifier to evaluate and selectively resample high-potential paths while preserving diversity. Notably, this approach approximates a reward-weighted distribution without modifying the base model or decoding schedule. Experiments demonstrate that $S^3$ significantly enhances performance on mathematical reasoning benchmarks such as MATH-500 and GSM8K when applied to LLaDA-8B-Instruct, validating the efficacy of test-time scaling.

0 citationsRead paper

Continuous-Utility Direct Preference Optimization

Jan 31, 2026

Traditional binary preference supervision struggles to capture fine-grained quality in reasoning processes, limiting the alignment efficacy of large language models on complex reasoning tasks. This work proposes CU-DPO, a novel framework that replaces binary labels with continuous utility scores to enable fine-grained alignment across diverse prompt-driven cognitive strategies. The approach employs a two-stage decoupled training procedure: first selecting strategies via a best-vs-all mechanism, then refining strategy execution through margin-stratified contrastive learning combined with entropy regularization. Theoretical analysis reveals a sample complexity improvement of Θ(K log K). Empirical results demonstrate that, across seven base models, strategy selection accuracy improves from 35–46% to 68–78%, with in-distribution mathematical reasoning performance gaining up to 6.6 points and strong generalization observed on out-of-distribution tasks.

0 citationsRead paper