Institution profile

Seeing Machines

Industry researchaustralasia · au
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Gromov Wasserstein Optimal Transport for Semantic Correspondences

Feb 03, 2026

This work proposes an efficient and highly accurate semantic correspondence method that addresses the high computational cost and reliance on multi-model ensembles in existing approaches. By leveraging features extracted from DINOv2, the method introduces a Gromov–Wasserstein optimal transport framework augmented with a spatial smoothness prior to explicitly model geometric consistency. This approach eliminates the need for complex models such as Stable Diffusion while maintaining an end-to-end pipeline. It substantially outperforms the DINOv2 baseline and achieves performance on par with or superior to current state-of-the-art methods that depend on model ensembles, all while offering a 5–10× speedup in inference time.

0 citationsRead paper

Small Talk, Big Impact? LLM-based Conversational Agents to Mitigate Passive Fatigue in Conditional Automated Driving

Oct 29, 2025

Passive fatigue in conditional automation poses safety risks by reducing driver alertness and degrading monitoring performance. Method: We propose a Large Language Model (LLM)-driven Context-Aware dialogue Agent (CA) that actively reorients drivers’ attention to the driving environment via natural language interaction. Our approach integrates on-road testing, multimodal data collection (in-vehicle video, drowsiness ratings, user interviews), and user-preference clustering to identify three prototypical driver personas, enabling adaptive dialogue design. Results: Experiments demonstrate that the CA significantly enhances driver alertness and sustains takeover readiness; users consistently affirm its monitoring-assistance value and exhibit high acceptance and usage intent. This work constitutes the first application of an LLM-based dynamic dialogue intervention system for in-vehicle passive fatigue mitigation, providing a scalable methodology and empirical foundation for intelligent human–machine shared-driving interface design.

0 citationsRead paper

DiSA: Diffusion Step Annealing in Autoregressive Image Generation

May 26, 2025

Autoregressive image generation via diffusion models suffers from high inference latency due to excessive denoising steps (50–100 per token). Method: This paper identifies, for the first time, that token distributions concentrate and denoising trajectories become increasingly linear in later generation stages. Leveraging this insight, we propose DiSA—a training-free, dynamic step annealing mechanism that adaptively schedules denoising steps via MLP-based prediction, variance estimation, and denoising trajectory monitoring. DiSA is orthogonal to existing diffusion acceleration techniques. Results: DiSA achieves 5–10× speedup on MAR and Harmon, and 1.4–2.5× on FlowAR and xAR, with no degradation in generation quality. It requires only a few lines of code for integration. Our core contribution is the discovery of the evolutionary规律 of denoising paths in autoregressive generation and the design of the first lightweight, training-free, plug-and-play diffusion step adaptation strategy.

0 citationsRead paper

Flexible Geometric Guidance for Probabilistic Human Pose Estimation with Diffusion Models

May 26, 2025IEEE International Conference on Automatic Face & Gesture Recognition

This work addresses the inherent ambiguities in 3D human pose estimation from 2D images—particularly depth uncertainty and occlusion-induced solution space indeterminacy—and overcomes the limited generalization of existing approaches that rely heavily on paired 2D-3D training data. The authors propose a novel paradigm based on an unconditional diffusion model, introducing for the first time a geometry-guided mechanism that leverages gradients from 2D keypoint heatmaps to steer the diffusion process. This enables the generation of multiple plausible, input-consistent 3D poses without requiring any paired 2D-3D training data. The method naturally supports multi-hypothesis prediction and pose completion without retraining conditional models. It achieves state-of-the-art performance among unpaired-data methods on Human3.6M and demonstrates exceptional generalization on MPI-INF-3DHP and 3DPW benchmarks.

0 citationsRead paper
Recent publications

Latest Papers

Gromov Wasserstein Optimal Transport for Semantic Correspondences

Feb 03, 2026

This work proposes an efficient and highly accurate semantic correspondence method that addresses the high computational cost and reliance on multi-model ensembles in existing approaches. By leveraging features extracted from DINOv2, the method introduces a Gromov–Wasserstein optimal transport framework augmented with a spatial smoothness prior to explicitly model geometric consistency. This approach eliminates the need for complex models such as Stable Diffusion while maintaining an end-to-end pipeline. It substantially outperforms the DINOv2 baseline and achieves performance on par with or superior to current state-of-the-art methods that depend on model ensembles, all while offering a 5–10× speedup in inference time.

0 citationsRead paper

Small Talk, Big Impact? LLM-based Conversational Agents to Mitigate Passive Fatigue in Conditional Automated Driving

Oct 29, 2025

Passive fatigue in conditional automation poses safety risks by reducing driver alertness and degrading monitoring performance. Method: We propose a Large Language Model (LLM)-driven Context-Aware dialogue Agent (CA) that actively reorients drivers’ attention to the driving environment via natural language interaction. Our approach integrates on-road testing, multimodal data collection (in-vehicle video, drowsiness ratings, user interviews), and user-preference clustering to identify three prototypical driver personas, enabling adaptive dialogue design. Results: Experiments demonstrate that the CA significantly enhances driver alertness and sustains takeover readiness; users consistently affirm its monitoring-assistance value and exhibit high acceptance and usage intent. This work constitutes the first application of an LLM-based dynamic dialogue intervention system for in-vehicle passive fatigue mitigation, providing a scalable methodology and empirical foundation for intelligent human–machine shared-driving interface design.

0 citationsRead paper

DiSA: Diffusion Step Annealing in Autoregressive Image Generation

May 26, 2025

Autoregressive image generation via diffusion models suffers from high inference latency due to excessive denoising steps (50–100 per token). Method: This paper identifies, for the first time, that token distributions concentrate and denoising trajectories become increasingly linear in later generation stages. Leveraging this insight, we propose DiSA—a training-free, dynamic step annealing mechanism that adaptively schedules denoising steps via MLP-based prediction, variance estimation, and denoising trajectory monitoring. DiSA is orthogonal to existing diffusion acceleration techniques. Results: DiSA achieves 5–10× speedup on MAR and Harmon, and 1.4–2.5× on FlowAR and xAR, with no degradation in generation quality. It requires only a few lines of code for integration. Our core contribution is the discovery of the evolutionary规律 of denoising paths in autoregressive generation and the design of the first lightweight, training-free, plug-and-play diffusion step adaptation strategy.

0 citationsRead paper

Flexible Geometric Guidance for Probabilistic Human Pose Estimation with Diffusion Models

May 26, 2025IEEE International Conference on Automatic Face & Gesture Recognition

This work addresses the inherent ambiguities in 3D human pose estimation from 2D images—particularly depth uncertainty and occlusion-induced solution space indeterminacy—and overcomes the limited generalization of existing approaches that rely heavily on paired 2D-3D training data. The authors propose a novel paradigm based on an unconditional diffusion model, introducing for the first time a geometry-guided mechanism that leverages gradients from 2D keypoint heatmaps to steer the diffusion process. This enables the generation of multiple plausible, input-consistent 3D poses without requiring any paired 2D-3D training data. The method naturally supports multi-hypothesis prediction and pose completion without retraining conditional models. It achieves state-of-the-art performance among unpaired-data methods on Human3.6M and demonstrates exceptional generalization on MPI-INF-3DHP and 3DPW benchmarks.

0 citationsRead paper