Institution profile

Target AI

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

Mar 30, 2026

Current evaluations of vision-language models are largely confined to single-image, single-turn tasks, which inadequately assess their capacity for object identification and reasoning in extended, multi-image interactive settings. To address this gap, this work introduces AMIGO—the first benchmark for multi-image, multi-turn visual grounding—where an agent must locate a hidden target within a visually similar image gallery by engaging in an attribute-guided sequence of Yes/No/Unsure questions. The benchmark enforces a strict interaction protocol, incorporates a Skip penalty mechanism, and provides controllable noisy feedback, enabling comprehensive evaluation of questioning strategies, cross-turn constraint tracking, and fine-grained discriminative ability. Experiments on the Guess My Preferred Dress task systematically measure model performance across identification accuracy, evidence verification, efficiency, protocol adherence, noise robustness, and interaction trajectory quality.

0 citationsRead paper

Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval

Mar 05, 2026

This work addresses the limitations of existing e-commerce retrieval systems, which predominantly rely on textual information and struggle to effectively incorporate visual semantics from product images, thereby constraining cross-modal representation capabilities. To overcome this, the authors propose a two-stage alignment strategy tailored for e-commerce scenarios and introduce a novel vision-language fusion network that jointly optimizes multimodal representations of queries and products within a dual-tower architecture. By integrating domain-adaptive fine-tuning with a cross-modal alignment mechanism, the approach significantly enhances semantic complementarity between text and image modalities. Extensive experiments on a large-scale real-world e-commerce dataset demonstrate that the proposed method substantially outperforms text-only baselines and alternative multimodal fusion approaches, confirming its effectiveness and practical applicability.

0 citationsRead paper

Unified Learning-to-Rank for Multi-Channel Retrieval in Large-Scale E-Commerce Search

Feb 26, 2026

This work addresses the challenge of effectively fusing multiple heterogeneous retrieval channels under strict latency constraints to optimize business metrics such as user conversion. We propose a channel-aware unified learning-to-rank framework that formulates multi-channel result fusion as a query-dependent multi-objective ranking problem, jointly optimizing for click-through, add-to-cart, and purchase outcomes. The approach explicitly incorporates channel-specific signals and users’ short-term behavioral sequences, and leverages query-adaptive fusion strategies alongside cross-channel interaction modeling to overcome the limitations of conventional fixed-weight fusion methods. Online A/B experiments demonstrate that the system achieves a 2.85% improvement in user conversion rate while maintaining a p95 latency below 50 milliseconds, and has been successfully deployed in the production environment of Target.com.

0 citationsRead paper

Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks

Sep 09, 2025

This study systematically evaluates the impact of six time-frequency representations—Mel spectrograms, MFCCs, STFT chromagrams, CQT chromagrams, CENS chromagrams, and cyclic tempograms—on environmental audio classification performance under a unified deep CNN architecture. Using the ESC-50 dataset, all features are trained end-to-end under identical experimental conditions, and their performance is rigorously compared across both coarse-grained (category-level) and fine-grained classification tasks using accuracy, precision, recall, and F1-score. Results demonstrate that Mel spectrograms and MFCCs significantly outperform the other representations, achieving top overall accuracies of 86.2% and 85.7%, respectively—confirming their robustness and discriminative power for modeling environmental acoustic structure. To our knowledge, this is the first work to conduct a controlled, cross-representation benchmark under consistent network architecture and training protocol. The findings provide reproducible, empirical guidance for feature selection in environmental audio analysis.

0 citationsRead paper
Recent publications

Latest Papers

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

Mar 30, 2026

Current evaluations of vision-language models are largely confined to single-image, single-turn tasks, which inadequately assess their capacity for object identification and reasoning in extended, multi-image interactive settings. To address this gap, this work introduces AMIGO—the first benchmark for multi-image, multi-turn visual grounding—where an agent must locate a hidden target within a visually similar image gallery by engaging in an attribute-guided sequence of Yes/No/Unsure questions. The benchmark enforces a strict interaction protocol, incorporates a Skip penalty mechanism, and provides controllable noisy feedback, enabling comprehensive evaluation of questioning strategies, cross-turn constraint tracking, and fine-grained discriminative ability. Experiments on the Guess My Preferred Dress task systematically measure model performance across identification accuracy, evidence verification, efficiency, protocol adherence, noise robustness, and interaction trajectory quality.

0 citationsRead paper

Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval

Mar 05, 2026

This work addresses the limitations of existing e-commerce retrieval systems, which predominantly rely on textual information and struggle to effectively incorporate visual semantics from product images, thereby constraining cross-modal representation capabilities. To overcome this, the authors propose a two-stage alignment strategy tailored for e-commerce scenarios and introduce a novel vision-language fusion network that jointly optimizes multimodal representations of queries and products within a dual-tower architecture. By integrating domain-adaptive fine-tuning with a cross-modal alignment mechanism, the approach significantly enhances semantic complementarity between text and image modalities. Extensive experiments on a large-scale real-world e-commerce dataset demonstrate that the proposed method substantially outperforms text-only baselines and alternative multimodal fusion approaches, confirming its effectiveness and practical applicability.

0 citationsRead paper

Unified Learning-to-Rank for Multi-Channel Retrieval in Large-Scale E-Commerce Search

Feb 26, 2026

This work addresses the challenge of effectively fusing multiple heterogeneous retrieval channels under strict latency constraints to optimize business metrics such as user conversion. We propose a channel-aware unified learning-to-rank framework that formulates multi-channel result fusion as a query-dependent multi-objective ranking problem, jointly optimizing for click-through, add-to-cart, and purchase outcomes. The approach explicitly incorporates channel-specific signals and users’ short-term behavioral sequences, and leverages query-adaptive fusion strategies alongside cross-channel interaction modeling to overcome the limitations of conventional fixed-weight fusion methods. Online A/B experiments demonstrate that the system achieves a 2.85% improvement in user conversion rate while maintaining a p95 latency below 50 milliseconds, and has been successfully deployed in the production environment of Target.com.

0 citationsRead paper

Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks

Sep 09, 2025

This study systematically evaluates the impact of six time-frequency representations—Mel spectrograms, MFCCs, STFT chromagrams, CQT chromagrams, CENS chromagrams, and cyclic tempograms—on environmental audio classification performance under a unified deep CNN architecture. Using the ESC-50 dataset, all features are trained end-to-end under identical experimental conditions, and their performance is rigorously compared across both coarse-grained (category-level) and fine-grained classification tasks using accuracy, precision, recall, and F1-score. Results demonstrate that Mel spectrograms and MFCCs significantly outperform the other representations, achieving top overall accuracies of 86.2% and 85.7%, respectively—confirming their robustness and discriminative power for modeling environmental acoustic structure. To our knowledge, this is the first work to conduct a controlled, cross-representation benchmark under consistent network architecture and training protocol. The findings provide reproducible, empirical guidance for feature selection in environmental audio analysis.

0 citationsRead paper