Institution profile

ObjectEye Inc.

Industry researchnorthamerica · us
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

Apr 09, 2026

Existing methods jointly optimize contrastive alignment and masked reconstruction objectives, which often introduces semantic noise and causes optimization interference, thereby limiting cross-modal representation learning performance. This work proposes the TG-DP framework, which decouples reconstruction and alignment tasks along separate optimization paths for the first time. Each path employs a visibility pattern tailored to its specific objective, and a teacher model is introduced to guide the organization of visible tokens in the contrastive path, reducing interference and enhancing representation quality. The proposed method achieves significant improvements in zero-shot retrieval on AudioSet—R@1 increases from 35.2% to 37.4% (video→audio) and from 27.9% to 37.1% (audio→video)—and attains state-of-the-art linear probe performance on both AS20K and VGGSound benchmarks.

0 citationsRead paper

GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation

Oct 08, 2025

Text-to-image generation often suffers from semantic inconsistency and missing details when processing complex, lengthy prompts. Existing approaches either require model fine-tuning or lack systematic error analysis and interpretable optimization mechanisms. This paper proposes a model-agnostic, interpretable test-time prompt optimization framework grounded in multi-agent collaboration. It integrates automated error diagnosis, clustering-driven adaptive exploration, fine-grained verification, and memory-augmented refinement to dynamically and iteratively correct input prompts. Crucially, no model fine-tuning is required—only prompt engineering and inference-time optimization are leveraged to significantly enhance generation quality. On DPG-bench and Geneval benchmarks, the method improves text–image alignment by 16.9% and 5.7%, respectively, while simultaneously boosting structural coherence and detail fidelity.

0 citationsRead paper

AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection

Aug 08, 2025

Existing anomaly detection methods are typically tailored to specific anomaly types (e.g., texture defects or logical inconsistencies), exhibiting limited generalization across domains. To address this, we propose AnomalyMoE—the first generic visual anomaly detection framework based on Mixture-of-Experts (MoE). It employs a three-level decoupled modeling strategy: local structural, component-level semantic, and global logical representations—enabling unified, cross-modal and cross-task detection. We introduce two novel mechanisms: Expert Information Repulsion to enhance expert diversity, and Expert Selection Balancing to improve expert utilization. Coupled with hierarchical feature reconstruction, the framework supports unsupervised anomaly localization and fine-grained classification. Extensive evaluation across eight heterogeneous benchmarks—including industrial images, 3D point clouds, medical imaging, video surveillance, and logical anomalies—demonstrates consistent superiority over domain-specific state-of-the-art methods, achieving new SOTA performance.

0 citationsRead paper

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

Aug 05, 2025

Existing anomaly synthesis methods suffer from discontinuous microstructures, coarse-grained semantic control, and low generation efficiency. To address these issues, we propose ARAS—a language-guided, locally autoregressive anomaly synthesis framework. ARAS enables text-driven, fine-grained anomaly injection via token-anchored latent editing and a hard-gated autoregressive operator. It further introduces a training-free masked sampling kernel and a dynamic Quality-Aware Re-weighting and Adaptation mechanism (QARAD) to enhance synthetic realism and detection robustness. The method integrates latent-space editing, dual-encoder vision-language similarity modeling, and efficient sampling strategies. Evaluated on MVTec AD, VisA, and BTAD, ARAS achieves state-of-the-art performance in both image-level and pixel-level anomaly detection—while operating five times faster than prior approaches—and significantly improves texture fidelity and semantic controllability.

0 citationsRead paper

FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition

Jun 28, 2025

In pedestrian attribute recognition (PAR), existing region-based feature methods suffer from loss of attribute-specific fine-grained patterns and poor generalization to unseen attributes. To address these limitations, we propose a semantic-guided adaptive attribute-level feature learning framework comprising three core components: (i) Multi-Granularity Mixed Tokens (MGMT) for cross-scale discriminative representation modeling; (ii) Attribute-Guided Visual Feature Extraction (AVFE) for fine-grained, attribute-specific feature learning; and (iii) Region-Aware Contrastive Learning (RACL) to enhance image–text alignment and interpretability. Our approach is the first to systematically support unified recognition of both seen and unseen attributes in PAR without requiring attribute label fine-tuning. Extensive experiments on PA100K, PETA, and RAPv1 demonstrate significant improvements in fine-grained accuracy and zero-shot generalization capability, achieving state-of-the-art performance and validating the effectiveness of the proposed framework.

0 citationsRead paper
Recent publications

Latest Papers

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

Apr 09, 2026

Existing methods jointly optimize contrastive alignment and masked reconstruction objectives, which often introduces semantic noise and causes optimization interference, thereby limiting cross-modal representation learning performance. This work proposes the TG-DP framework, which decouples reconstruction and alignment tasks along separate optimization paths for the first time. Each path employs a visibility pattern tailored to its specific objective, and a teacher model is introduced to guide the organization of visible tokens in the contrastive path, reducing interference and enhancing representation quality. The proposed method achieves significant improvements in zero-shot retrieval on AudioSet—R@1 increases from 35.2% to 37.4% (video→audio) and from 27.9% to 37.1% (audio→video)—and attains state-of-the-art linear probe performance on both AS20K and VGGSound benchmarks.

0 citationsRead paper

GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation

Oct 08, 2025

Text-to-image generation often suffers from semantic inconsistency and missing details when processing complex, lengthy prompts. Existing approaches either require model fine-tuning or lack systematic error analysis and interpretable optimization mechanisms. This paper proposes a model-agnostic, interpretable test-time prompt optimization framework grounded in multi-agent collaboration. It integrates automated error diagnosis, clustering-driven adaptive exploration, fine-grained verification, and memory-augmented refinement to dynamically and iteratively correct input prompts. Crucially, no model fine-tuning is required—only prompt engineering and inference-time optimization are leveraged to significantly enhance generation quality. On DPG-bench and Geneval benchmarks, the method improves text–image alignment by 16.9% and 5.7%, respectively, while simultaneously boosting structural coherence and detail fidelity.

0 citationsRead paper

AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection

Aug 08, 2025

Existing anomaly detection methods are typically tailored to specific anomaly types (e.g., texture defects or logical inconsistencies), exhibiting limited generalization across domains. To address this, we propose AnomalyMoE—the first generic visual anomaly detection framework based on Mixture-of-Experts (MoE). It employs a three-level decoupled modeling strategy: local structural, component-level semantic, and global logical representations—enabling unified, cross-modal and cross-task detection. We introduce two novel mechanisms: Expert Information Repulsion to enhance expert diversity, and Expert Selection Balancing to improve expert utilization. Coupled with hierarchical feature reconstruction, the framework supports unsupervised anomaly localization and fine-grained classification. Extensive evaluation across eight heterogeneous benchmarks—including industrial images, 3D point clouds, medical imaging, video surveillance, and logical anomalies—demonstrates consistent superiority over domain-specific state-of-the-art methods, achieving new SOTA performance.

0 citationsRead paper

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

Aug 05, 2025

Existing anomaly synthesis methods suffer from discontinuous microstructures, coarse-grained semantic control, and low generation efficiency. To address these issues, we propose ARAS—a language-guided, locally autoregressive anomaly synthesis framework. ARAS enables text-driven, fine-grained anomaly injection via token-anchored latent editing and a hard-gated autoregressive operator. It further introduces a training-free masked sampling kernel and a dynamic Quality-Aware Re-weighting and Adaptation mechanism (QARAD) to enhance synthetic realism and detection robustness. The method integrates latent-space editing, dual-encoder vision-language similarity modeling, and efficient sampling strategies. Evaluated on MVTec AD, VisA, and BTAD, ARAS achieves state-of-the-art performance in both image-level and pixel-level anomaly detection—while operating five times faster than prior approaches—and significantly improves texture fidelity and semantic controllability.

0 citationsRead paper

FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition

Jun 28, 2025

In pedestrian attribute recognition (PAR), existing region-based feature methods suffer from loss of attribute-specific fine-grained patterns and poor generalization to unseen attributes. To address these limitations, we propose a semantic-guided adaptive attribute-level feature learning framework comprising three core components: (i) Multi-Granularity Mixed Tokens (MGMT) for cross-scale discriminative representation modeling; (ii) Attribute-Guided Visual Feature Extraction (AVFE) for fine-grained, attribute-specific feature learning; and (iii) Region-Aware Contrastive Learning (RACL) to enhance image–text alignment and interpretability. Our approach is the first to systematically support unified recognition of both seen and unseen attributes in PAR without requiring attribute label fine-tuning. Extensive experiments on PA100K, PETA, and RAPv1 demonstrate significant improvements in fine-grained accuracy and zero-shot generalization capability, achieving state-of-the-art performance and validating the effectiveness of the proposed framework.

0 citationsRead paper