Institution profile

Inceptio Technology

Industry researchasia · cn
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Jul 13, 2026

This work addresses the vulnerability of Vision-Language Agents (VLAs) in autonomous driving under multimodal adversarial attacks by organizing a challenge centered on DriveLM-style multi-view visual question answering. Participants are tasked with generating high-fidelity adversarial images and text perturbations with minimal textual distortion to mislead models into producing answers that deviate from reference responses, while evaluating transferability under both white-box and black-box settings. The study presents the first systematic assessment of multi-view multimodal attacks, introducing novel techniques such as QA-graph-guided budget allocation, feature-space optimization, and suffix-constrained textual perturbations. Key findings include the dominance of image-side attacks, the efficacy of scene-level optimization, and the susceptibility of layout-sensitive content, thereby establishing a benchmark for robustness evaluation and defense development in VLAs.

0 citationsRead paper

SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

Apr 15, 2026

This work addresses the challenge of 3D reconstruction of close-proximity human interactions from monocular video, which is often hindered by severe occlusions leading to pose ambiguity, temporal inconsistency, and erroneous spatial relationships. To tackle these issues, the authors propose the first diffusion-based framework that integrates semantic guidance with geometric constraints. The approach first leverages a vision-language model to generate high-level interaction semantics that guide motion inpainting, then employs a sequence-level temporal optimizer incorporating contact-aware geometric constraints to ensure smooth and physically plausible reconstructions. Evaluated on multiple interaction benchmarks, the method significantly outperforms existing approaches and demonstrates strong generalization capabilities on both unseen datasets and real-world scenarios.

0 citationsRead paper

Improvement on LiDAR-Camera Calibration Using Square Targets

Jun 23, 2025

To address the poor robustness and deployment difficulty of LiDAR–camera extrinsic calibration in mass production and after-sales scenarios for autonomous driving, this paper proposes a fully automatic calibration method based on square calibration targets. The method introduces a purely geometry-driven, multi-stage target detection pipeline, a hierarchical coarse-search mechanism insensitive to initial pose errors, and a direct optimization algorithm incorporating photometric consistency constraints—collectively enhancing robustness against sensor noise, sparse or incomplete point clouds, and large initial misalignments. Crucially, it requires no specialized retroreflective materials and achieves rapid (<1 minute), high-precision calibration (rotational error <0.1°, translational error <2 mm). Extensive experiments validate its stability and deployability in real-world manufacturing lines and after-sales service environments, demonstrating strong scalability for large-scale deployment of multi-sensor systems in production settings.

0 citationsRead paper

Fine-Grained Controllable Apparel Showcase Image Generation via Garment-Centric Outpainting

Mar 03, 2025

This work addresses fine-grained controllable fashion image generation by proposing a garment-centric diffusion outpainting method that jointly synthesizes high-fidelity outfit display images from an input garment image, text prompts, and a face image—without explicit cloth deformation modeling. The method builds upon a latent diffusion model (LDM) to construct a multi-condition controllable generation framework. Its core contributions are: (1) a novel garment-adaptive pose prediction module enabling geometric alignment between garment and human body; (2) a multi-scale appearance customization module (MS-ACM) supporting fine-grained textual control over global style and local texture; and (3) a lightweight cross-modal feature fusion mechanism that eliminates the need for auxiliary encoders. Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in garment detail fidelity, text–vision alignment accuracy, and controllability, exhibiting strong potential for commercial deployment.

0 citationsRead paper

Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

Jan 01, 2025

This work addresses test-time adaptation (TTA), investigating how structured semantic priors implicitly encoded in diffusion model score functions can enhance the out-of-distribution generalization of pre-trained discriminative models. We first reveal that diffusion scores inherently contain transferable, structured semantic information—previously unrecognized in TTA contexts. To leverage this, we propose a single-step denoising knowledge extraction mechanism, circumventing the high computational cost of multi-step Monte Carlo score estimation. Our approach integrates score-matching-driven semantic prior modeling with lightweight feature distillation, enabling efficient TTA without model fine-tuning or access to source-domain data. Extensive experiments on diverse image classification and dense prediction benchmarks demonstrate substantial improvements over state-of-the-art TTA methods. Ablation studies comprehensively validate the effectiveness and necessity of each component, confirming that structured semantic priors from diffusion scores significantly boost robustness under distribution shift.

0 citationsRead paper
Recent publications

Latest Papers

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Jul 13, 2026

This work addresses the vulnerability of Vision-Language Agents (VLAs) in autonomous driving under multimodal adversarial attacks by organizing a challenge centered on DriveLM-style multi-view visual question answering. Participants are tasked with generating high-fidelity adversarial images and text perturbations with minimal textual distortion to mislead models into producing answers that deviate from reference responses, while evaluating transferability under both white-box and black-box settings. The study presents the first systematic assessment of multi-view multimodal attacks, introducing novel techniques such as QA-graph-guided budget allocation, feature-space optimization, and suffix-constrained textual perturbations. Key findings include the dominance of image-side attacks, the efficacy of scene-level optimization, and the susceptibility of layout-sensitive content, thereby establishing a benchmark for robustness evaluation and defense development in VLAs.

0 citationsRead paper

SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

Apr 15, 2026

This work addresses the challenge of 3D reconstruction of close-proximity human interactions from monocular video, which is often hindered by severe occlusions leading to pose ambiguity, temporal inconsistency, and erroneous spatial relationships. To tackle these issues, the authors propose the first diffusion-based framework that integrates semantic guidance with geometric constraints. The approach first leverages a vision-language model to generate high-level interaction semantics that guide motion inpainting, then employs a sequence-level temporal optimizer incorporating contact-aware geometric constraints to ensure smooth and physically plausible reconstructions. Evaluated on multiple interaction benchmarks, the method significantly outperforms existing approaches and demonstrates strong generalization capabilities on both unseen datasets and real-world scenarios.

0 citationsRead paper

Improvement on LiDAR-Camera Calibration Using Square Targets

Jun 23, 2025

To address the poor robustness and deployment difficulty of LiDAR–camera extrinsic calibration in mass production and after-sales scenarios for autonomous driving, this paper proposes a fully automatic calibration method based on square calibration targets. The method introduces a purely geometry-driven, multi-stage target detection pipeline, a hierarchical coarse-search mechanism insensitive to initial pose errors, and a direct optimization algorithm incorporating photometric consistency constraints—collectively enhancing robustness against sensor noise, sparse or incomplete point clouds, and large initial misalignments. Crucially, it requires no specialized retroreflective materials and achieves rapid (<1 minute), high-precision calibration (rotational error <0.1°, translational error <2 mm). Extensive experiments validate its stability and deployability in real-world manufacturing lines and after-sales service environments, demonstrating strong scalability for large-scale deployment of multi-sensor systems in production settings.

0 citationsRead paper

Fine-Grained Controllable Apparel Showcase Image Generation via Garment-Centric Outpainting

Mar 03, 2025

This work addresses fine-grained controllable fashion image generation by proposing a garment-centric diffusion outpainting method that jointly synthesizes high-fidelity outfit display images from an input garment image, text prompts, and a face image—without explicit cloth deformation modeling. The method builds upon a latent diffusion model (LDM) to construct a multi-condition controllable generation framework. Its core contributions are: (1) a novel garment-adaptive pose prediction module enabling geometric alignment between garment and human body; (2) a multi-scale appearance customization module (MS-ACM) supporting fine-grained textual control over global style and local texture; and (3) a lightweight cross-modal feature fusion mechanism that eliminates the need for auxiliary encoders. Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in garment detail fidelity, text–vision alignment accuracy, and controllability, exhibiting strong potential for commercial deployment.

0 citationsRead paper

Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

Jan 01, 2025

This work addresses test-time adaptation (TTA), investigating how structured semantic priors implicitly encoded in diffusion model score functions can enhance the out-of-distribution generalization of pre-trained discriminative models. We first reveal that diffusion scores inherently contain transferable, structured semantic information—previously unrecognized in TTA contexts. To leverage this, we propose a single-step denoising knowledge extraction mechanism, circumventing the high computational cost of multi-step Monte Carlo score estimation. Our approach integrates score-matching-driven semantic prior modeling with lightweight feature distillation, enabling efficient TTA without model fine-tuning or access to source-domain data. Extensive experiments on diverse image classification and dense prediction benchmarks demonstrate substantial improvements over state-of-the-art TTA methods. Ablation studies comprehensively validate the effectiveness and necessity of each component, confirming that structured semantic priors from diffusion scores significantly boost robustness under distribution shift.

0 citationsRead paper