Institution profile

Shanghai Transsion Co., Ltd

Industry researchasia · cn
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

Aug 05, 2026

This work addresses the semantic channel gap in multimodal large language models, which struggle to effectively leverage visual text for reasoning about task instructions embedded in images—such as those found in screenshots. The study is the first to explicitly identify and quantify this gap, introducing an end-to-end prompt-region grounding method that operates without OCR or region-level metadata. By embedding task instructions directly into images, the authors construct a Visualized Task Semantics (VTS) benchmark and align question-relevant regions with their semantic meanings through masked image modeling and typed semantic representations, thereby recovering clean visual features from occluded views. Evaluated across four benchmarks, the proposed approach improves VTS accuracy from 58.0 to 66.3 (+8.3 percentage points) while preserving performance on original text-interface tasks.

0 citationsRead paper

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

Jul 04, 2026

This work addresses the computational and memory bottlenecks hindering efficient, cinematic camera-motion video generation (e.g., bullet time, dolly zoom) from images on mobile devices. The authors propose a three-stage optimization strategy: distillation-guided pruning to obtain a compact model, combined diffusion distillation and reinforcement learning to compress the generator into four denoising steps, and mixed post-training quantization to reduce the model size below 1 GB. Built upon the Diffusion Transformer architecture, this approach achieves the first on-device image-to-video diffusion model capable of generating cinematic camera motions. Compared to the Wan 2.1 teacher model, it offers a 40× speedup in inference, enabling generation of 49-frame 480p videos on a MediaTek Dimensity 8400 chipset with only 20 seconds per denoising step and a peak memory footprint of 1.8 GB, significantly advancing practical high-quality video synthesis on mobile platforms.

0 citationsRead paper

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

Jun 26, 2026

This work addresses the challenges of object removal in real-world scenarios, where modeling non-local effects and handling inaccurate user-provided masks remain difficult, while existing diffusion models incur prohibitive computational costs that hinder deployment on interactive or edge devices. The authors propose OSOR, a novel approach that achieves stable training within a single-step diffusion framework for the first time. OSOR integrates an occupancy-guided discriminator, a lightweight alpha head, and a Semantic Anchor Validation Pipeline (SAVP) to jointly optimize generation quality under imperfect masks. Evaluated on the large-scale CORNE dataset and the AnimeEraseBench/TextEraseBench benchmarks, OSOR surpasses multi-step diffusion baselines in perceptual quality while accelerating inference by 4–30×, substantially improving both efficiency and robustness for interactive and edge-based applications.

0 citationsRead paper

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

Jun 18, 2026

This work addresses the challenge in streaming zero-shot voice conversion of disentangling speaker identity from linguistic content, where existing approaches struggle to simultaneously suppress speaker leakage and preserve vocal expressiveness. The study introduces speaker anonymization into this task for the first time, proposing a novel perturbation mechanism that explicitly protects speaker identity while retaining prosodic information. Furthermore, it designs a strictly causal, non-lookahead generative network to enable truly zero-latency streaming conversion. Without relying on future-frame buffering, the method effectively balances speaker privacy and speech utility, significantly enhancing both the naturalness and real-time performance of converted speech.

0 citationsRead paper

An Effective Solution for the CVPR 2026 8th UG2+ Challenge Track 3: Dynamic Object Segmentation in Turbulence

May 30, 2026

This work addresses the challenge of dynamic object segmentation under severe geometric distortions and noise induced by atmospheric turbulence. Building upon the Segment Any Motion framework, the authors propose a domain adaptation strategy that leverages simulated turbulence perturbations during training, coupled with a spatiotemporal post-processing module designed to preserve small targets while enforcing label consistency across frames. The proposed approach effectively suppresses boundary artifacts and transient noise, achieving second place in Track 3 of the CVPR 2026 UG2+ Challenge. This demonstrates a significant improvement in both accuracy and robustness for segmenting moving objects in turbulent imaging conditions.

0 citationsRead paper
Recent publications

Latest Papers

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

Aug 05, 2026

This work addresses the semantic channel gap in multimodal large language models, which struggle to effectively leverage visual text for reasoning about task instructions embedded in images—such as those found in screenshots. The study is the first to explicitly identify and quantify this gap, introducing an end-to-end prompt-region grounding method that operates without OCR or region-level metadata. By embedding task instructions directly into images, the authors construct a Visualized Task Semantics (VTS) benchmark and align question-relevant regions with their semantic meanings through masked image modeling and typed semantic representations, thereby recovering clean visual features from occluded views. Evaluated across four benchmarks, the proposed approach improves VTS accuracy from 58.0 to 66.3 (+8.3 percentage points) while preserving performance on original text-interface tasks.

0 citationsRead paper

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

Jul 04, 2026

This work addresses the computational and memory bottlenecks hindering efficient, cinematic camera-motion video generation (e.g., bullet time, dolly zoom) from images on mobile devices. The authors propose a three-stage optimization strategy: distillation-guided pruning to obtain a compact model, combined diffusion distillation and reinforcement learning to compress the generator into four denoising steps, and mixed post-training quantization to reduce the model size below 1 GB. Built upon the Diffusion Transformer architecture, this approach achieves the first on-device image-to-video diffusion model capable of generating cinematic camera motions. Compared to the Wan 2.1 teacher model, it offers a 40× speedup in inference, enabling generation of 49-frame 480p videos on a MediaTek Dimensity 8400 chipset with only 20 seconds per denoising step and a peak memory footprint of 1.8 GB, significantly advancing practical high-quality video synthesis on mobile platforms.

0 citationsRead paper

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

Jun 26, 2026

This work addresses the challenges of object removal in real-world scenarios, where modeling non-local effects and handling inaccurate user-provided masks remain difficult, while existing diffusion models incur prohibitive computational costs that hinder deployment on interactive or edge devices. The authors propose OSOR, a novel approach that achieves stable training within a single-step diffusion framework for the first time. OSOR integrates an occupancy-guided discriminator, a lightweight alpha head, and a Semantic Anchor Validation Pipeline (SAVP) to jointly optimize generation quality under imperfect masks. Evaluated on the large-scale CORNE dataset and the AnimeEraseBench/TextEraseBench benchmarks, OSOR surpasses multi-step diffusion baselines in perceptual quality while accelerating inference by 4–30×, substantially improving both efficiency and robustness for interactive and edge-based applications.

0 citationsRead paper

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

Jun 18, 2026

This work addresses the challenge in streaming zero-shot voice conversion of disentangling speaker identity from linguistic content, where existing approaches struggle to simultaneously suppress speaker leakage and preserve vocal expressiveness. The study introduces speaker anonymization into this task for the first time, proposing a novel perturbation mechanism that explicitly protects speaker identity while retaining prosodic information. Furthermore, it designs a strictly causal, non-lookahead generative network to enable truly zero-latency streaming conversion. Without relying on future-frame buffering, the method effectively balances speaker privacy and speech utility, significantly enhancing both the naturalness and real-time performance of converted speech.

0 citationsRead paper

An Effective Solution for the CVPR 2026 8th UG2+ Challenge Track 3: Dynamic Object Segmentation in Turbulence

May 30, 2026

This work addresses the challenge of dynamic object segmentation under severe geometric distortions and noise induced by atmospheric turbulence. Building upon the Segment Any Motion framework, the authors propose a domain adaptation strategy that leverages simulated turbulence perturbations during training, coupled with a spatiotemporal post-processing module designed to preserve small targets while enforcing label consistency across frames. The proposed approach effectively suppresses boundary artifacts and transient noise, achieving second place in Track 3 of the CVPR 2026 UG2+ Challenge. This demonstrates a significant improvement in both accuracy and robustness for segmenting moving objects in turbulent imaging conditions.

0 citationsRead paper