Institution profile

Shanghai Collaborative Innovation Center of Intelligent Visual Computing

Academic institutionasia · cn
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

OmniGen-AR: AutoRegressive Any-to-Image Generation

Jun 08, 2026

Existing autoregressive vision generation models typically support only a single modality of conditioning, limiting their ability to meet real-world demands for image synthesis driven by multiple types of control signals. This work proposes OmniGen-AR, a unified autoregressive framework that discretizes diverse multimodal conditions—including text, spatial layouts, and visual context—into a shared token space via a common visual-text tokenizer. To prevent information leakage during training, the model introduces Decoupled Causal Attention (DCA), which separates causal dependencies between conditioning and content tokens while preserving the standard autoregressive prediction pipeline during inference. Evaluated on benchmarks such as GenEval (0.63) and VBench (80.02), OmniGen-AR achieves state-of-the-art or competitive performance, demonstrating its effectiveness in flexible, high-fidelity multimodal image generation.

0 citationsRead paper

TDR: Task-Decoupled Retrieval with Fine-Grained LLM Feedback for In-Context Learning

Jul 24, 2025

In-context learning (ICL), cross-task example retrieval faces two key challenges: (1) inter-task data distribution entanglement, and (2) misalignment between retrieved examples and large language model (LLM) feedback at a fine-grained level. To address these, we propose Task-Decoupled Retrieval (TDR), the first framework to explicitly decouple cross-task data distributions and establish an LLM-based fine-grained feedback mechanism for retriever training. TDR supports plug-and-play deployment and multi-model adaptation. It integrates task-aware retrieval strategies with fine-grained feedback modeling, achieving significant ICL performance gains across 30 diverse NLP tasks. Empirical results show consistent improvements in average accuracy over state-of-the-art methods, with strong generalization across unseen tasks and models—establishing new SOTA performance.

0 citationsRead paper

Making Large Language Models Better Reasoners with Orchestrated Streaming Experiences

Apr 01, 2025

Zero-shot chain-of-thought (CoT) reasoning suffers from poor performance, while few-shot approaches rely on manually crafted examples. To address this, we propose RoSE—a retrieval-augmented self-enhancement framework that improves zero-shot CoT reasoning without human annotation or model fine-tuning. RoSE dynamically constructs and maintains a streaming experience pool of historical question-answer pairs (including CoT traces), automatically retrieving and orchestrating relevant instances to augment inference. Its core innovation is a problem-aware bucketed diversity sampling mechanism that jointly optimizes for semantic similarity, prediction uncertainty, and reasoning complexity to generate high-quality prompts. RoSE enables end-to-end self-improvement via LLM-driven experience storage, multi-dimensional similarity computation, uniform bucketed sampling, and explicit complexity modeling. Extensive experiments demonstrate consistent and significant gains in zero-shot reasoning accuracy across diverse benchmark tasks, multiple large language models, and CoT variants—achieving strong generalization with minimal deployment overhead.

0 citationsRead paper
Recent publications

Latest Papers

OmniGen-AR: AutoRegressive Any-to-Image Generation

Jun 08, 2026

Existing autoregressive vision generation models typically support only a single modality of conditioning, limiting their ability to meet real-world demands for image synthesis driven by multiple types of control signals. This work proposes OmniGen-AR, a unified autoregressive framework that discretizes diverse multimodal conditions—including text, spatial layouts, and visual context—into a shared token space via a common visual-text tokenizer. To prevent information leakage during training, the model introduces Decoupled Causal Attention (DCA), which separates causal dependencies between conditioning and content tokens while preserving the standard autoregressive prediction pipeline during inference. Evaluated on benchmarks such as GenEval (0.63) and VBench (80.02), OmniGen-AR achieves state-of-the-art or competitive performance, demonstrating its effectiveness in flexible, high-fidelity multimodal image generation.

0 citationsRead paper

TDR: Task-Decoupled Retrieval with Fine-Grained LLM Feedback for In-Context Learning

Jul 24, 2025

In-context learning (ICL), cross-task example retrieval faces two key challenges: (1) inter-task data distribution entanglement, and (2) misalignment between retrieved examples and large language model (LLM) feedback at a fine-grained level. To address these, we propose Task-Decoupled Retrieval (TDR), the first framework to explicitly decouple cross-task data distributions and establish an LLM-based fine-grained feedback mechanism for retriever training. TDR supports plug-and-play deployment and multi-model adaptation. It integrates task-aware retrieval strategies with fine-grained feedback modeling, achieving significant ICL performance gains across 30 diverse NLP tasks. Empirical results show consistent improvements in average accuracy over state-of-the-art methods, with strong generalization across unseen tasks and models—establishing new SOTA performance.

0 citationsRead paper

Making Large Language Models Better Reasoners with Orchestrated Streaming Experiences

Apr 01, 2025

Zero-shot chain-of-thought (CoT) reasoning suffers from poor performance, while few-shot approaches rely on manually crafted examples. To address this, we propose RoSE—a retrieval-augmented self-enhancement framework that improves zero-shot CoT reasoning without human annotation or model fine-tuning. RoSE dynamically constructs and maintains a streaming experience pool of historical question-answer pairs (including CoT traces), automatically retrieving and orchestrating relevant instances to augment inference. Its core innovation is a problem-aware bucketed diversity sampling mechanism that jointly optimizes for semantic similarity, prediction uncertainty, and reasoning complexity to generate high-quality prompts. RoSE enables end-to-end self-improvement via LLM-driven experience storage, multi-dimensional similarity computation, uniform bucketed sampling, and explicit complexity modeling. Extensive experiments demonstrate consistent and significant gains in zero-shot reasoning accuracy across diverse benchmark tasks, multiple large language models, and CoT variants—achieving strong generalization with minimal deployment overhead.

0 citationsRead paper