Institution profile

SIONIC AI

Industry research
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Is Position Bias in Dense Retrievers Built In-or Learned from Data?

May 26, 2026

This work investigates position bias in dense retrievers, which tend to favor documents where relevant evidence appears near the beginning, often overlooking information located later. To systematically examine how the distribution of evidence positions in training data influences this bias, the authors construct synthetic training sets with controlled evidence placement (at the beginning, middle, or end) and fine-tune eight distinct pre-trained architectures. They demonstrate for the first time that the positional distribution of evidence in training data is a key controllable factor driving position bias. By introducing a position-balanced data construction strategy, they effectively mitigate this bias without compromising average retrieval performance. Experimental results show that their approach reduces model sensitivity to evidence position by 57%–87% on position-aware evaluation benchmarks.

0 citationsRead paper

Stop learning it all to mitigate visual hallucination, Focus on the hallucination target

Jun 13, 2025

Multimodal large language models (MLLMs) frequently exhibit visual hallucinations—generating objects not present in the input image—thereby severely compromising factual consistency and reliability in vision-language tasks. To address this, we propose a hallucination-targeted fine-grained preference learning framework. Our method is the first to localize preference optimization to specific hallucinated response segments and their corresponding image regions, enabling pixel-level supervision. We construct a novel dataset containing paired hallucinated/correct responses with precise pixel-level grounding annotations. By integrating multimodal alignment modeling with customized response chunking and region-aware labeling, our approach achieves interpretable and spatially grounded hallucination suppression. Extensive experiments demonstrate that our method significantly reduces hallucination rates across multiple visual hallucination benchmarks, substantially improving model factuality and reliability without degrading overall task performance.

0 citationsRead paper
Recent publications

Latest Papers

Is Position Bias in Dense Retrievers Built In-or Learned from Data?

May 26, 2026

This work investigates position bias in dense retrievers, which tend to favor documents where relevant evidence appears near the beginning, often overlooking information located later. To systematically examine how the distribution of evidence positions in training data influences this bias, the authors construct synthetic training sets with controlled evidence placement (at the beginning, middle, or end) and fine-tune eight distinct pre-trained architectures. They demonstrate for the first time that the positional distribution of evidence in training data is a key controllable factor driving position bias. By introducing a position-balanced data construction strategy, they effectively mitigate this bias without compromising average retrieval performance. Experimental results show that their approach reduces model sensitivity to evidence position by 57%–87% on position-aware evaluation benchmarks.

0 citationsRead paper

Stop learning it all to mitigate visual hallucination, Focus on the hallucination target

Jun 13, 2025

Multimodal large language models (MLLMs) frequently exhibit visual hallucinations—generating objects not present in the input image—thereby severely compromising factual consistency and reliability in vision-language tasks. To address this, we propose a hallucination-targeted fine-grained preference learning framework. Our method is the first to localize preference optimization to specific hallucinated response segments and their corresponding image regions, enabling pixel-level supervision. We construct a novel dataset containing paired hallucinated/correct responses with precise pixel-level grounding annotations. By integrating multimodal alignment modeling with customized response chunking and region-aware labeling, our approach achieves interpretable and spatially grounded hallucination suppression. Extensive experiments demonstrate that our method significantly reduces hallucination rates across multiple visual hallucination benchmarks, substantially improving model factuality and reliability without degrading overall task performance.

0 citationsRead paper