Institution profile

Stony Brook University

Academic institutionnorthamerica · us
Official website
Research library721linked papers
Opportunities0open roles
Selected work

Representative Papers

From Uncertainty to Trust: Enhancing Reliability in Vision-Language Models with Uncertainty-Guided Dropout Decoding

Dec 09, 2024arXiv.org

LVLMs frequently suffer from hallucinations and unreliable outputs due to misinterpretation of visual inputs. To address this, we propose an uncertainty-guided inference-time visual token dropout method: (1) the first adaptation of dropout to the visual token level during inference; (2) decoupled modeling of epistemic and aleatoric uncertainty, with explicit focus on quantifying perceptual errors; (3) uncertainty estimation via projection of visual tokens into the text embedding space, followed by weighted masking; and (4) robust, training-free correction via multi-context masked decoding and ensemble prediction. Evaluated on CHAIR, THRONE, and MMBench, our method significantly reduces object hallucination (OH) while substantially improving output reliability and cross-scenario generation quality.

13 citationsRead paper

Too Many Frames, not all Useful: Efficient Strategies for Long-Form Video QA

Jun 13, 2024arXiv.org

Long-form video question answering (LVQA) suffers from high visual redundancy and sparse salient information, while existing methods inefficiently process uniformly sampled frames via independent vision-language model (VLM) descriptions, leading to poor semantic utilization. To address this, we propose the Hierarchical Keyframe Selector (HKFS), the first framework to jointly perform question-guided dynamic temporal segment localization and semantic keyframe selection. HKFS integrates multi-granularity temporal modeling, question-driven visual attention, and a lightweight VLM adaptation architecture—LVNet—to substantially reduce visual-language modeling overhead. Our approach achieves state-of-the-art performance on three major LVQA benchmarks—EgoSchema, NExT-QA, and IntentQA—and demonstrates strong generalization on VideoMME. Notably, it supports LVQA over videos up to one hour in length, enabling scalable, efficient, and semantically grounded long-video understanding.

12 citations2 influentialRead paper

AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science

May 25, 2025arXiv.org

Large language models (LLMs) struggle to critically leverage external domain knowledge in automated data science. Method: We introduce AssistedDS—the first benchmark for domain-knowledge-assisted evaluation—comprising synthetic and real Kaggle datasets paired with beneficial or adversarial domain documents. Our “interpretable synthesis + real-world scenarios” dual-track framework employs multi-stage prompting to assess end-to-end capabilities: document retrieval, knowledge filtering, code generation, and execution validation. Contribution/Results: We uncover a critical “blind adoption” flaw in LLMs: they fail significantly in time-series modeling, cross-fold consistency, and categorical variable handling. Experiments show state-of-the-art models suffer sharp performance degradation under adversarial documents; beneficial knowledge fails to mitigate harmful information, exposing severe deficiencies in domain knowledge discrimination and robust application.

3 citations1 influentialRead paper

A Multimodal Approach Combining Structural and Cross-domain Textual Guidance for Weakly Supervised OCT Segmentation

Nov 19, 2024IEEE journal of biomedical and health informatics

To address the high cost of pixel-level annotations and low-quality pseudo-labels in weakly supervised OCT image segmentation, this paper proposes a dual-guided (structural and textual) pseudo-label generation framework. Methodologically: (1) a structure-aware layer enhancement module is designed to improve anatomical layer segmentation robustness; (2) a dual-path text-guided mechanism integrates image-level label-derived textual descriptions with synthetically generated descriptive texts to achieve vision–semantics cross-modal alignment; (3) the framework incorporates CLIP-driven cross-domain text embeddings, a dual-branch visual encoder, and an iterative pseudo-label refinement strategy. Evaluated on three public OCT datasets, the method achieves significant mIoU improvements over existing weakly supervised approaches, establishing new state-of-the-art performance. The source code and pretrained models are publicly released.

3 citationsRead paper
Recent publications

Latest Papers