SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning
为解决复杂推理中示例选择问题,提出SALA框架,通过学习任务特定推理操作并用DTW对齐序列,实现灵活匹配。
为解决复杂推理中示例选择问题,提出SALA框架,通过学习任务特定推理操作并用DTW对齐序列,实现灵活匹配。
Tongue image diagnosis faces challenges including fine-grained vision–semantics modeling, scarce expert annotations, severe class imbalance, and insufficient clinical interpretability. Method: We propose MIRNet—a multimodal interpretable framework featuring (i) mask autoencoder (MAE)-based self-supervised pretraining for robust feature learning; (ii) a clinician-constructed constraint graph integrated with graph attention networks (GAT) and KL-divergence regularization to model label correlations and embed clinical priors; and (iii) an asymmetric loss (ASL) with dedicated regularization to mitigate class imbalance. We further introduce TongueAtlas-4K, a large-scale tongue diagnosis dataset comprising over 4,000 high-quality expert-annotated images. Results: MIRNet achieves state-of-the-art performance on multi-label tongue diagnosis, significantly improving model interpretability, cross-scenario generalization, and clinical plausibility. The framework demonstrates strong transferability to other medical imaging diagnostic tasks.
Video large language models (VideoLLMs) suffer from temporal hallucination—generating factually inconsistent descriptions misaligned with video content. To address this, we propose the first activation engineering framework explicitly designed for video temporal dynamics, requiring no model fine-tuning. Our method identifies time-sensitive neural modules via data-driven neuron activation analysis, then applies module-level dynamic activation scaling or masking guided by temporal variability quantification. Crucially, we empirically establish that temporal hallucination stems primarily from insufficient sensitivity to temporal dynamics—not task-specific factors—enabling targeted intervention. Evaluated across diverse VideoLLM architectures and standard benchmarks, our approach significantly reduces hallucination rates, improves factual consistency and temporal reasoning reliability, and preserves original model performance on non-hallucination metrics. This work introduces a paradigm shift in mitigating temporal hallucinations through interpretable, architecture-agnostic activation modulation.
为解决复杂推理中示例选择问题,提出SALA框架,通过学习任务特定推理操作并用DTW对齐序列,实现灵活匹配。
Tongue image diagnosis faces challenges including fine-grained vision–semantics modeling, scarce expert annotations, severe class imbalance, and insufficient clinical interpretability. Method: We propose MIRNet—a multimodal interpretable framework featuring (i) mask autoencoder (MAE)-based self-supervised pretraining for robust feature learning; (ii) a clinician-constructed constraint graph integrated with graph attention networks (GAT) and KL-divergence regularization to model label correlations and embed clinical priors; and (iii) an asymmetric loss (ASL) with dedicated regularization to mitigate class imbalance. We further introduce TongueAtlas-4K, a large-scale tongue diagnosis dataset comprising over 4,000 high-quality expert-annotated images. Results: MIRNet achieves state-of-the-art performance on multi-label tongue diagnosis, significantly improving model interpretability, cross-scenario generalization, and clinical plausibility. The framework demonstrates strong transferability to other medical imaging diagnostic tasks.
Video large language models (VideoLLMs) suffer from temporal hallucination—generating factually inconsistent descriptions misaligned with video content. To address this, we propose the first activation engineering framework explicitly designed for video temporal dynamics, requiring no model fine-tuning. Our method identifies time-sensitive neural modules via data-driven neuron activation analysis, then applies module-level dynamic activation scaling or masking guided by temporal variability quantification. Crucially, we empirically establish that temporal hallucination stems primarily from insufficient sensitivity to temporal dynamics—not task-specific factors—enabling targeted intervention. Evaluated across diverse VideoLLM architectures and standard benchmarks, our approach significantly reduces hallucination rates, improves factual consistency and temporal reasoning reliability, and preserves original model performance on non-hallucination metrics. This work introduces a paradigm shift in mitigating temporal hallucinations through interpretable, architecture-agnostic activation modulation.