VeriDx: Earning the Right to Diagnose with Disease-Centric Verification
本文提出VeriDx框架,通过疾病中心验证解决医学LLM诊断中因忽略假设引起的义务而导致的错误,如遗漏关键测试等。
本文提出VeriDx框架,通过疾病中心验证解决医学LLM诊断中因忽略假设引起的义务而导致的错误,如遗漏关键测试等。
本文提出了一种无需深度图的RGB-D显著物体检测方法,通过从冻结的Depth Anything V2模型中蒸馏几何信息,并结合像素级可靠性估计器来提升RGB-only模型性能。
研究了在证据选择性披露情况下,如何分配曝光度以平衡偏好不同决策的双方。通过调整曝光度,影响证据选择和权重,从而优化决策过程。
This work addresses the challenge of visual query localization in first-person videos, where ambiguous object boundaries and insufficient global contextual guidance hinder performance. Inspired by the hierarchical perceptual mechanisms of the human cortex, the authors propose a unified 2D/3D visual query localization framework. The method leverages segmentation priors to extract foreground-aware query representations and employs deformable correlation filters for robust localization, further refining boundaries through multi-scale region-adaptive contextual feedback. In 3D scenes, it innovatively introduces a geometry–semantics joint confidence measure to evaluate and fuse multi-view information. To the best of our knowledge, this is the first study to integrate hierarchical perception with geometry–semantics confidence modeling for this task, achieving state-of-the-art performance on both VQL-2D and VQL-3D benchmarks.
This work addresses the suboptimal decision-making of existing Transformer-based agents in non-stationary, partially observable environments, where reliance solely on observation similarity for attention retrieval fails to distinguish between distinct action–reward histories under identical observations. To overcome this limitation, the paper introduces the Utility-Augmented Transformer (UAT), which formally characterizes the “feedback-blind retrieval” problem and incorporates a utility-state-modulated attention mechanism. This mechanism explicitly integrates action–reward history into query, key, and value projections to guide contextual retrieval. The proposed architecture features a zero-gate degeneracy property, strictly generalizing the representational capacity of observation-only Transformers, and offers theoretical guarantees under Lipschitz continuity and finite-horizon assumptions. Empirical results demonstrate that UAT significantly outperforms current baselines across four non-stationary benchmark tasks, particularly excelling in high-noise regimes and rapid adaptation scenarios.
本文提出VeriDx框架,通过疾病中心验证解决医学LLM诊断中因忽略假设引起的义务而导致的错误,如遗漏关键测试等。
本文提出了一种无需深度图的RGB-D显著物体检测方法,通过从冻结的Depth Anything V2模型中蒸馏几何信息,并结合像素级可靠性估计器来提升RGB-only模型性能。
研究了在证据选择性披露情况下,如何分配曝光度以平衡偏好不同决策的双方。通过调整曝光度,影响证据选择和权重,从而优化决策过程。
This work addresses the challenge of visual query localization in first-person videos, where ambiguous object boundaries and insufficient global contextual guidance hinder performance. Inspired by the hierarchical perceptual mechanisms of the human cortex, the authors propose a unified 2D/3D visual query localization framework. The method leverages segmentation priors to extract foreground-aware query representations and employs deformable correlation filters for robust localization, further refining boundaries through multi-scale region-adaptive contextual feedback. In 3D scenes, it innovatively introduces a geometry–semantics joint confidence measure to evaluate and fuse multi-view information. To the best of our knowledge, this is the first study to integrate hierarchical perception with geometry–semantics confidence modeling for this task, achieving state-of-the-art performance on both VQL-2D and VQL-3D benchmarks.
This work addresses the suboptimal decision-making of existing Transformer-based agents in non-stationary, partially observable environments, where reliance solely on observation similarity for attention retrieval fails to distinguish between distinct action–reward histories under identical observations. To overcome this limitation, the paper introduces the Utility-Augmented Transformer (UAT), which formally characterizes the “feedback-blind retrieval” problem and incorporates a utility-state-modulated attention mechanism. This mechanism explicitly integrates action–reward history into query, key, and value projections to guide contextual retrieval. The proposed architecture features a zero-gate degeneracy property, strictly generalizing the representational capacity of observation-only Transformers, and offers theoretical guarantees under Lipschitz continuity and finite-horizon assumptions. Empirical results demonstrate that UAT significantly outperforms current baselines across four non-stationary benchmark tasks, particularly excelling in high-noise regimes and rapid adaptation scenarios.