Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study
本文通过定义和分解多元同意指数Gamma,量化分析了多数投票在提高LLM答案准确性时可能失效的问题,并以GPT-4.1为例进行了研究。
本文通过定义和分解多元同意指数Gamma,量化分析了多数投票在提高LLM答案准确性时可能失效的问题,并以GPT-4.1为例进行了研究。
In semi-supervised medical image segmentation under low-labeling ratios, excessive perturbations degrade consistency learning and blur decision boundaries. To address this, we propose a confidence-aware adaptive displacement mechanism: local patches are dynamically selected for replacement based on confidence maps, integrated with learnable thresholding, consistency regularization, and uncertainty-aware training to enable progressive pseudo-label refinement. Our method is the first to jointly couple confidence-guided geometric perturbation with adaptive thresholding, effectively mitigating label noise propagation in low-confidence regions. Evaluated on multiple public medical benchmarks, it achieves state-of-the-art performance—improving average Dice score by 2.1% and reducing Hausdorff distance (HD95) by 18.7%, with particularly notable gains in robustness for ambiguous boundary segmentation.
本文通过定义和分解多元同意指数Gamma,量化分析了多数投票在提高LLM答案准确性时可能失效的问题,并以GPT-4.1为例进行了研究。
In semi-supervised medical image segmentation under low-labeling ratios, excessive perturbations degrade consistency learning and blur decision boundaries. To address this, we propose a confidence-aware adaptive displacement mechanism: local patches are dynamically selected for replacement based on confidence maps, integrated with learnable thresholding, consistency regularization, and uncertainty-aware training to enable progressive pseudo-label refinement. Our method is the first to jointly couple confidence-guided geometric perturbation with adaptive thresholding, effectively mitigating label noise propagation in low-confidence regions. Evaluated on multiple public medical benchmarks, it achieves state-of-the-art performance—improving average Dice score by 2.1% and reducing Hausdorff distance (HD95) by 18.7%, with particularly notable gains in robustness for ambiguous boundary segmentation.