Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection
研究通过独立分析推理拓扑和代理间多样性来源,解决多语言低资源情感检测问题,使用不同配置的多代理LLM系统进行评估。
研究通过独立分析推理拓扑和代理间多样性来源,解决多语言低资源情感检测问题,使用不同配置的多代理LLM系统进行评估。
本文针对像素文本编码器存在的问题,通过设计四关键组件并训练Pixel Linguist II模型,提升了多语言视觉文本理解能力。
本文提出Traceable Trust框架,以评估和设计AI输出到生物科学研究行动的转化过程,确保其可信度。
This study addresses the prevalent issues of coarse, misaligned, or incomplete manual annotations in remote sensing semantic segmentation, which often distort model evaluation. To tackle this, the authors propose a training-free, reference-free mask fidelity assessment method that constructs counterfactual image pairs—preserving and erasing the region within the mask—and leverages a frozen vision-language model to evaluate whether class-specific evidence is concentrated inside the mask and absent outside it. This approach enables, for the first time, reference-free auditing of annotation quality in remote sensing segmentation, revealing systematic labeling biases across categories and facilitating automatic refinement of supervision signals. The proposed Contrastive Mask Fidelity (CMF) metric achieves 81% agreement with expert judgments across ten remote sensing datasets, substantially outperforming existing methods, and CMF-guided supervision significantly enhances cross-domain transfer performance.
This paper exposes a fundamental gap between statistical significance (e.g., one-sided *p* = 0.025) and actual replicability: under identical sample sizes, the probability of replicating an effect in the same direction is only ~0.975, while the probability of reproducing statistical significance is markedly lower (~0.283). Conventional power analysis overestimates replicability by ignoring sampling variance in the original effect estimate. To address this, we develop a replication probability model grounded in variance propagation—formally integrating the sampling variances of both original and replication effect estimates. Our framework unifies frequentist and Bayesian perspectives, discarding noninformative priors in favor of discretized probability mass analysis. The resulting theory yields novel, high-confidence replication sample-size criteria, providing both theoretical foundations and practical tools for robust biomedical validation studies.
研究通过独立分析推理拓扑和代理间多样性来源,解决多语言低资源情感检测问题,使用不同配置的多代理LLM系统进行评估。
本文针对像素文本编码器存在的问题,通过设计四关键组件并训练Pixel Linguist II模型,提升了多语言视觉文本理解能力。
本文提出Traceable Trust框架,以评估和设计AI输出到生物科学研究行动的转化过程,确保其可信度。
This study addresses the prevalent issues of coarse, misaligned, or incomplete manual annotations in remote sensing semantic segmentation, which often distort model evaluation. To tackle this, the authors propose a training-free, reference-free mask fidelity assessment method that constructs counterfactual image pairs—preserving and erasing the region within the mask—and leverages a frozen vision-language model to evaluate whether class-specific evidence is concentrated inside the mask and absent outside it. This approach enables, for the first time, reference-free auditing of annotation quality in remote sensing segmentation, revealing systematic labeling biases across categories and facilitating automatic refinement of supervision signals. The proposed Contrastive Mask Fidelity (CMF) metric achieves 81% agreement with expert judgments across ten remote sensing datasets, substantially outperforming existing methods, and CMF-guided supervision significantly enhances cross-domain transfer performance.
This paper exposes a fundamental gap between statistical significance (e.g., one-sided *p* = 0.025) and actual replicability: under identical sample sizes, the probability of replicating an effect in the same direction is only ~0.975, while the probability of reproducing statistical significance is markedly lower (~0.283). Conventional power analysis overestimates replicability by ignoring sampling variance in the original effect estimate. To address this, we develop a replication probability model grounded in variance propagation—formally integrating the sampling variances of both original and replication effect estimates. Our framework unifies frequentist and Bayesian perspectives, discarding noninformative priors in favor of discretized probability mass analysis. The resulting theory yields novel, high-confidence replication sample-size criteria, providing both theoretical foundations and practical tools for robust biomedical validation studies.