Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs
研究解决了大型语言模型在多选题中偏好流行但错误选项的问题,通过引入PopMCQ基准测试和PopDebias方法有效缓解了这一倾向。
研究解决了大型语言模型在多选题中偏好流行但错误选项的问题,通过引入PopMCQ基准测试和PopDebias方法有效缓解了这一倾向。
针对文本属性图上的少样本节点分类问题,提出了一种基于置信度的双教师学习框架CoTeach,动态选择更可靠的教师以提高性能并降低成本。
This study addresses the absence of public benchmarks integrating language models with hypergraph learning by introducing TAHB, the first text-attributed hypergraph benchmark. Comprising ten real-world datasets, TAHB supports dual evaluation paradigms encompassing both LLM augmentation and prediction tasks. Systematic experiments validate the efficacy of text-aware hypergraph representation learning, demonstrating that LLM-enhanced semantics significantly improve model performance and that joint structural-textual modeling constitutes the optimal predictive strategy. By filling a critical gap in the literature, TAHB provides a standardized evaluation platform and essential empirical evidence for the deep integration of hypergraph learning and large language models, thereby facilitating future research at this emerging intersection.
This study investigates how individuals respond to AI-generated financial advice in real-world pension investment decisions and its causal impact on asset allocation. Through a 2×2 randomized controlled trial involving 400 Korean workplace pension participants, the research delivered either aggressive or conservative AI recommendations—with or without explanatory rationales—and combined behavioral economics tasks with econometric analysis to identify the causal effects of AI advice in an authentic pension context. The findings reveal that approximately 37% of the recommended portfolio shifts were transmitted to participants’ final allocations, significantly altering expected returns, volatility, and risk profiles. While 81% of participants adjusted their choices—95% of whom moved in the direction of the advice—they implemented only about half of the suggested change on average. Notably, providing explanatory rationales did not significantly enhance compliance. The results highlight selective adoption and partial adherence to AI advice, offering empirical insights for the design of robo-advisory systems.
This study addresses the inefficiency and poor generalization of traditional trial-and-error approaches in laser cutting of optical thin films. To overcome these limitations, the authors propose RL²C, a reinforcement learning framework that integrates Q-learning with an ε-greedy strategy and incorporates a dynamic state-space expansion mechanism to adaptively tune critical parameters such as focal length and laser power. This work represents the first application of reinforcement learning with dynamic environmental adaptability to industrial laser cutting, significantly enhancing optimization efficiency and cross-material generalization. Experimental results demonstrate that RL²C reduces the number of optimization steps by 12.5% and processing time by 81.8% compared to existing methods, while effectively minimizing cut taper and film loss.
研究解决了大型语言模型在多选题中偏好流行但错误选项的问题,通过引入PopMCQ基准测试和PopDebias方法有效缓解了这一倾向。
针对文本属性图上的少样本节点分类问题,提出了一种基于置信度的双教师学习框架CoTeach,动态选择更可靠的教师以提高性能并降低成本。
This study addresses the absence of public benchmarks integrating language models with hypergraph learning by introducing TAHB, the first text-attributed hypergraph benchmark. Comprising ten real-world datasets, TAHB supports dual evaluation paradigms encompassing both LLM augmentation and prediction tasks. Systematic experiments validate the efficacy of text-aware hypergraph representation learning, demonstrating that LLM-enhanced semantics significantly improve model performance and that joint structural-textual modeling constitutes the optimal predictive strategy. By filling a critical gap in the literature, TAHB provides a standardized evaluation platform and essential empirical evidence for the deep integration of hypergraph learning and large language models, thereby facilitating future research at this emerging intersection.
This study investigates how individuals respond to AI-generated financial advice in real-world pension investment decisions and its causal impact on asset allocation. Through a 2×2 randomized controlled trial involving 400 Korean workplace pension participants, the research delivered either aggressive or conservative AI recommendations—with or without explanatory rationales—and combined behavioral economics tasks with econometric analysis to identify the causal effects of AI advice in an authentic pension context. The findings reveal that approximately 37% of the recommended portfolio shifts were transmitted to participants’ final allocations, significantly altering expected returns, volatility, and risk profiles. While 81% of participants adjusted their choices—95% of whom moved in the direction of the advice—they implemented only about half of the suggested change on average. Notably, providing explanatory rationales did not significantly enhance compliance. The results highlight selective adoption and partial adherence to AI advice, offering empirical insights for the design of robo-advisory systems.
This study addresses the inefficiency and poor generalization of traditional trial-and-error approaches in laser cutting of optical thin films. To overcome these limitations, the authors propose RL²C, a reinforcement learning framework that integrates Q-learning with an ε-greedy strategy and incorporates a dynamic state-space expansion mechanism to adaptively tune critical parameters such as focal length and laser power. This work represents the first application of reinforcement learning with dynamic environmental adaptability to industrial laser cutting, significantly enhancing optimization efficiency and cross-material generalization. Experimental results demonstrate that RL²C reduces the number of optimization steps by 12.5% and processing time by 81.8% compared to existing methods, while effectively minimizing cut taper and film loss.