Adaptive Incentive Design in Dynamic Principal-Agent Problem via Kernelized Bandits
本文通过引入随机性改进了动态委托-代理问题中的激励设计,使用Heteroscedastic GP-UCB算法解决了因不确定性导致的计算难题。
本文通过引入随机性改进了动态委托-代理问题中的激励设计,使用Heteroscedastic GP-UCB算法解决了因不确定性导致的计算难题。
This study systematically evaluates the mathematical reasoning capabilities of large language models (LLMs) on both solved and open problems in graph theory, as well as their applicability boundaries in computational education. Employing an eight-stage evaluation protocol that simulates authentic mathematical inquiry—integrating interactive prompt engineering, a mathematical reasoning assessment framework, and expert validation—the work presents the first comparative analysis of LLM performance across these two problem categories. Results demonstrate that LLMs can generate expert-verified correct proofs for solved problems. When confronted with open problems, LLMs propose plausible exploration strategies without exhibiting hallucinations, reflecting appropriate handling of uncertainty, yet they fail to achieve substantive breakthroughs. This work thus delineates both the capabilities and limitations of LLMs in rigorous mathematical reasoning.
This work identifies a systemic amplification of intersectional social biases in vision-language models (VLMs), stemming from their reliance on spurious statistical correlations rather than socially grounded contextual reasoning—particularly undermining fairness in occupation prediction. Using the FairFace dataset, we conduct experiments on five open-source VLMs, sampling reasoning trajectories via three distinct prompt styles. Integrating quantitative predictive evaluation with qualitative attribution analysis, we establish—for the first time—the causal linkage between model reasoning pathways and intersectional bias generation. Results reveal consistent and statistically significant bias patterns across 32 occupational categories, demonstrating that the inference process itself constitutes a critical stage for bias propagation and amplification. We thus argue for “reasoning alignment”: prior to deployment, VLMs must be calibrated so their internal reasoning logic aligns with human values and sociocultural knowledge. This reframes fairness governance for VLMs, proposing a novel paradigm centered on interpretability-aware bias mitigation.
本文通过引入随机性改进了动态委托-代理问题中的激励设计,使用Heteroscedastic GP-UCB算法解决了因不确定性导致的计算难题。
This study systematically evaluates the mathematical reasoning capabilities of large language models (LLMs) on both solved and open problems in graph theory, as well as their applicability boundaries in computational education. Employing an eight-stage evaluation protocol that simulates authentic mathematical inquiry—integrating interactive prompt engineering, a mathematical reasoning assessment framework, and expert validation—the work presents the first comparative analysis of LLM performance across these two problem categories. Results demonstrate that LLMs can generate expert-verified correct proofs for solved problems. When confronted with open problems, LLMs propose plausible exploration strategies without exhibiting hallucinations, reflecting appropriate handling of uncertainty, yet they fail to achieve substantive breakthroughs. This work thus delineates both the capabilities and limitations of LLMs in rigorous mathematical reasoning.
This work identifies a systemic amplification of intersectional social biases in vision-language models (VLMs), stemming from their reliance on spurious statistical correlations rather than socially grounded contextual reasoning—particularly undermining fairness in occupation prediction. Using the FairFace dataset, we conduct experiments on five open-source VLMs, sampling reasoning trajectories via three distinct prompt styles. Integrating quantitative predictive evaluation with qualitative attribution analysis, we establish—for the first time—the causal linkage between model reasoning pathways and intersectional bias generation. Results reveal consistent and statistically significant bias patterns across 32 occupational categories, demonstrating that the inference process itself constitutes a critical stage for bias propagation and amplification. We thus argue for “reasoning alignment”: prior to deployment, VLMs must be calibrated so their internal reasoning logic aligns with human values and sociocultural knowledge. This reframes fairness governance for VLMs, proposing a novel paradigm centered on interpretability-aware bias mitigation.