Robust Nonparametric Testing for Structural Changes in Multivariate Volatility via Multiple Quantiles
本文提出一种非参数检验方法,通过聚合多个分位数的广义分位数得分来检测多变量波动率矩阵中的结构变化,适用于重尾分布且无需指定备择假设下的波动动态。
本文提出一种非参数检验方法,通过聚合多个分位数的广义分位数得分来检测多变量波动率矩阵中的结构变化,适用于重尾分布且无需指定备择假设下的波动动态。
本文提出了一种基于经验似然和协变量平衡约束的方法,在随机实验中调整协变量的同时保持估计量的单调性,提高了估计效率。
This work addresses the tension in multivariate time series forecasting between modeling inter-channel dependencies and preserving model flexibility: channel-dependent approaches are prone to overfitting due to sensitivity to channel ordering, while channel-independent models neglect inter-channel relationships. To resolve this, we propose CPiRi, a novel framework that introduces permutation invariance over channels into multivariate time series modeling for the first time. CPiRi employs a spatiotemporal decoupling architecture, a frozen pre-trained temporal encoder, a lightweight spatial relation module, and a channel-shuffling training strategy to adaptively infer channel relationships from data. Grounded in permutation equivariance theory, our approach ensures strong inductive generalization to unseen channel configurations. Experiments demonstrate that CPiRi achieves state-of-the-art performance across multiple benchmarks, exhibits robustness to channel order perturbations, generalizes to full-channel settings using only half the channels during training, and maintains efficiency at scale.
Scopus contains millions of “homeless” publications—authored by researchers with complete institutional affiliations yet erroneously labeled “country-undefined”—undermining database reliability and compromising the accuracy and fairness of research evaluation. This study systematically identifies four primary root causes: incomplete address information, failure to recognize national name variants, typographical errors, and deficiencies in address parsing algorithms. Leveraging 124 years of Scopus metadata, we integrate bibliometric analysis, multilingual standardization of country names, and fine-grained data cleaning to classify and quantitatively trace these causes. Our findings yield a reproducible methodological framework for metadata quality enhancement, enabling institutional affiliation calibration, cross-national research performance assessment, and optimization of scholarly infrastructure. The approach advances best practices in bibliographic data curation and supports equitable, evidence-based science policy.
Large language models (LLMs) frequently exhibit hallucination and produce unreliable outputs in multiple-choice question answering (MCQA). Method: This paper proposes the first trustworthy reasoning framework for MCQA that jointly integrates statistical significance testing and calibration-preserving prediction. It constructs a response distribution via self-consistent sampling, uses response frequency as a test statistic for p-value computation, and builds a minimum prediction set with theoretically guaranteed miscoverage rate ≤ α via empirical risk control. Contribution/Results: It is the first work to introduce hypothesis testing into LLM uncertainty quantification, ensuring strict calibration of prediction sets; it further proves that prediction set size serves as a valid uncertainty measure. Experiments on MMLU and MMLU-Pro demonstrate precise α-level miscoverage control and monotonic reduction of prediction set size with decreasing α, significantly enhancing reliability and interpretability—especially in high-risk scenarios.
本文提出一种非参数检验方法,通过聚合多个分位数的广义分位数得分来检测多变量波动率矩阵中的结构变化,适用于重尾分布且无需指定备择假设下的波动动态。
本文提出了一种基于经验似然和协变量平衡约束的方法,在随机实验中调整协变量的同时保持估计量的单调性,提高了估计效率。
This work addresses the tension in multivariate time series forecasting between modeling inter-channel dependencies and preserving model flexibility: channel-dependent approaches are prone to overfitting due to sensitivity to channel ordering, while channel-independent models neglect inter-channel relationships. To resolve this, we propose CPiRi, a novel framework that introduces permutation invariance over channels into multivariate time series modeling for the first time. CPiRi employs a spatiotemporal decoupling architecture, a frozen pre-trained temporal encoder, a lightweight spatial relation module, and a channel-shuffling training strategy to adaptively infer channel relationships from data. Grounded in permutation equivariance theory, our approach ensures strong inductive generalization to unseen channel configurations. Experiments demonstrate that CPiRi achieves state-of-the-art performance across multiple benchmarks, exhibits robustness to channel order perturbations, generalizes to full-channel settings using only half the channels during training, and maintains efficiency at scale.
Scopus contains millions of “homeless” publications—authored by researchers with complete institutional affiliations yet erroneously labeled “country-undefined”—undermining database reliability and compromising the accuracy and fairness of research evaluation. This study systematically identifies four primary root causes: incomplete address information, failure to recognize national name variants, typographical errors, and deficiencies in address parsing algorithms. Leveraging 124 years of Scopus metadata, we integrate bibliometric analysis, multilingual standardization of country names, and fine-grained data cleaning to classify and quantitatively trace these causes. Our findings yield a reproducible methodological framework for metadata quality enhancement, enabling institutional affiliation calibration, cross-national research performance assessment, and optimization of scholarly infrastructure. The approach advances best practices in bibliographic data curation and supports equitable, evidence-based science policy.
Large language models (LLMs) frequently exhibit hallucination and produce unreliable outputs in multiple-choice question answering (MCQA). Method: This paper proposes the first trustworthy reasoning framework for MCQA that jointly integrates statistical significance testing and calibration-preserving prediction. It constructs a response distribution via self-consistent sampling, uses response frequency as a test statistic for p-value computation, and builds a minimum prediction set with theoretically guaranteed miscoverage rate ≤ α via empirical risk control. Contribution/Results: It is the first work to introduce hypothesis testing into LLM uncertainty quantification, ensuring strict calibration of prediction sets; it further proves that prediction set size serves as a valid uncertainty measure. Experiments on MMLU and MMLU-Pro demonstrate precise α-level miscoverage control and monotonic reduction of prediction set size with decreasing α, significantly enhancing reliability and interpretability—especially in high-risk scenarios.