Estimating Hierarchically Rank Structured Covariance Matrices
本文针对高维协方差矩阵估计问题,提出了一种基于分层秩结构的正则化方法,以少量样本高效估计协方差矩阵,减少采样误差。
本文针对高维协方差矩阵估计问题,提出了一种基于分层秩结构的正则化方法,以少量样本高效估计协方差矩阵,减少采样误差。
This study investigates the performance limitations of local search for the vertex coloring problem on bipartite graphs, revealing that it can become trapped in arbitrarily poor local optima on certain graph structures. To address this issue, the authors propose a gray-box mutation operator based on color frequency that leverages structural information of the problem to guide the search process. Theoretical analysis demonstrates that on complete bipartite graphs, this operator reduces the expected runtime from exponential to Θ(n log n), enabling efficient discovery of globally optimal colorings. Beyond characterizing the fitness landscape of local search in this context, the work introduces an effective gray-box strategy for combinatorial optimization that exploits problem-specific knowledge to enhance search efficiency.
This study addresses a long-standing open problem concerning language class inclusions in cooperative distributed grammar systems (CDGS) by systematically integrating two regulatory mechanisms—forbidding random context and ordered rules—into the CDGS framework for the first time. The authors construct four novel model variants that operate along two dimensions: rule ordering within components and execution ordering across components. Employing formal language-theoretic methods, they precisely characterize the generative power of these models, establishing exact equivalences with five well-known classes of controlled rewriting languages. This work not only resolves several previously unsettled inclusion relationships reported in the literature but also clarifies the formerly ambiguous hierarchy of language classes and substantially streamlines the theoretical foundation of CDGS.
This study addresses the challenge of identifying market-predictive statements in cryptocurrency-related tweets and analyzing their implicit sentiment. To this end, the authors propose a two-stage classification framework: the first stage distinguishes predictive from non-predictive tweets, while the second categorizes predictive tweets into bullish, bearish, or neutral classes. A balanced dataset is constructed by combining human-annotated data with synthetic examples generated by GPT, and sentiment features are extracted using SenticNet. Experimental results show that Transformer-based models achieve the highest F1 score in the first stage, whereas traditional machine learning methods perform best in the second stage. Data augmentation with GPT effectively mitigates class imbalance and enhances overall performance, further revealing significant associations between prediction categories and sentiment features.
This study investigates how dimensionality influences satisfiability, solver hardness, and unsatisfiability proof size in random geometric SAT instances, aiming to bridge the gap between theoretical complexity and the empirical tractability of industrial SAT problems. By generating SAT instances from random geometric graphs and conducting large-scale experiments with modern solvers and proof complexity tools, the work reveals—for the first time—that low-dimensional geometric instances exhibit no peak in solver difficulty at the satisfiability threshold, and that solving time is uncorrelated with proof size. Moreover, low-dimensional instances are easier to solve and have a lower threshold clause density, gradually converging toward uniformly hard random instances as dimensionality increases. These findings demonstrate that the model continuously generates a spectrum of instances ranging from easy to hard, effectively capturing key characteristics of real-world industrial SAT benchmarks.
本文针对高维协方差矩阵估计问题,提出了一种基于分层秩结构的正则化方法,以少量样本高效估计协方差矩阵,减少采样误差。
This study investigates the performance limitations of local search for the vertex coloring problem on bipartite graphs, revealing that it can become trapped in arbitrarily poor local optima on certain graph structures. To address this issue, the authors propose a gray-box mutation operator based on color frequency that leverages structural information of the problem to guide the search process. Theoretical analysis demonstrates that on complete bipartite graphs, this operator reduces the expected runtime from exponential to Θ(n log n), enabling efficient discovery of globally optimal colorings. Beyond characterizing the fitness landscape of local search in this context, the work introduces an effective gray-box strategy for combinatorial optimization that exploits problem-specific knowledge to enhance search efficiency.
This study addresses a long-standing open problem concerning language class inclusions in cooperative distributed grammar systems (CDGS) by systematically integrating two regulatory mechanisms—forbidding random context and ordered rules—into the CDGS framework for the first time. The authors construct four novel model variants that operate along two dimensions: rule ordering within components and execution ordering across components. Employing formal language-theoretic methods, they precisely characterize the generative power of these models, establishing exact equivalences with five well-known classes of controlled rewriting languages. This work not only resolves several previously unsettled inclusion relationships reported in the literature but also clarifies the formerly ambiguous hierarchy of language classes and substantially streamlines the theoretical foundation of CDGS.
This study addresses the challenge of identifying market-predictive statements in cryptocurrency-related tweets and analyzing their implicit sentiment. To this end, the authors propose a two-stage classification framework: the first stage distinguishes predictive from non-predictive tweets, while the second categorizes predictive tweets into bullish, bearish, or neutral classes. A balanced dataset is constructed by combining human-annotated data with synthetic examples generated by GPT, and sentiment features are extracted using SenticNet. Experimental results show that Transformer-based models achieve the highest F1 score in the first stage, whereas traditional machine learning methods perform best in the second stage. Data augmentation with GPT effectively mitigates class imbalance and enhances overall performance, further revealing significant associations between prediction categories and sentiment features.
This study investigates how dimensionality influences satisfiability, solver hardness, and unsatisfiability proof size in random geometric SAT instances, aiming to bridge the gap between theoretical complexity and the empirical tractability of industrial SAT problems. By generating SAT instances from random geometric graphs and conducting large-scale experiments with modern solvers and proof complexity tools, the work reveals—for the first time—that low-dimensional geometric instances exhibit no peak in solver difficulty at the satisfiability threshold, and that solving time is uncorrelated with proof size. Moreover, low-dimensional instances are easier to solve and have a lower threshold clause density, gradually converging toward uniformly hard random instances as dimensionality increases. These findings demonstrate that the model continuously generates a spectrum of instances ranging from easy to hard, effectively capturing key characteristics of real-world industrial SAT benchmarks.