A Deep Latent Variable Framework for Jointly Modeling Missingness, Measurement Error, and Heterogeneity
本文提出一种深度潜在变量框架,综合处理缺失数据、测量误差和人群异质性问题,采用分层树路由变分自编码器和校准去噪方法。
本文提出一种深度潜在变量框架,综合处理缺失数据、测量误差和人群异质性问题,采用分层树路由变分自编码器和校准去噪方法。
This study investigates the finiteness of $k$-vertex-critical graphs within several subclasses of $(P_4 + \ell P_1)$-free graphs. By forbidding specific induced subgraphs such as the chair, bull, and cricket, and leveraging structural characterizations of the graph families $B_n(m)$ and their extensions $B_n(m)^+$, the authors establish for the first time that multiple such subclasses contain only finitely many $k$-vertex-critical graphs. A key contribution is the derivation of an upper bound on the chromatic number, $\chi(G) \leq \ell + 2$, along with an improved general bound of $O(\ell^{\omega-1})$. These results yield polynomial-time certifying algorithms for $k$-colorability in the corresponding graph classes.
This work addresses the lack of a solid statistical foundation in traditional TF-IDF and its inability to capture term burstiness. The authors propose a statistical framework based on a penalized likelihood ratio test, modeling term frequencies with a Beta-Binomial distribution and incorporating a Gamma prior as a regularizer. This formulation yields a term-weighting statistic that aligns with existing TF-IDF variants. Notably, the study provides the first theoretical interpretation of TF-IDF from a hypothesis testing perspective, uncovering its intrinsic connection to burstiness modeling. Experimental results on document classification tasks demonstrate that the proposed method achieves performance comparable to classical TF-IDF, thereby validating the effectiveness and soundness of the proposed statistical framework.
TF-IDF, a cornerstone heuristic in information retrieval, lacks a formal statistical foundation for assessing term significance. Method: Modeling term–document co-occurrence as binomial sampling, we show that TF-IDF asymptotically approximates the negative logarithm of the Fisher exact test p-value under large-sample conditions. We further introduce TF-ICF (inverse corpus frequency), rigorously proving its high correlation with the Fisher p-value under mild assumptions and its convergence to classical TF-IDF in the infinite-document limit. Contribution/Results: This work establishes, for the first time, a formal statistical link between TF-IDF—the most influential term-weighting heuristic—and classical hypothesis testing. It endows TF-IDF with an interpretable statistical semantics: quantifying the significance of deviation from random term distribution. Moreover, it provides a principled theoretical framework for designing novel term-weighting schemes grounded in statistical inference, thereby bridging empirical IR practice with foundational statistical theory.
This study addresses the challenge of identifying minor defects during early software development stages, where labeled data are scarce. We present the first systematic evaluation of Quantum Support Vector Classifiers (QSVC/PQSVC) for defect prediction on real-world open-source projects. Leveraging 30,924 commit records across 14 projects, we propose a subset-aggregation framework for QSVC-based defect prediction and an incremental testing mechanism: (i) subset partitioning with weighted voting enhances prediction consistency; (ii) an incremental feature-mapping strategy mitigates mapping failures induced by quantum hardware resource constraints. Experimental results demonstrate that QSVC/PQSVC achieve accuracy comparable to—and in some cases exceeding—that of classical SVC, validating the feasibility and practical potential of quantum SVMs in STAF (Software Testing at Early-stage and with Few labels) scenarios.
本文提出一种深度潜在变量框架,综合处理缺失数据、测量误差和人群异质性问题,采用分层树路由变分自编码器和校准去噪方法。
This study investigates the finiteness of $k$-vertex-critical graphs within several subclasses of $(P_4 + \ell P_1)$-free graphs. By forbidding specific induced subgraphs such as the chair, bull, and cricket, and leveraging structural characterizations of the graph families $B_n(m)$ and their extensions $B_n(m)^+$, the authors establish for the first time that multiple such subclasses contain only finitely many $k$-vertex-critical graphs. A key contribution is the derivation of an upper bound on the chromatic number, $\chi(G) \leq \ell + 2$, along with an improved general bound of $O(\ell^{\omega-1})$. These results yield polynomial-time certifying algorithms for $k$-colorability in the corresponding graph classes.
This work addresses the lack of a solid statistical foundation in traditional TF-IDF and its inability to capture term burstiness. The authors propose a statistical framework based on a penalized likelihood ratio test, modeling term frequencies with a Beta-Binomial distribution and incorporating a Gamma prior as a regularizer. This formulation yields a term-weighting statistic that aligns with existing TF-IDF variants. Notably, the study provides the first theoretical interpretation of TF-IDF from a hypothesis testing perspective, uncovering its intrinsic connection to burstiness modeling. Experimental results on document classification tasks demonstrate that the proposed method achieves performance comparable to classical TF-IDF, thereby validating the effectiveness and soundness of the proposed statistical framework.
TF-IDF, a cornerstone heuristic in information retrieval, lacks a formal statistical foundation for assessing term significance. Method: Modeling term–document co-occurrence as binomial sampling, we show that TF-IDF asymptotically approximates the negative logarithm of the Fisher exact test p-value under large-sample conditions. We further introduce TF-ICF (inverse corpus frequency), rigorously proving its high correlation with the Fisher p-value under mild assumptions and its convergence to classical TF-IDF in the infinite-document limit. Contribution/Results: This work establishes, for the first time, a formal statistical link between TF-IDF—the most influential term-weighting heuristic—and classical hypothesis testing. It endows TF-IDF with an interpretable statistical semantics: quantifying the significance of deviation from random term distribution. Moreover, it provides a principled theoretical framework for designing novel term-weighting schemes grounded in statistical inference, thereby bridging empirical IR practice with foundational statistical theory.
This study addresses the challenge of identifying minor defects during early software development stages, where labeled data are scarce. We present the first systematic evaluation of Quantum Support Vector Classifiers (QSVC/PQSVC) for defect prediction on real-world open-source projects. Leveraging 30,924 commit records across 14 projects, we propose a subset-aggregation framework for QSVC-based defect prediction and an incremental testing mechanism: (i) subset partitioning with weighted voting enhances prediction consistency; (ii) an incremental feature-mapping strategy mitigates mapping failures induced by quantum hardware resource constraints. Experimental results demonstrate that QSVC/PQSVC achieve accuracy comparable to—and in some cases exceeding—that of classical SVC, validating the feasibility and practical potential of quantum SVMs in STAF (Software Testing at Early-stage and with Few labels) scenarios.