skchange: Fast and Flexible Algorithms for Changepoint Detection
该论文介绍了一个名为skchange的开源Python库,用于时间序列中的结构变化检测,采用成本最小化和统计测试等方法,并具有快速搜索、高维数据处理等功能。
该论文介绍了一个名为skchange的开源Python库,用于时间序列中的结构变化检测,采用成本最小化和统计测试等方法,并具有快速搜索、高维数据处理等功能。
本文解决了数据流中实时检测分布变化的问题,通过gridcp包使用网格方法和多种统计测试实现高效在线检测。
This study addresses the limitation of conditional view independence in existing variational incomplete multi-view clustering methods by proposing a novel variational framework that explicitly models cross-view correlations. By parameterizing the covariance matrix via normalized Cholesky decomposition, the approach adaptively captures intrinsic data structures through posterior estimation errors under a unified optimization objective. This method overcomes the independence bottleneck with minimal additional parameters and consistently outperforms state-of-the-art baselines across multiple benchmark datasets. These results effectively validate the critical role of explicit cross-view correlation modeling in enhancing clustering performance for incomplete multi-view learning scenarios.
This work addresses the challenge of efficiently and accurately detecting whether a given text is partially or fully contained—particularly via near verbatim copying—within massive web-scale corpora. To this end, the authors introduce FindMyText, an open-source tool that leverages document fingerprinting with a novel mechanism for identifying contiguous matching fingerprint sequences, explicitly capturing near-exact copied segments rather than relying on holistic text similarity. The system incorporates a distributed disk-based index to enable scalable processing. The study also establishes the first benchmark specifically designed for text containment tasks, demonstrating that FindMyText significantly outperforms existing methods across diverse datasets including arXiv, Wikipedia, and general web corpora, thereby validating its efficiency, robustness, and practical utility.
This study addresses a critical limitation in clinical prediction: conventional approaches for assessing the importance of high-dimensional features—such as genomic data—rely solely on changes in model performance, thereby neglecting collinearity and causal directionality among variables, which introduces estimation bias. To overcome this, the work introduces asymmetric Shapley values into the domain for the first time, explicitly modeling causal dependencies among variables under causal graph assumptions. It proposes an efficient algorithm to compute both local and global feature contributions and enables decomposition with respect to arbitrary predictive performance metrics. Evaluated on a colorectal cancer progression-free survival prediction task, the method not only enhances the reliability of feature importance estimation but also provides a powerful tool for individualized inference and interpretable analysis.
该论文介绍了一个名为skchange的开源Python库,用于时间序列中的结构变化检测,采用成本最小化和统计测试等方法,并具有快速搜索、高维数据处理等功能。
本文解决了数据流中实时检测分布变化的问题,通过gridcp包使用网格方法和多种统计测试实现高效在线检测。
This study addresses the limitation of conditional view independence in existing variational incomplete multi-view clustering methods by proposing a novel variational framework that explicitly models cross-view correlations. By parameterizing the covariance matrix via normalized Cholesky decomposition, the approach adaptively captures intrinsic data structures through posterior estimation errors under a unified optimization objective. This method overcomes the independence bottleneck with minimal additional parameters and consistently outperforms state-of-the-art baselines across multiple benchmark datasets. These results effectively validate the critical role of explicit cross-view correlation modeling in enhancing clustering performance for incomplete multi-view learning scenarios.
This work addresses the challenge of efficiently and accurately detecting whether a given text is partially or fully contained—particularly via near verbatim copying—within massive web-scale corpora. To this end, the authors introduce FindMyText, an open-source tool that leverages document fingerprinting with a novel mechanism for identifying contiguous matching fingerprint sequences, explicitly capturing near-exact copied segments rather than relying on holistic text similarity. The system incorporates a distributed disk-based index to enable scalable processing. The study also establishes the first benchmark specifically designed for text containment tasks, demonstrating that FindMyText significantly outperforms existing methods across diverse datasets including arXiv, Wikipedia, and general web corpora, thereby validating its efficiency, robustness, and practical utility.
This study addresses a critical limitation in clinical prediction: conventional approaches for assessing the importance of high-dimensional features—such as genomic data—rely solely on changes in model performance, thereby neglecting collinearity and causal directionality among variables, which introduces estimation bias. To overcome this, the work introduces asymmetric Shapley values into the domain for the first time, explicitly modeling causal dependencies among variables under causal graph assumptions. It proposes an efficient algorithm to compute both local and global feature contributions and enables decomposition with respect to arbitrary predictive performance metrics. Evaluated on a colorectal cancer progression-free survival prediction task, the method not only enhances the reliability of feature importance estimation but also provides a powerful tool for individualized inference and interpretable analysis.