Visualizing Class Specific Heterogeneous Tendencies using R
本文介绍了一种使用R包mccca实现的方法MCCCA,用于识别并可视化不同类别(如性别、国籍)特有的异质性趋势。
本文介绍了一种使用R包mccca实现的方法MCCCA,用于识别并可视化不同类别(如性别、国籍)特有的异质性趋势。
研究通过分析Visual Studio Code社区的讨论,使用关键词检索和语义过滤结合的方法,探讨了开发者如何讨论生成式AI工具及其在软件开发中的实际应用问题。
本文提出了一种基于后验协方差表示的高效敏感性分析方法(PosCoSeA),用于近似留k交叉验证和自助法重采样,以解决贝叶斯模型中不确定性度量不准确的问题。
This study addresses the challenge of imperfect detection and sample-loss bias in estimating population abundance when multiple survey protocols interfere with one another. The authors propose a two-stage estimation framework that first calibrates for sample loss and then estimates detection probability, thereby avoiding the feedback loops and weak identifiability inherent in joint modeling approaches. A novel sandwich-type robust variance estimator is introduced to effectively propagate uncertainty from the first stage and mitigate the impact of potential model misspecification. Compared to Bayesian joint models, this method yields more reliable uncertainty quantification, as demonstrated by its robustness and practical utility in an analysis of squirrel ectoparasite data.
This study addresses the challenge of detecting large language model (LLM)-generated code that evades conventional plagiarism detection through semantics-preserving rewrites. It presents the first systematic evaluation of Java bytecode-based k-gram software watermarking (with k ranging from 1 to 6) in the context of LLM-generated code. The approach integrates five similarity metrics—cosine similarity, Dice coefficient, Jaccard index, Simpson index, and edit distance–based similarity—and is evaluated on code produced by three prominent LLMs. Results demonstrate that the proposed watermarking technique effectively identifies LLM-assisted plagiarism, with code generated by domain-specialized models (e.g., ChatGPT-5.1-Codex-Mini) exhibiting greater stealthiness, thereby confirming that model specialization enhances the concealment of plagiarized content.
本文介绍了一种使用R包mccca实现的方法MCCCA,用于识别并可视化不同类别(如性别、国籍)特有的异质性趋势。
研究通过分析Visual Studio Code社区的讨论,使用关键词检索和语义过滤结合的方法,探讨了开发者如何讨论生成式AI工具及其在软件开发中的实际应用问题。
本文提出了一种基于后验协方差表示的高效敏感性分析方法(PosCoSeA),用于近似留k交叉验证和自助法重采样,以解决贝叶斯模型中不确定性度量不准确的问题。
This study addresses the challenge of imperfect detection and sample-loss bias in estimating population abundance when multiple survey protocols interfere with one another. The authors propose a two-stage estimation framework that first calibrates for sample loss and then estimates detection probability, thereby avoiding the feedback loops and weak identifiability inherent in joint modeling approaches. A novel sandwich-type robust variance estimator is introduced to effectively propagate uncertainty from the first stage and mitigate the impact of potential model misspecification. Compared to Bayesian joint models, this method yields more reliable uncertainty quantification, as demonstrated by its robustness and practical utility in an analysis of squirrel ectoparasite data.
This study addresses the challenge of detecting large language model (LLM)-generated code that evades conventional plagiarism detection through semantics-preserving rewrites. It presents the first systematic evaluation of Java bytecode-based k-gram software watermarking (with k ranging from 1 to 6) in the context of LLM-generated code. The approach integrates five similarity metrics—cosine similarity, Dice coefficient, Jaccard index, Simpson index, and edit distance–based similarity—and is evaluated on code produced by three prominent LLMs. Results demonstrate that the proposed watermarking technique effectively identifies LLM-assisted plagiarism, with code generated by domain-specialized models (e.g., ChatGPT-5.1-Codex-Mini) exhibiting greater stealthiness, thereby confirming that model specialization enhances the concealment of plagiarized content.