When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
研究通过构建包含复杂组成因素的图像缺陷数据集CO-AID,利用人类评估来识别AI生成图像中的缺陷,并训练深度模型以预测和优化这些缺陷。
研究通过构建包含复杂组成因素的图像缺陷数据集CO-AID,利用人类评估来识别AI生成图像中的缺陷,并训练深度模型以预测和优化这些缺陷。
This study addresses the challenge of poor image quality in low-dose cone-beam computed tomography (CBCT), which suffers from scatter, noise, and artifacts that hinder accurate dose calculation and adaptive radiotherapy. To overcome this limitation, the authors propose the first supervised CBCT-to-CT synthesis framework based on a conditional denoising diffusion probabilistic model (DDPM), capable of generating high-fidelity CT-equivalent images. A key innovation lies in the systematic comparison of clinical DICOM images versus FDK-reconstructed images from raw projections as inputs, revealing that the latter significantly enhances CT equivalence. Experimental results demonstrate that the proposed method effectively supports precise patient positioning and accurate dose computation within low-dose radiotherapy workflows.
This work addresses the vulnerability of existing large language model (LLM) authorship attribution methods, which rely on surface-level linguistic features and are thus susceptible to paraphrasing, back-translation, and other obfuscation attacks, limiting their generalizability. To overcome this limitation, the authors propose a novel approach that integrates reasoning graphs with graph neural networks for the first time. Specifically, they employ an argument mining pipeline to extract deep reasoning structures from LLM-generated texts, construct corresponding reasoning graphs, and apply graph neural networks for authorship attribution. Experimental results demonstrate that the proposed method achieves up to a 27-percentage-point improvement in accuracy under obfuscation attacks and a 19-percentage-point gain on unseen LLM versions, substantially outperforming Longformer-based baselines and surpassing the robustness ceiling of conventional surface-feature approaches.
This work proposes an efficient Boltzmann sampling algorithm for the power set of combinatorial structures with bounded counting sequences, eliminating the need for generating function evaluations or external oracles. By leveraging the intrinsic counting properties of the underlying structures, the method achieves, for the first time, Boltzmann sampling of power sets without relying on generating functions, thereby overcoming a key limitation of traditional approaches. Experimental results demonstrate that the proposed algorithm matches the runtime performance of existing Boltzmann samplers, confirming its efficiency and practical feasibility.
This study addresses optic disc segmentation in retinal images—a fundamental task in medical image analysis—by pioneering the adaptation of the RETFound vision foundation model to medical image segmentation. We propose a lightweight segmentation head coupled with end-to-end fine-tuning, enabling efficient transfer learning with only a small number of annotated samples. Unlike prior work focused on diagnostic applications, this is the first effort to evaluate RETFound’s generalization capability on non-diagnostic downstream vision tasks, demonstrating its potential to replace task-specific architectures. Evaluated on five public datasets, our method achieves a Dice coefficient of approximately 96%, substantially outperforming most existing approaches. It further establishes new state-of-the-art performance in internal validation, cross-domain generalization, and domain adaptation. Collectively, this work introduces a novel paradigm for the general-purpose deployment of medical foundation models.
研究通过构建包含复杂组成因素的图像缺陷数据集CO-AID,利用人类评估来识别AI生成图像中的缺陷,并训练深度模型以预测和优化这些缺陷。
This study addresses the challenge of poor image quality in low-dose cone-beam computed tomography (CBCT), which suffers from scatter, noise, and artifacts that hinder accurate dose calculation and adaptive radiotherapy. To overcome this limitation, the authors propose the first supervised CBCT-to-CT synthesis framework based on a conditional denoising diffusion probabilistic model (DDPM), capable of generating high-fidelity CT-equivalent images. A key innovation lies in the systematic comparison of clinical DICOM images versus FDK-reconstructed images from raw projections as inputs, revealing that the latter significantly enhances CT equivalence. Experimental results demonstrate that the proposed method effectively supports precise patient positioning and accurate dose computation within low-dose radiotherapy workflows.
This work addresses the vulnerability of existing large language model (LLM) authorship attribution methods, which rely on surface-level linguistic features and are thus susceptible to paraphrasing, back-translation, and other obfuscation attacks, limiting their generalizability. To overcome this limitation, the authors propose a novel approach that integrates reasoning graphs with graph neural networks for the first time. Specifically, they employ an argument mining pipeline to extract deep reasoning structures from LLM-generated texts, construct corresponding reasoning graphs, and apply graph neural networks for authorship attribution. Experimental results demonstrate that the proposed method achieves up to a 27-percentage-point improvement in accuracy under obfuscation attacks and a 19-percentage-point gain on unseen LLM versions, substantially outperforming Longformer-based baselines and surpassing the robustness ceiling of conventional surface-feature approaches.
This work proposes an efficient Boltzmann sampling algorithm for the power set of combinatorial structures with bounded counting sequences, eliminating the need for generating function evaluations or external oracles. By leveraging the intrinsic counting properties of the underlying structures, the method achieves, for the first time, Boltzmann sampling of power sets without relying on generating functions, thereby overcoming a key limitation of traditional approaches. Experimental results demonstrate that the proposed algorithm matches the runtime performance of existing Boltzmann samplers, confirming its efficiency and practical feasibility.
This study addresses optic disc segmentation in retinal images—a fundamental task in medical image analysis—by pioneering the adaptation of the RETFound vision foundation model to medical image segmentation. We propose a lightweight segmentation head coupled with end-to-end fine-tuning, enabling efficient transfer learning with only a small number of annotated samples. Unlike prior work focused on diagnostic applications, this is the first effort to evaluate RETFound’s generalization capability on non-diagnostic downstream vision tasks, demonstrating its potential to replace task-specific architectures. Evaluated on five public datasets, our method achieves a Dice coefficient of approximately 96%, substantially outperforming most existing approaches. It further establishes new state-of-the-art performance in internal validation, cross-domain generalization, and domain adaptation. Collectively, this work introduces a novel paradigm for the general-purpose deployment of medical foundation models.