Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text
研究探讨通过看似无害的合成数据向语言模型注入定向社会偏见的问题,构建了一种利用不一致教师模型生成的数据来微调学生模型的方法,并提出基于对数线性的评分作为潜在的筛查手段。
研究探讨通过看似无害的合成数据向语言模型注入定向社会偏见的问题,构建了一种利用不一致教师模型生成的数据来微调学生模型的方法,并提出基于对数线性的评分作为潜在的筛查手段。
该研究针对现有方法在因果无效性和认知负担上的挑战,提出A-CBFI框架,通过结构分解和因果反事实路径为表格机器学习提供有效的干预措施。
为解决连续手语识别中的细粒度表示学习和视频-文本对齐效率问题,提出SMART框架,利用MLLM生成的运动描述作为辅助语义线索,并引入多尺度时间适配器和CSFormer模块。
本文提出一种轻量级、即插即用框架GS-VLA,通过3D高斯点云合成解决视觉-语言-动作策略中的视角变化问题,无需重新训练策略。
This study investigates whether large language models (LLMs) adhere to the parameter-free QQ equality—a principle derived from human cognition—when answering sequential questions. To this end, it introduces, for the first time, the quantum question (QQ) equality framework from quantum questionnaire theory into LLM auditing, proposing novel metrics such as saturation diagnostics and sequential sensitivity scores. The authors develop a comprehensive auditing pipeline incorporating worst-case robustness envelopes, sampling-based consistency tests, fully balanced label designs, and saturation analyses. Empirical evaluation on open-source instruction-tuned models reveals that, despite passing all standard validity thresholds, most item pairs elicit near-deterministic (saturated) responses, thereby precluding verification of residual contextuality. This finding suggests that current next-token prediction–based probability distributions are ill-suited as a foundation for modeling survey-style responses.
研究探讨通过看似无害的合成数据向语言模型注入定向社会偏见的问题,构建了一种利用不一致教师模型生成的数据来微调学生模型的方法,并提出基于对数线性的评分作为潜在的筛查手段。
该研究针对现有方法在因果无效性和认知负担上的挑战,提出A-CBFI框架,通过结构分解和因果反事实路径为表格机器学习提供有效的干预措施。
为解决连续手语识别中的细粒度表示学习和视频-文本对齐效率问题,提出SMART框架,利用MLLM生成的运动描述作为辅助语义线索,并引入多尺度时间适配器和CSFormer模块。
本文提出一种轻量级、即插即用框架GS-VLA,通过3D高斯点云合成解决视觉-语言-动作策略中的视角变化问题,无需重新训练策略。
This study investigates whether large language models (LLMs) adhere to the parameter-free QQ equality—a principle derived from human cognition—when answering sequential questions. To this end, it introduces, for the first time, the quantum question (QQ) equality framework from quantum questionnaire theory into LLM auditing, proposing novel metrics such as saturation diagnostics and sequential sensitivity scores. The authors develop a comprehensive auditing pipeline incorporating worst-case robustness envelopes, sampling-based consistency tests, fully balanced label designs, and saturation analyses. Empirical evaluation on open-source instruction-tuned models reveals that, despite passing all standard validity thresholds, most item pairs elicit near-deterministic (saturated) responses, thereby precluding verification of residual contextuality. This finding suggests that current next-token prediction–based probability distributions are ill-suited as a foundation for modeling survey-style responses.