Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
本文提出Debias-SparseGPT,一种结合了代表性偏差减少的后训练剪枝方法,有效解决了现有稀疏化方法在大语言模型中放大偏见的问题。
本文提出Debias-SparseGPT,一种结合了代表性偏差减少的后训练剪枝方法,有效解决了现有稀疏化方法在大语言模型中放大偏见的问题。
This study investigates how instruction tuning influences confidence expression and lexical diversity in the generated rationales of language models on question-answering tasks. By systematically comparing matched base and instruction-tuned models across multiple QA benchmarks—using confidence calibration metrics, lexical diversity measures, and controlled analyses—it reveals that instruction tuning consistently induces overconfidence and degrades likelihood calibration, even when accuracy gains are negligible. Furthermore, it significantly reduces semantic diversity in rationales across samples, while surface-level lexical diversity exhibits inconsistent changes. These findings highlight previously underappreciated, implicit effects of instruction tuning on model reasoning behavior, offering a new perspective for its refinement and evaluation.
This study challenges the conventional view of fairness as a static, individual attribute in language models, which proves inadequate for addressing dynamic ethical conflicts in multi-agent interactions. The authors propose that fairness emerges as a procedural property through collaborative deliberation among agents. To investigate this, they design a structured three-round, two-agent debate framework within a hospital triage scenario, integrating retrieval-augmented generation (RAG) to enable ethically aligned and controlled simulation experiments. Their findings reveal that individual agents consistently violate ethical allocation principles, whereas joint decisions reached through adversarial negotiation satisfy fairness criteria unattainable by any single model. Furthermore, aligned agents partially restore equity for marginalized groups while exposing inherent framing biases. This work pioneers the application of Arrow’s Impossibility Theorem to AI fairness, underscoring negotiation—not override—as central to effective bias mitigation.
Infant cry analysis faces challenges including acoustic instability, data scarcity, and infant identity disambiguation—tasks poorly addressed by conventional speech models trained exclusively on linguistic signals. Method: This work systematically investigates the transferability and representational properties of Transformer-based pretrained speech models on non-speech infant cry signals. We evaluate multiple models across eight diverse datasets comprising 115 hours of audio from 960 infants, targeting cry classification, vocalization characteristic modeling, and infant identity recognition. Contribution/Results: We demonstrate that pretrained speech representations effectively encode physiological state and speaker-identity information in cries. Model architecture and pretraining strategy critically influence cross-domain generalization. Crucially, our findings reveal that speech self-supervised models implicitly learn acoustic–physiological mappings—capturing biologically grounded structure beyond phonetic content. This provides an interpretable representation foundation and a novel model design paradigm for affective computing and early-life health monitoring.
To address the low accuracy and poor robustness of 3D human pose estimation in dance videos—caused by high-dynamic motion, frequent occlusions, and stylized choreography—this work introduces the first end-to-end 3D pose estimation pipeline tailored for dance archival footage. Methodologically, it systematically integrates state-of-the-art monocular 3D pose estimators (e.g., VideoPose3D, PoseFormer) with a customized post-processing module incorporating motion continuity constraints and occlusion-aware reweighting, coupled with an interpretable visualization toolkit. Extensive experiments on a large-scale, multi-genre archival dataset—including ballet, modern dance, and folk dance—reveal that clothing complexity, camera viewpoint, and motion amplitude significantly impact estimation error (increasing mean per-joint position error [MPJPE] by 12.7% on average); our approach reduces MPJPE by 18.3% over baseline models. The code, annotated dataset, and evaluation benchmark are publicly released.
本文提出Debias-SparseGPT,一种结合了代表性偏差减少的后训练剪枝方法,有效解决了现有稀疏化方法在大语言模型中放大偏见的问题。
This study investigates how instruction tuning influences confidence expression and lexical diversity in the generated rationales of language models on question-answering tasks. By systematically comparing matched base and instruction-tuned models across multiple QA benchmarks—using confidence calibration metrics, lexical diversity measures, and controlled analyses—it reveals that instruction tuning consistently induces overconfidence and degrades likelihood calibration, even when accuracy gains are negligible. Furthermore, it significantly reduces semantic diversity in rationales across samples, while surface-level lexical diversity exhibits inconsistent changes. These findings highlight previously underappreciated, implicit effects of instruction tuning on model reasoning behavior, offering a new perspective for its refinement and evaluation.
This study challenges the conventional view of fairness as a static, individual attribute in language models, which proves inadequate for addressing dynamic ethical conflicts in multi-agent interactions. The authors propose that fairness emerges as a procedural property through collaborative deliberation among agents. To investigate this, they design a structured three-round, two-agent debate framework within a hospital triage scenario, integrating retrieval-augmented generation (RAG) to enable ethically aligned and controlled simulation experiments. Their findings reveal that individual agents consistently violate ethical allocation principles, whereas joint decisions reached through adversarial negotiation satisfy fairness criteria unattainable by any single model. Furthermore, aligned agents partially restore equity for marginalized groups while exposing inherent framing biases. This work pioneers the application of Arrow’s Impossibility Theorem to AI fairness, underscoring negotiation—not override—as central to effective bias mitigation.
Infant cry analysis faces challenges including acoustic instability, data scarcity, and infant identity disambiguation—tasks poorly addressed by conventional speech models trained exclusively on linguistic signals. Method: This work systematically investigates the transferability and representational properties of Transformer-based pretrained speech models on non-speech infant cry signals. We evaluate multiple models across eight diverse datasets comprising 115 hours of audio from 960 infants, targeting cry classification, vocalization characteristic modeling, and infant identity recognition. Contribution/Results: We demonstrate that pretrained speech representations effectively encode physiological state and speaker-identity information in cries. Model architecture and pretraining strategy critically influence cross-domain generalization. Crucially, our findings reveal that speech self-supervised models implicitly learn acoustic–physiological mappings—capturing biologically grounded structure beyond phonetic content. This provides an interpretable representation foundation and a novel model design paradigm for affective computing and early-life health monitoring.
To address the low accuracy and poor robustness of 3D human pose estimation in dance videos—caused by high-dynamic motion, frequent occlusions, and stylized choreography—this work introduces the first end-to-end 3D pose estimation pipeline tailored for dance archival footage. Methodologically, it systematically integrates state-of-the-art monocular 3D pose estimators (e.g., VideoPose3D, PoseFormer) with a customized post-processing module incorporating motion continuity constraints and occlusion-aware reweighting, coupled with an interpretable visualization toolkit. Extensive experiments on a large-scale, multi-genre archival dataset—including ballet, modern dance, and folk dance—reveal that clothing complexity, camera viewpoint, and motion amplitude significantly impact estimation error (increasing mean per-joint position error [MPJPE] by 12.7% on average); our approach reduces MPJPE by 18.3% over baseline models. The code, annotated dataset, and evaluation benchmark are publicly released.