A Removal Based Approach to Improve LLM Faithfulness at Test-Time
该研究提出一种测试时移除未在解释中提及的概念的方法,以提高大型语言模型解释的完整性,从而改善其决策的可靠性和安全性。
该研究提出一种测试时移除未在解释中提及的概念的方法,以提高大型语言模型解释的完整性,从而改善其决策的可靠性和安全性。
Underrepresented students—particularly those from humanities backgrounds and non-STEM disciplines—often face significant barriers to engaging with probabilistic machine learning (ML) due to high mathematical thresholds and perceived irrelevance to societal concerns. Method: This study introduces a “framework-focused” pedagogy for an undergraduate probability ML course, anchored in an interdisciplinary narrative—the fictional “Interstellar Hypothetical Hospital”—and integrating probabilistic programming to lower mathematical entry barriers. The curriculum employs open, real-world case studies and counter-narrative discussions to systematically connect foundational ML concepts (e.g., Bayesian modeling) with sociotechnical implications. Contribution/Results: The approach innovatively merges whimsical storytelling with rigorous probabilistic reasoning to enhance accessibility and engagement; uses ethical dilemmas as anchors for developing dialectical AI literacy; and concurrently cultivates modeling competence, critical thinking, and technocivic identity. Empirical evaluation demonstrates significant gains in students’ integrated understanding of ML principles and their social dimensions, alongside increased confidence and capacity to participate diversely in public AI discourse.
Traditional topic modeling methods suffer from poor interpretability and limited capacity to address domain-specific research questions in the social sciences. Method: This paper proposes a large language model (LLM)-based data augmentation framework that integrates controllable semantic text generation into unsupervised topic modeling. Leveraging GPT-4 to synthesize domain-relevant textual data, the approach couples generated corpora with LDA and BERTopic for guided, question-oriented topic discovery—requiring minimal human intervention. A political science–specific corpus and evaluation framework are constructed to support rigorous validation. Contribution/Results: Experiments demonstrate substantial improvements in topic interpretability and task relevance; the method enables direct answering of domain research questions and reduces manual annotation effort by over 70%. By bridging generative AI with social science–driven topic modeling, this work establishes a novel paradigm for theory-informed, question-centered thematic analysis.
该研究提出一种测试时移除未在解释中提及的概念的方法,以提高大型语言模型解释的完整性,从而改善其决策的可靠性和安全性。
Underrepresented students—particularly those from humanities backgrounds and non-STEM disciplines—often face significant barriers to engaging with probabilistic machine learning (ML) due to high mathematical thresholds and perceived irrelevance to societal concerns. Method: This study introduces a “framework-focused” pedagogy for an undergraduate probability ML course, anchored in an interdisciplinary narrative—the fictional “Interstellar Hypothetical Hospital”—and integrating probabilistic programming to lower mathematical entry barriers. The curriculum employs open, real-world case studies and counter-narrative discussions to systematically connect foundational ML concepts (e.g., Bayesian modeling) with sociotechnical implications. Contribution/Results: The approach innovatively merges whimsical storytelling with rigorous probabilistic reasoning to enhance accessibility and engagement; uses ethical dilemmas as anchors for developing dialectical AI literacy; and concurrently cultivates modeling competence, critical thinking, and technocivic identity. Empirical evaluation demonstrates significant gains in students’ integrated understanding of ML principles and their social dimensions, alongside increased confidence and capacity to participate diversely in public AI discourse.
Traditional topic modeling methods suffer from poor interpretability and limited capacity to address domain-specific research questions in the social sciences. Method: This paper proposes a large language model (LLM)-based data augmentation framework that integrates controllable semantic text generation into unsupervised topic modeling. Leveraging GPT-4 to synthesize domain-relevant textual data, the approach couples generated corpora with LDA and BERTopic for guided, question-oriented topic discovery—requiring minimal human intervention. A political science–specific corpus and evaluation framework are constructed to support rigorous validation. Contribution/Results: Experiments demonstrate substantial improvements in topic interpretability and task relevance; the method enables direct answering of domain research questions and reduces manual annotation effort by over 70%. By bridging generative AI with social science–driven topic modeling, this work establishes a novel paradigm for theory-informed, question-centered thematic analysis.