Institution profile

Kunumi

Research institution
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

Aug 05, 2026

Online misinformation detection faces challenges in scalability and reliance on external knowledge. This work proposes a novel method that requires no fine-tuning, retrieval, or task-specific supervision, treating truthfulness as a geometric property within the representation space of pretrained language models. By contrasting activation patterns between true and false statements, the approach identifies a “falsehood direction” in the residual stream and classifies inputs via projection of their final-layer activations. Combining contrastive activation addition (CAA) with an MLP classifier, the method demonstrates strong performance across mainstream architectures—including Gemma, Llama, and Qwen—matching or surpassing zero- and few-shot prompting on benchmarks LIAR and FACTors, with particularly notable gains for smaller models. Its performance is limited on AVeriTeC, which relies on annotated evidence, underscoring the method’s paradigm of evidence-free detection.

0 citationsRead paper
Recent publications

Latest Papers

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

Aug 05, 2026

Online misinformation detection faces challenges in scalability and reliance on external knowledge. This work proposes a novel method that requires no fine-tuning, retrieval, or task-specific supervision, treating truthfulness as a geometric property within the representation space of pretrained language models. By contrasting activation patterns between true and false statements, the approach identifies a “falsehood direction” in the residual stream and classifies inputs via projection of their final-layer activations. Combining contrastive activation addition (CAA) with an MLP classifier, the method demonstrates strong performance across mainstream architectures—including Gemma, Llama, and Qwen—matching or surpassing zero- and few-shot prompting on benchmarks LIAR and FACTors, with particularly notable gains for smaller models. Its performance is limited on AVeriTeC, which relies on annotated evidence, underscoring the method’s paradigm of evidence-free detection.

0 citationsRead paper