Institution profile

Tangible Research Inc.

Industry research
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering

Jan 13, 2026

This study addresses the challenge of automatically evaluating large language models in political question-answering scenarios, where assessments must jointly consider factual correctness, response clarity, and evasion detection. The impact of prompt design on such high-level semantic tasks remains underexplored. Leveraging the CLARITY dataset from SemEval 2026, this work presents the first systematic evaluation of different prompting strategies—namely zero-shot prompting, chain-of-thought, and few-shot chain-of-thought—on GPT-3.5 and GPT-5.2 for clarity scoring and topic detection. Results show that GPT-5.2 achieves a clarity prediction accuracy of 63% under few-shot chain-of-thought prompting, up from 56%, and reaches 74% accuracy in topic identification. Evasion detection remains challenging, with peak performance at 34% and substantial difficulties in fine-grained category discrimination. The findings highlight both the efficacy and limitations of structured prompting in complex semantic evaluation tasks.

0 citationsRead paper
Recent publications

Latest Papers

Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering

Jan 13, 2026

This study addresses the challenge of automatically evaluating large language models in political question-answering scenarios, where assessments must jointly consider factual correctness, response clarity, and evasion detection. The impact of prompt design on such high-level semantic tasks remains underexplored. Leveraging the CLARITY dataset from SemEval 2026, this work presents the first systematic evaluation of different prompting strategies—namely zero-shot prompting, chain-of-thought, and few-shot chain-of-thought—on GPT-3.5 and GPT-5.2 for clarity scoring and topic detection. Results show that GPT-5.2 achieves a clarity prediction accuracy of 63% under few-shot chain-of-thought prompting, up from 56%, and reaches 74% accuracy in topic identification. Evasion detection remains challenging, with peak performance at 34% and substantial difficulties in fine-grained category discrimination. The findings highlight both the efficacy and limitations of structured prompting in complex semantic evaluation tasks.

0 citationsRead paper