Institution profile

Hyperbots

Industry researchnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs

Jan 21, 2025

This study addresses the underexplored role of cognitive rigidity—specifically, “mental set”—in large language models (LLMs), investigating its impact on strategy switching and adaptive reasoning. Method: We systematically introduce mental set concepts from cognitive psychology into LLM evaluation, constructing a novel mental-set-inducing benchmark. Using parameter-efficient fine-tuning (PEFT), in-context learning (ICL), and cross-model analysis (Llama-3.1-8B/70B, GPT-4o), we evaluate performance on strategy-reversal tasks. Contribution/Results: We observe significant performance degradation (23–41%) across mainstream LLMs on reversal tasks, confirming entrenched pattern dependence. Critically, we move beyond static benchmarks (e.g., MMLU, GSM8K) to propose the first dynamic evaluation framework explicitly targeting cognitive flexibility in reasoning. This framework provides a new paradigm for diagnosing LLM reasoning bottlenecks and advancing generalization capabilities through adaptive inference.

0 citationsRead paper
Recent publications

Latest Papers

Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs

Jan 21, 2025

This study addresses the underexplored role of cognitive rigidity—specifically, “mental set”—in large language models (LLMs), investigating its impact on strategy switching and adaptive reasoning. Method: We systematically introduce mental set concepts from cognitive psychology into LLM evaluation, constructing a novel mental-set-inducing benchmark. Using parameter-efficient fine-tuning (PEFT), in-context learning (ICL), and cross-model analysis (Llama-3.1-8B/70B, GPT-4o), we evaluate performance on strategy-reversal tasks. Contribution/Results: We observe significant performance degradation (23–41%) across mainstream LLMs on reversal tasks, confirming entrenched pattern dependence. Critically, we move beyond static benchmarks (e.g., MMLU, GSM8K) to propose the first dynamic evaluation framework explicitly targeting cognitive flexibility in reasoning. This framework provides a new paradigm for diagnosing LLM reasoning bottlenecks and advancing generalization capabilities through adaptive inference.

0 citationsRead paper