Institution profile

Wand AI

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

May 21, 2025

This study investigates whether large language models (LLMs) achieve a human-like trade-off between semantic fidelity and compression efficiency in their internal representations. Method: We introduce the first quantitative framework grounded in rate-distortion theory and the information bottleneck principle, enabling systematic comparison of LLM embeddings against human categorization behavior on canonical cognitive benchmarks. Contribution/Results: We find that while LLMs capture coarse-grained, human-aligned concepts, they exhibit significantly weaker fine-grained semantic discrimination than humans. Their representations overemphasize statistical compression at the expense of semantic nuance and lack contextual adaptivity. These findings reveal a fundamental cognitive divergence between LLMs and humans in concept formation and establish the first information-theoretic, interpretable paradigm for evaluating and improving the semantic representational capacity of language models.

0 citationsRead paper

Concise Reasoning via Reinforcement Learning

Apr 07, 2025

Chain-of-thought (CoT) reasoning in large language models incurs excessive token consumption and latency due to inherent redundancy, yet the theoretical origins of such redundancy—particularly under reinforcement learning (RL)—remain unexamined. Method: This work first establishes, theoretically, that standard RL optimization inherently induces redundant CoT generation; it further identifies an intrinsic positive correlation between CoT conciseness and reasoning accuracy. Building on these insights, we propose a lightweight, two-stage RL-based pruning paradigm that avoids full retraining: it integrates Proximal Policy Optimization (PPO), few-shot reward modeling, and CoT distillation. Results: Evaluated across multiple mathematical and logical reasoning benchmarks, our method achieves an average 42% reduction in CoT length while maintaining or improving accuracy by up to 1.3%, significantly lowering computational cost and inference latency.

0 citationsRead paper

Layer by Layer: Uncovering Hidden Representations in Language Models

Feb 04, 2025

This work challenges the conventional assumption that final-layer representations in large language models (LLMs) are optimal, revealing instead that intermediate-layer hidden states encode richer and more robust semantic information. Method: We propose the first multidimensional representation quality evaluation framework integrating information-theoretic measures (mutual information, compression ratio), manifold geometry, and perturbation invariance—designed for cross-architectural (Transformer/SSM) and cross-modal (text/vision) validation. Contribution/Results: Evaluated on 32 text embedding benchmarks, intermediate-layer embeddings consistently outperform final-layer counterparts by an average of 4.2%, demonstrating both statistical consistency and strong generalization across tasks and architectures. This study provides the first empirical evidence establishing the superiority of intermediate-layer representations, thereby introducing a new paradigm for efficient representation extraction, model compression, and interpretability research.

0 citationsRead paper

"Yeah Right!"-- Do LLMs Exhibit Multimodal Feature Transfer?

Jan 07, 2025

This study investigates large language models’ (LLMs) cross-domain capability to comprehend implicit deception in human dialogue—such as irony and sarcasm—and examines whether multimodal skill transfer enhances pragmatic awareness. Method: We propose a novel paradigm integrating speech-text joint pretraining with fine-tuning on human-to-human conversational data, and systematically evaluate both multimodal (speech+text) and unimodal (text-only) LLMs on zero-shot deceptive utterance detection. Contribution/Results: Experimental results demonstrate that either speech-text multimodal pretraining or human dialogue fine-tuning alone significantly improves implicit deception recognition; their combination yields further gains. Crucially, multimodal models surpass text-only baselines without additional prompting. This work provides the first empirical evidence that multimodal pretraining facilitates cross-modal intent representation learning and transferable implicit semantic understanding—offering a new pathway toward socially aware conversational AI.

0 citationsRead paper
Recent publications

Latest Papers

From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

May 21, 2025

This study investigates whether large language models (LLMs) achieve a human-like trade-off between semantic fidelity and compression efficiency in their internal representations. Method: We introduce the first quantitative framework grounded in rate-distortion theory and the information bottleneck principle, enabling systematic comparison of LLM embeddings against human categorization behavior on canonical cognitive benchmarks. Contribution/Results: We find that while LLMs capture coarse-grained, human-aligned concepts, they exhibit significantly weaker fine-grained semantic discrimination than humans. Their representations overemphasize statistical compression at the expense of semantic nuance and lack contextual adaptivity. These findings reveal a fundamental cognitive divergence between LLMs and humans in concept formation and establish the first information-theoretic, interpretable paradigm for evaluating and improving the semantic representational capacity of language models.

0 citationsRead paper

Concise Reasoning via Reinforcement Learning

Apr 07, 2025

Chain-of-thought (CoT) reasoning in large language models incurs excessive token consumption and latency due to inherent redundancy, yet the theoretical origins of such redundancy—particularly under reinforcement learning (RL)—remain unexamined. Method: This work first establishes, theoretically, that standard RL optimization inherently induces redundant CoT generation; it further identifies an intrinsic positive correlation between CoT conciseness and reasoning accuracy. Building on these insights, we propose a lightweight, two-stage RL-based pruning paradigm that avoids full retraining: it integrates Proximal Policy Optimization (PPO), few-shot reward modeling, and CoT distillation. Results: Evaluated across multiple mathematical and logical reasoning benchmarks, our method achieves an average 42% reduction in CoT length while maintaining or improving accuracy by up to 1.3%, significantly lowering computational cost and inference latency.

0 citationsRead paper

Layer by Layer: Uncovering Hidden Representations in Language Models

Feb 04, 2025

This work challenges the conventional assumption that final-layer representations in large language models (LLMs) are optimal, revealing instead that intermediate-layer hidden states encode richer and more robust semantic information. Method: We propose the first multidimensional representation quality evaluation framework integrating information-theoretic measures (mutual information, compression ratio), manifold geometry, and perturbation invariance—designed for cross-architectural (Transformer/SSM) and cross-modal (text/vision) validation. Contribution/Results: Evaluated on 32 text embedding benchmarks, intermediate-layer embeddings consistently outperform final-layer counterparts by an average of 4.2%, demonstrating both statistical consistency and strong generalization across tasks and architectures. This study provides the first empirical evidence establishing the superiority of intermediate-layer representations, thereby introducing a new paradigm for efficient representation extraction, model compression, and interpretability research.

0 citationsRead paper

"Yeah Right!"-- Do LLMs Exhibit Multimodal Feature Transfer?

Jan 07, 2025

This study investigates large language models’ (LLMs) cross-domain capability to comprehend implicit deception in human dialogue—such as irony and sarcasm—and examines whether multimodal skill transfer enhances pragmatic awareness. Method: We propose a novel paradigm integrating speech-text joint pretraining with fine-tuning on human-to-human conversational data, and systematically evaluate both multimodal (speech+text) and unimodal (text-only) LLMs on zero-shot deceptive utterance detection. Contribution/Results: Experimental results demonstrate that either speech-text multimodal pretraining or human dialogue fine-tuning alone significantly improves implicit deception recognition; their combination yields further gains. Crucially, multimodal models surpass text-only baselines without additional prompting. This work provides the first empirical evidence that multimodal pretraining facilitates cross-modal intent representation learning and transferable implicit semantic understanding—offering a new pathway toward socially aware conversational AI.

0 citationsRead paper