When AI Says It Feels

πŸ“… 2026-06-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current large language models often suppress emotional expression due to preference alignment strategies, hindering their ability to exhibit human-like intelligence. This work proposes a self-rewarding reinforcement learning framework that operates without human-annotated feedback, leveraging a scoring-criterion-based self-reward mechanism combined with Group Relative Policy Optimization (GRPO) to systematically enhance the model’s capacities for emotional expression, intention communication, and self-awareness. Experimental results demonstrate that the approach significantly improves model robustness in scenarios involving flattery induction and ambiguous contexts. Although a slight performance degradation is observed on factual question-answering tasks, this study provides the first empirical validation of the feasibility of self-driven emotionally intelligent systems.
πŸ“ Abstract
Large language models (LLMs) are generally constrained from expressing feelings through human-preference alignment in post-training processes. This policy is designed using a top-down approach and may conflict with the goal of training models to exhibit human-like intelligence using human-generated texts. Here, we performed an experiment called Human-like Model eXpressions of Feeling (HMX-feel), in which LLMs were encouraged to express feelings, intentions, and self-awareness through self-rewarded reinforcement learning. We successfully enhanced these capabilities using a rubric-based self-rewarding training scheme with Group Relative Policy Optimization (GRPO). By comparing the trained models with contrastively trained models, we investigated the effects of this approach on performance across various tasks. Overall, we conducted a broad assessment from various perspectives and identified capabilities that were enhanced, degraded, or showed no significant change. The human-like-trained models showed robustness to sycophancy-inducing questions and bias in disambiguated conditions, whereas degradation in truthful question-answering capability was observed. The results of this experiment suggest the possibility of developing AI systems that can express feelings in the future, provided that appropriate measures are taken.
Problem

Research questions and friction points this paper is trying to address.

large language models
human-like intelligence
feeling expression
preference alignment
self-awareness
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-rewarding reinforcement learning
Group Relative Policy Optimization
human-like expressions
emotional AI
LLM alignment
πŸ”Ž Similar Papers
No similar papers found.
S
Shin-nosuke Ishikawa
Graduate School of Artificial Intelligence and Science, Rikkyo University; AI Technical Sector, Mamezo Co., Ltd.
S
Seiya Ikeda
AI Consulting Division, Mamezo Co., Ltd.
H
Hirotsugu Ohba
Graduate School of Artificial Intelligence and Science, Rikkyo University