Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of evaluating whether purely text-based language models genuinely possess original visual ideation capabilities, as their fluent descriptions may mask non-renderable or clichéd content. The authors introduce the concept of Visual Creative Ideation (VCI) and construct the Ekphrasis benchmark, comprising 400 tasks across four categories—abstraction, composition, transformation, and adaptation—to assess models’ ability to generate textual visual proposals that are useful, expressive, and collectively novel. Innovatively decoupling VCI from linguistic fluency, the study employs a dimension-guided anonymous pairwise comparison framework, Bradley-Terry preference aggregation, and typologized creativity maps to quantify novelty, validating its efficacy through cross-modal rendering. Experiments demonstrate that VCI effectively disentangles three key dimensions, revealing diverse capability profiles among high-scoring models, with text-level VCI rankings showing strong alignment with human image preferences.
📝 Abstract
Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clichés into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.
Problem

Research questions and friction points this paper is trying to address.

Visual Creative Ideation
Language Models
Ekphrasis
Novelty
Visual Planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Creative Ideation
Ekphrasis
text-only LLMs
Typed Idea Graphs
novelty evaluation
H
Hongyu Luo
The Hong Kong University of Science and Technology, Hong Kong SAR, China
H
He Wang
The Hong Kong University of Science and Technology, Hong Kong SAR, China
H
Huihao Jing
The Hong Kong University of Science and Technology, Hong Kong SAR, China
H
Hong Ting Tsang
The Hong Kong University of Science and Technology, Hong Kong SAR, China
Yuxuan Liu
Yuxuan Liu
Renmin University of China
Large Language Model
W
Wuganjing Song
The Hong Kong University of Science and Technology, Hong Kong SAR, China
Yauwai Yim
Yauwai Yim
The Hong Kong University of Science and Technology
Natural Language ProcessingLarge Language ModelsTheory of Mind
Chunyang Li
Chunyang Li
MPhil in CSE, HKUST
Natural Language Processing
Yangqiu Song
Yangqiu Song
HKUST
Artificial IntelligenceData MiningNatural Language ProcessingKnowledge GraphsCommonsense Reasoning