Object Counting with GPT-4o and GPT-5: A Comparative Study

📅 2025-12-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Zero-shot object counting aims to estimate instance counts for unseen categories, yet existing approaches heavily rely on annotated data or visual exemplars. This paper introduces the first purely text-prompt-driven unsupervised zero-shot counting paradigm, leveraging the native multimodal perception capabilities of large multimodal language models (MLLMs)—specifically GPT-4o and GPT-5—without fine-tuning, training, or in-context image examples. We systematically design counting-specific prompting strategies enabling end-to-end inference on FSC-147 and CARPK. To our knowledge, this work presents the first comparative evaluation of GPT-4o and GPT-5 for zero-shot counting. Experiments demonstrate that our method achieves competitive, and in some cases superior, accuracy relative to state-of-the-art zero-shot baselines on FSC-147, validating that MLLMs can effectively model visual density and semantic relationships solely through textual prompts. This establishes a novel, training-free, example-free framework for general-purpose visual counting.

Technology Category

Application Category

📝 Abstract
Zero-shot object counting attempts to estimate the number of object instances belonging to novel categories that the vision model performing the counting has never encountered during training. Existing methods typically require large amount of annotated data and often require visual exemplars to guide the counting process. However, large language models (LLMs) are powerful tools with remarkable reasoning and data understanding abilities, which suggest the possibility of utilizing them for counting tasks without any supervision. In this work we aim to leverage the visual capabilities of two multi-modal LLMs, GPT-4o and GPT-5, to perform object counting in a zero-shot manner using only textual prompts. We evaluate both models on the FSC-147 and CARPK datasets and provide a comparative analysis. Our findings show that the models achieve performance comparable to the state-of-the-art zero-shot approaches on FSC-147, in some cases, even surpass them.
Problem

Research questions and friction points this paper is trying to address.

Zero-shot object counting without visual exemplars
Using multimodal LLMs for unsupervised counting tasks
Comparing GPT-4o and GPT-5 on benchmark datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using GPT-4o and GPT-5 for zero-shot object counting
Counting objects with only textual prompts, no visual exemplars
Achieving state-of-the-art performance on benchmark datasets
🔎 Similar Papers
R
Richard Fuzesséry
Software Engineering Institute, Obuda University, Budapest, Hungary
K
Kaziwa Saleh
John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary
S
Sándor Szénási
Faculty of Economics and Informatics, J. Selye University, Komárno, Slovakia
Z
Zoltán Vámossy
John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary