Institution profile

J. Selye University

Academic institutioneurope · sk
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Object Counting with GPT-4o and GPT-5: A Comparative Study

Dec 02, 2025

Zero-shot object counting aims to estimate instance counts for unseen categories, yet existing approaches heavily rely on annotated data or visual exemplars. This paper introduces the first purely text-prompt-driven unsupervised zero-shot counting paradigm, leveraging the native multimodal perception capabilities of large multimodal language models (MLLMs)—specifically GPT-4o and GPT-5—without fine-tuning, training, or in-context image examples. We systematically design counting-specific prompting strategies enabling end-to-end inference on FSC-147 and CARPK. To our knowledge, this work presents the first comparative evaluation of GPT-4o and GPT-5 for zero-shot counting. Experiments demonstrate that our method achieves competitive, and in some cases superior, accuracy relative to state-of-the-art zero-shot baselines on FSC-147, validating that MLLMs can effectively model visual density and semantic relationships solely through textual prompts. This establishes a novel, training-free, example-free framework for general-purpose visual counting.

0 citationsRead paper
Recent publications

Latest Papers

Object Counting with GPT-4o and GPT-5: A Comparative Study

Dec 02, 2025

Zero-shot object counting aims to estimate instance counts for unseen categories, yet existing approaches heavily rely on annotated data or visual exemplars. This paper introduces the first purely text-prompt-driven unsupervised zero-shot counting paradigm, leveraging the native multimodal perception capabilities of large multimodal language models (MLLMs)—specifically GPT-4o and GPT-5—without fine-tuning, training, or in-context image examples. We systematically design counting-specific prompting strategies enabling end-to-end inference on FSC-147 and CARPK. To our knowledge, this work presents the first comparative evaluation of GPT-4o and GPT-5 for zero-shot counting. Experiments demonstrate that our method achieves competitive, and in some cases superior, accuracy relative to state-of-the-art zero-shot baselines on FSC-147, validating that MLLMs can effectively model visual density and semantic relationships solely through textual prompts. This establishes a novel, training-free, example-free framework for general-purpose visual counting.

0 citationsRead paper