Visual General Intelligence: A White Paper

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了视觉模态在通向通用人工智能(AGI)路径上的潜力,通过分析视觉智能的定义、原则、输入形式、基准及与其他模态如语言的关系。
📝 Abstract
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.
Problem

Research questions and friction points this paper is trying to address.

Visual General Intelligence
AGI
visual modalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual General Intelligence
Vision-centered Perspective
Autoregressive Modeling
Web-scale Learning
Intelligence from Visual Experience
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Hirokatsu Kataoka
Hirokatsu Kataoka
AIST / University of Oxford
Computer VisionAction RecognitionAction PredictionVisual Pre-trainingFDSL
Y
Yoshihiro Fukuhara
National Institute of Advanced Industrial Science and Technology (AIST)
Yonglong Tian
Yonglong Tian
Research Scientist, OpenAI
Artificial IntelligenceDeep LearningMachine LearningComputer Vision
Shangzhe Wu
Shangzhe Wu
University of Cambridge
Computer VisionInverse Rendering
O
Oishi Deb
Visual Geometry Group (VGG), University of Oxford
Ryousuke Yamada
Ryousuke Yamada
University of Tsukuba
Computer Vision3D Object Recognition
Christian Rupprecht
Christian Rupprecht
University of Oxford
Machine LearningComputer Vision
Jianyuan Wang
Jianyuan Wang
Oxford Visual Geometry Group & FAIR
K
Kohsuke Ide
National Institute of Advanced Industrial Science and Technology (AIST)
K
Koichi Namekata
Visual Geometry Group (VGG), University of Oxford
Xianzheng Ma
Xianzheng Ma
VGG & AVL, University of Oxford
3D Computer VisionEmbodied AILarge Language Models
Y
Yiming Chen
Visual Geometry Group (VGG), University of Oxford
Robert Geirhos
Robert Geirhos
Research Scientist, Google DeepMind
Understanding DNNsDeep LearningVideoHuman VisionPsychophysics
Aditi Raghunathan
Aditi Raghunathan
Assistant professor, Carnegie Mellon University
Yuki M. Asano
Yuki M. Asano
Full Professor, Head of FunAI Lab, University of Technology Nuremberg
Deep LearningMultimodal LearningSelf-supervised LearningLarge Model AdaptationLLMs
Deva Ramanan
Deva Ramanan
Professor, Robotics Institute, Carnegie Mellon University
Computer VisionMachine Learning
David Fouhey
David Fouhey
New York University
Computer VisionMachine LearningAI for ScienceSolar Physics
A
Andrew J. Davison
Imperial College London
Yilun Du
Yilun Du
Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
Jiajun Wu
Jiajun Wu
Stanford University
Computer VisionRoboticsArtificial IntelligenceMachine LearningCognitive Science
Zhuang Liu
Zhuang Liu
Assistant Professor, Princeton University
Deep LearningComputer VisionMachine Learning