AGITB: A Signal-Level Benchmark for Evaluating Artificial General Intelligence

📅 2025-04-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing AGI evaluation frameworks lack practicality, incremental scalability, and interpretability, failing to capture the essential emergence of cognitive capabilities. Method: We introduce AGITB—the first signal-level benchmark for general intelligence—comprising 12 binary time-series prediction tasks. It abstracts away semantic and symbolic dependencies, evaluating cognition through computation-invariant dimensions: determinism, sensitivity, and generalization. Crucially, AGITB operates exclusively on raw, semantically uninterpreted signals, prohibiting pretraining, memorization, and brute-force search, and aligning evaluation with core computational principles observed in biological intelligence. Contribution/Results: Experiments show that humans universally pass all tasks, whereas no current AI system—including state-of-the-art LLMs—achieves full proficiency. AGITB thus establishes the first quantifiable, inescapable, and temporally evolving metric for measuring AGI capability progression.

Technology Category

Application Category

📝 Abstract
Despite remarkable progress in machine learning, current AI systems continue to fall short of true human-like intelligence. While Large Language Models (LLMs) excel in pattern recognition and response generation, they lack genuine understanding - an essential hallmark of Artificial General Intelligence (AGI). Existing AGI evaluation methods fail to offer a practical, gradual, and informative metric. This paper introduces the Artificial General Intelligence Test Bed (AGITB), comprising twelve rigorous tests that form a signal-processing-level foundation for the potential emergence of cognitive capabilities. AGITB evaluates intelligence through a model's ability to predict binary signals across time without relying on symbolic representations or pretraining. Unlike high-level tests grounded in language or perception, AGITB focuses on core computational invariants reflective of biological intelligence, such as determinism, sensitivity, and generalisation. The test bed assumes no prior bias, operates independently of semantic meaning, and ensures unsolvability through brute force or memorization. While humans pass AGITB by design, no current AI system has met its criteria, making AGITB a compelling benchmark for guiding and recognizing progress toward AGI.
Problem

Research questions and friction points this paper is trying to address.

Evaluating AI systems for true human-like intelligence capabilities
Providing a practical and gradual metric for AGI assessment
Focusing on core computational invariants of biological intelligence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Signal-level benchmark for AGI evaluation
Tests core computational invariants of intelligence
Avoids symbolic representations and pretraining
🔎 Similar Papers
No similar papers found.