The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work challenges the prevailing view of large language models as mere “stochastic parrots” by introducing the Sequence-level Interactive Dynamic Parallel Processing (SIDPP) framework, which conceptualizes Transformers as systems that dynamically generate transformation parameters from input prompts to perform concept-to-concept mappings. The framework incorporates an output-weight interconnection mechanism that reveals a strong prompt sensitivity—where dynamic processing capacity intensifies with longer prompts—and suggests a potential correspondence with human cortical language processing. Experimental results demonstrate that such dynamic processing can contribute comparably to, or even surpass, static processing in model performance. These findings not only open new avenues for model interpretability and controllability but also provide theoretical foundations for developing compact, efficient architectures and advancing our understanding of human language cognition.
📝 Abstract
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models merely reproduce statistical regularities learned in training, we argue that Transformers construct and apply prompt-dependent transformations whose parameters are generated during inference. We call this form of computation SIDPP: Sequence-level Interactive Dynamic Parallel Processing. The Transformer is interpreted as a system that transforms concepts by means of concepts. Token vectors are the concepts to be transformed; parameterized transformations defined by matrices and vectors are the transforming concepts. These may be static, when fixed through training, or dynamic, when generated from the input sequence. Mechanically, they correspond to groups of simple neural networks. The Transformer's architectural novelty lies in output-weight interconnections, through which the outputs of some networks determine the weights of others, alongside ordinary output-input interconnections. By means of these interconnections, the system constructs transformations from the prompt and uses them to modify token representations. The contribution of dynamic processing grows with prompt length and may equal or exceed that of static processing, a phenomenon we call strong prompt sensitivity. This account bears on interpretability, predictability, control, and the design of smaller, more sustainable systems. Finally, since the human neural system possesses the mechanisms required to implement SIDPP, we argue that a form of SIDPP may, in principle, be neurally realized in the cerebral cortex. We therefore conjecture that human language processing may itself be a form of SIDPP produced by a functional architecture relevantly similar to that of the Transformer.
Problem

Research questions and friction points this paper is trying to address.

Transformer
dynamic processing
prompt sensitivity
SIDPP
output-weight interconnections
Innovation

Methods, ideas, or system contributions that make the work stand out.

SIDPP
output-weight interconnections
dynamic processing
strong prompt sensitivity
Transformer architecture
🔎 Similar Papers
No similar papers found.
Marco Giunti
Marco Giunti
University of Oxford
Concurrency TheoryProgramming LanguagesNetwork Security
F
Fabrizia Giulia Garavaglia
Dip. di Pedagogia, Psicologia, Filosofia, Università di Cagliari