A Smaller Transformer in Your Transformer

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过提出Transformer-Within-Transformer方法,解决了Vision Transformers中的计算冗余问题,减少了参数数量和推理计算量,同时保持了模型性能。
📝 Abstract
Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
computational redundancy
inference compute
model expressivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer-Within-Transformer
redundant layers
inference compute
parameter count
model expressivity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dhananjay Tomar
Institute for Cancer Genetics and Informatics, Oslo University Hospital, Oslo, Norway; Department of Informatics, University of Oslo, Oslo, Norway; SFI Visual Intelligence, UiT The Arctic University of Norway, Tromsø, Norway
Marius Aasan
Marius Aasan
University of Oslo
Machine LearningImagingProbabilistic Machine Learning
A
Andreas Kleppe
Institute for Cancer Genetics and Informatics, Oslo University Hospital, Oslo, Norway; Department of Informatics, University of Oslo, Oslo, Norway; SFI Visual Intelligence, UiT The Arctic University of Norway, Tromsø, Norway
Adín Ramírez Rivera
Adín Ramírez Rivera
Professor, University of Oslo
Image ProcessingComputer VisionMachine Learning