HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出HiPerViT,通过在视觉变换器中引入二阶统计先验来解决多尺度纹理识别问题,提高了模型对纹理的敏感性。
📝 Abstract
Texture recognition remains challenging for modern vision models because discriminative evidence is often carried by higher-order spatial statistics rather than by object shape alone. While Vision Transformers provide strong long-range modeling capacity, their standard object-centric representations do not explicitly expose such statistical structure, which limits texture sensitivity in fine-grained recognition settings. We present HiPerViT, a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based recognition pipeline. The method combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. This design enables direct interaction between spatial tokens and second-order feature co-occurrence statistics, providing the model with explicit access to texture-relevant information without requiring multimodal pretraining or ensemble construction. Across six texture recognition benchmarks, HiPerViT achieves consistent improvements over strong vision-only baselines under the reported evaluation protocols, including gains of +3.05 percentage points on DTD, +10.48 on GTOS-Mobile, and +10.10 on 1200Tex. Beyond benchmark performance, our analyses show that these gains are largely invariant to the backbone depth used to extract second-order statistics and to the ordering of interaction and distillation stages. This pattern suggests that the primary source of improvement is not a specific fusion topology, but the explicit availability of second-order statistical information as a first-class representational signal. These results support explicit statistical tokenization as an effective and robust design principle for texture-centric visual recognition.
Problem

Research questions and friction points this paper is trying to address.

Texture Recognition
Higher-order Spatial Statistics
Vision Transformers
Statistical Structure
Fine-grained Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Perceiver-Vision Transformer
Second-order Statistical Prior
Bilinear Descriptor
Perceiver-style Latent Distillation
Texture Recognition
🔎 Similar Papers
No similar papers found.
J
João Pedro C. A. de Sá
Institute of Mathematics and Computer Sciences (ICMC), University of São Paulo (USP), Av. Trabalhador São-carlense, 400, São Carlos, SP 13566-590, Brazil
Odemir Martinez Bruno
Odemir Martinez Bruno
Professor, University of S. Paulo
Discrete Dynamical SystemsData ScienceNonlinear ScienceFractalsPattern Recognition