HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives
本文提出HiLRP方法,通过分解注意力和分辨率降低操作为四种类型,统一解释不同架构的Vision Transformer,解决了现有归因方法在多样性ViT上的不适用问题。
本文提出HiLRP方法,通过分解注意力和分辨率降低操作为四种类型,统一解释不同架构的Vision Transformer,解决了现有归因方法在多样性ViT上的不适用问题。
This study addresses the serious threat posed by urea adulteration in milk to food safety by proposing a rapid, non-destructive, and low-cost detection method. For the first time, transmissive multispectral imaging is combined with controlled milk sample density (specific gravity of 1.032) to acquire images across twelve spectral bands ranging from 365 to 940 nm, enabling precise quantification of urea content without reagents or complex pretreatment. Quantitative models based on multiple linear regression (R² = 0.9599) and a feedforward neural network (R² = 0.9773) demonstrate high accuracy under controlled conditions and strong potential for on-site application. This approach offers a novel and practical strategy for screening urea adulteration in dairy products.
This study addresses the limitations of traditional regression methods, which often rely on restrictive assumptions of homoscedasticity and normally distributed residuals, and are prone to retransformation bias when log-transformations are applied. To overcome these issues, the authors propose a Copula-based regression framework that flexibly models the joint distribution of the response and covariates along with their marginal distributions, explicitly capturing heteroscedasticity and asymmetric dependence structures without imposing strong parametric assumptions. This work represents the first systematic application of Copulas to heteroscedastic regression prediction. The proposed method demonstrates substantially improved accuracy: in simulation studies, it achieves an average MAPE of 0.21, outperforming linear and log-linear models by 6%–33% and 24%–57%, respectively; its superior performance is further confirmed on the real-world Wages dataset.
This study addresses the challenge of fairly evaluating the contribution of visual backbone networks in solar irradiance forecasting, as existing approaches often modify multiple components simultaneously. To enable reproducible and equitable comparisons, the authors propose a controlled benchmarking protocol that isolates the visual encoder by fixing all other elements of the multimodal prediction pipeline—including historical weather encoding, clear-sky index normalization, and the fusion and regression heads—while only swapping the visual backbone. Experiments on the Folsom and NREL datasets validate the effectiveness of this approach: all tested backbones (ConvNeXt, Swin Transformer, VMamba, Spatial Mamba, and MambaVision) outperform the smart persistence baseline. Notably, VMamba Small and Swin Base achieve the lowest RMSE (~65.4 W/m²) on Folsom, while Swin Tiny yields the best performance on NREL (23.76 W/m²).
Modeling multivariate temporal data involving multiple entities poses significant challenges in efficiently capturing complex dependencies across entities, over time, and their intricate couplings. This work proposes a structure-aware single-block Transformer architecture that explicitly decomposes and models spatial, temporal, and cross-domain interactions through parallel spatial and temporal self-attention mechanisms coupled with bidirectional cross-attention. A learnable gating mechanism is introduced to effectively fuse information from multiple sources. By adopting explicit factorization instead of deep stacking, the proposed model achieves competitive or superior performance compared to more complex architectures across multiple tasks, despite having only 1.76 million parameters, thereby demonstrating its efficiency and effectiveness.
本文提出HiLRP方法,通过分解注意力和分辨率降低操作为四种类型,统一解释不同架构的Vision Transformer,解决了现有归因方法在多样性ViT上的不适用问题。
This study addresses the serious threat posed by urea adulteration in milk to food safety by proposing a rapid, non-destructive, and low-cost detection method. For the first time, transmissive multispectral imaging is combined with controlled milk sample density (specific gravity of 1.032) to acquire images across twelve spectral bands ranging from 365 to 940 nm, enabling precise quantification of urea content without reagents or complex pretreatment. Quantitative models based on multiple linear regression (R² = 0.9599) and a feedforward neural network (R² = 0.9773) demonstrate high accuracy under controlled conditions and strong potential for on-site application. This approach offers a novel and practical strategy for screening urea adulteration in dairy products.
This study addresses the limitations of traditional regression methods, which often rely on restrictive assumptions of homoscedasticity and normally distributed residuals, and are prone to retransformation bias when log-transformations are applied. To overcome these issues, the authors propose a Copula-based regression framework that flexibly models the joint distribution of the response and covariates along with their marginal distributions, explicitly capturing heteroscedasticity and asymmetric dependence structures without imposing strong parametric assumptions. This work represents the first systematic application of Copulas to heteroscedastic regression prediction. The proposed method demonstrates substantially improved accuracy: in simulation studies, it achieves an average MAPE of 0.21, outperforming linear and log-linear models by 6%–33% and 24%–57%, respectively; its superior performance is further confirmed on the real-world Wages dataset.
This study addresses the challenge of fairly evaluating the contribution of visual backbone networks in solar irradiance forecasting, as existing approaches often modify multiple components simultaneously. To enable reproducible and equitable comparisons, the authors propose a controlled benchmarking protocol that isolates the visual encoder by fixing all other elements of the multimodal prediction pipeline—including historical weather encoding, clear-sky index normalization, and the fusion and regression heads—while only swapping the visual backbone. Experiments on the Folsom and NREL datasets validate the effectiveness of this approach: all tested backbones (ConvNeXt, Swin Transformer, VMamba, Spatial Mamba, and MambaVision) outperform the smart persistence baseline. Notably, VMamba Small and Swin Base achieve the lowest RMSE (~65.4 W/m²) on Folsom, while Swin Tiny yields the best performance on NREL (23.76 W/m²).
Modeling multivariate temporal data involving multiple entities poses significant challenges in efficiently capturing complex dependencies across entities, over time, and their intricate couplings. This work proposes a structure-aware single-block Transformer architecture that explicitly decomposes and models spatial, temporal, and cross-domain interactions through parallel spatial and temporal self-attention mechanisms coupled with bidirectional cross-attention. A learnable gating mechanism is introduced to effectively fuse information from multiple sources. By adopting explicit factorization instead of deep stacking, the proposed model achieves competitive or superior performance compared to more complex architectures across multiple tasks, despite having only 1.76 million parameters, thereby demonstrating its efficiency and effectiveness.