Topology inside NC$^1$
本文探讨了ACC^0和NC^1电路复杂性类,通过研究不同拓扑结构如多对数属、交叉数和厚度,证明了特定条件下这些属性对计算能力的影响。
本文探讨了ACC^0和NC^1电路复杂性类,通过研究不同拓扑结构如多对数属、交叉数和厚度,证明了特定条件下这些属性对计算能力的影响。
It remains unclear whether current video-language models genuinely comprehend the temporal dynamics, motion, and semantics of videos. To address this, this work proposes REVEAL—the first systematic diagnostic benchmark—comprising five controlled stress tests: reversed playback, false statement injection, spatiotemporal occlusion, camera motion simulation, and language shortcut interference. Coupled with an automated pipeline for generating controllable perturbations, REVEAL comprehensively evaluates model robustness in fundamental perception and reasoning. Experiments reveal that both leading open-source and proprietary models exhibit significant fragility across basic tasks, frequently misdescribing reversed videos, overlooking visual evidence, or accepting incorrect claims, whereas human participants perform these tasks with ease. These findings expose fundamental deficiencies in existing models’ video understanding capabilities.
This work addresses the high computational and maintenance costs associated with full retraining of multilingual large language models when adding or updating languages. To overcome this limitation, the authors propose a language-specific model merging strategy that efficiently integrates new languages or updated data by fusing dedicated language submodels, eliminating the need for full-scale retraining. As the first systematic study to evaluate multilingual model merging from an efficiency perspective, the approach maintains competitive model performance while reducing initial training time by up to 50% and cutting the cost of post-update re-merging for individual languages by over 60%. The effectiveness of the method is validated on both public and industrial datasets.
To address the computational and memory bottlenecks of standard Transformers—stemming from the quadratic complexity of self-attention—in real-time long-dialogue understanding, this work systematically evaluates efficient Transformer variants (e.g., Performer, Reformer) and lightweight CNN encoders. We empirically demonstrate, for the first time, that CNN-based architectures achieve both superior efficiency—2.6× faster training, 80% faster inference, and 72% lower memory consumption—and state-of-the-art generalization performance on both real-world customer service dialogues and the Long Range Arena (LRA) benchmark. This challenges the prevailing architectural dependency on Transformers for long-sequence modeling and establishes CNNs as a highly efficient, scalable alternative for real-time semantic understanding under resource constraints.
本文探讨了ACC^0和NC^1电路复杂性类,通过研究不同拓扑结构如多对数属、交叉数和厚度,证明了特定条件下这些属性对计算能力的影响。
It remains unclear whether current video-language models genuinely comprehend the temporal dynamics, motion, and semantics of videos. To address this, this work proposes REVEAL—the first systematic diagnostic benchmark—comprising five controlled stress tests: reversed playback, false statement injection, spatiotemporal occlusion, camera motion simulation, and language shortcut interference. Coupled with an automated pipeline for generating controllable perturbations, REVEAL comprehensively evaluates model robustness in fundamental perception and reasoning. Experiments reveal that both leading open-source and proprietary models exhibit significant fragility across basic tasks, frequently misdescribing reversed videos, overlooking visual evidence, or accepting incorrect claims, whereas human participants perform these tasks with ease. These findings expose fundamental deficiencies in existing models’ video understanding capabilities.
This work addresses the high computational and maintenance costs associated with full retraining of multilingual large language models when adding or updating languages. To overcome this limitation, the authors propose a language-specific model merging strategy that efficiently integrates new languages or updated data by fusing dedicated language submodels, eliminating the need for full-scale retraining. As the first systematic study to evaluate multilingual model merging from an efficiency perspective, the approach maintains competitive model performance while reducing initial training time by up to 50% and cutting the cost of post-update re-merging for individual languages by over 60%. The effectiveness of the method is validated on both public and industrial datasets.
To address the computational and memory bottlenecks of standard Transformers—stemming from the quadratic complexity of self-attention—in real-time long-dialogue understanding, this work systematically evaluates efficient Transformer variants (e.g., Performer, Reformer) and lightweight CNN encoders. We empirically demonstrate, for the first time, that CNN-based architectures achieve both superior efficiency—2.6× faster training, 80% faster inference, and 72% lower memory consumption—and state-of-the-art generalization performance on both real-world customer service dialogues and the Long Range Arena (LRA) benchmark. This challenges the prevailing architectural dependency on Transformers for long-sequence modeling and establishes CNNs as a highly efficient, scalable alternative for real-time semantic understanding under resource constraints.