Multidimensional Task Learning: A Unified Tensor Framework for Computer Vision Tasks

📅 2026-02-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a multidimensional task learning (MTL) framework grounded in the generalized Einstein MLP (GE-MLP), which operates directly in tensor space via Einstein products, eliminating the need to flatten high-dimensional data and thereby preserving its intrinsic structure. In contrast to conventional computer vision approaches that rely on matrix-based modeling and inherently disrupt data geometry through vectorization, the proposed framework unifies diverse tasks—such as classification, segmentation, and detection—as distinct dimensional configurations of tensors. This formulation not only strictly subsumes the representational capacity of traditional matrix methods but also theoretically demonstrates that mainstream vision tasks are special cases of MTL. Consequently, the framework enables the natural construction of more complex spatiotemporal or cross-modal tasks within a coherent end-to-end learning paradigm.

Technology Category

Application Category

📝 Abstract
This paper introduces Multidimensional Task Learning (MTL), a unified mathematical framework based on Generalized Einstein MLPs (GE-MLPs) that operate directly on tensors via the Einstein product. We argue that current computer vision task formulations are inherently constrained by matrix-based thinking: standard architectures rely on matrix-valued weights and vectorvalued biases, requiring structural flattening that restricts the space of naturally expressible tasks. GE-MLPs lift this constraint by operating with tensor-valued parameters, enabling explicit control over which dimensions are preserved or contracted without information loss. Through rigorous mathematical derivations, we demonstrate that classification, segmentation, and detection are special cases of MTL, differing only in their dimensional configuration within a formally defined task space. We further prove that this task space is strictly larger than what matrix-based formulations can natively express, enabling principled task configurations such as spatiotemporal or cross modal predictions that require destructive flattening under conventional approaches. This work provides a mathematical foundation for understanding, comparing, and designing computer vision tasks through the lens of tensor algebra.
Problem

Research questions and friction points this paper is trying to address.

Multidimensional Task Learning
tensor-based representation
matrix-based limitation
computer vision tasks
structural flattening
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multidimensional Task Learning
Tensor Framework
Einstein Product
GE-MLP
Task Space
🔎 Similar Papers
No similar papers found.
A
Alaa El Ichi
Université du Littoral Côte d'Opale, LMPA, 50 rue F. Buisson, 62228 Calais-Cedex, France
K
Khalide Jbilou
Université du Littoral Côte d'Opale, LMPA, 50 rue F. Buisson, 62228 Calais-Cedex, France