Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态临床AI中输入对齐差及缺乏领域特定可解释表示的问题,本文提出基于运动学知识图谱和模板文本的ScoliDetect框架,通过双向交叉注意与潜在瓶颈聚合方法提高了青少年特发性脊柱侧弯筛查的准确性和可解释性。
📝 Abstract
Multimodal clinical AI is limited by weakly aligned inputs and the absence of domain-specific interpretable representations, particularly when learning from dense video stream, structured time-series, and template-based kinematic text. Here we present ScoliDetect, an explainable framework for adolescent idiopathic scoliosis screening from monocular gait video, built around a kinematic knowledge map (KKM) and complementary template-based kinematic text derived from per-sequence pose statics. KKM is a fixed-index structured representation that encodes gait features across absolute motion, self-skeleton configuration and joint-joint signal correlation, providing anchor-referenced multimodal fusion and factor-level interpretation. We integrate video, KKM, and template-based kinematic text through bidirectional cross-attention with latent-bottleneck aggregation. In a multicenter cohort (n = 1,858 after exclusions), prespecified supervised ablations on an external screening cohort show that KKM-mediated multimodal fusion outperforms unimodal models and late concatenation. Under a staged training protocol, trimodal contrastive pretraining is applied after architecture selection as representation initialization, improving external ROC-AUC from 0.961 to 0.972. Furthermore, the structured nature of the KKM provides inherent, factor-level attributions mapped directly to specific kinematic phases and skeletal indices, offering verifiable interpretability. The results demonstrate that embedding explicit structural topologies into latent spaces significantly enhances both the generalization and explainability of multimodal pattern analysis systems.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Clinical AI
Weakly Aligned Inputs
Interpretable Representations
Gait Analysis
Scoliosis Screening
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kinematic Knowledge Map
Multimodal Fusion
Explainability
Bidirectional Cross-Attention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Chen Dong
Chen Dong
Beijing University of Posts and Telecommunications
wireless communicationssemanticapplied math
H
He Zonglin
Department of Orthopaedics & Traumatology, University of Hong Kong, Hong Kong, China
C
Cheung Kenneth M. C.
Department of Orthopaedics & Traumatology, University of Hong Kong, Hong Kong, China; and the University of Hong Kong - Shenzhen Hospital, Shenzhen, China