Global-Local Contextual Progressive Expansion Network for Martian Landslide Segmentation in Multimodal Remote Sensing Imagery
本文针对火星滑坡分割问题,提出了一种结合上下文渐进层扩展特征提取和基于Transformer上下文推理的U形网络TransCPLES,以提高分割精度和计算效率。
本文针对火星滑坡分割问题,提出了一种结合上下文渐进层扩展特征提取和基于Transformer上下文推理的U形网络TransCPLES,以提高分割精度和计算效率。
Traditional Online Public Access Catalogs (OPACs) rely on keyword-based and Boolean queries, which struggle to support efficient knowledge discovery in vast collections of scholarly literature and often lead to information overload. This work proposes an intelligent OPAC framework that integrates artificial intelligence and knowledge graphs, incorporating semantic embeddings and multi-source open scholarly data into the OPAC system for the first time. By leveraging knowledge graph construction, context-aware semantic search, dynamic topic navigation, and interactive visualizations, the framework transcends the limitations of linear querying. The approach significantly enhances retrieval efficiency and result relevance while effectively mitigating information overload, offering a viable pathway toward modernizing digital libraries and supporting next-generation scholarly workflows.
This study systematically investigates boundary vertex theory for strongly connected directed graphs under the sum metric, clarifying inclusion relations and structural properties among boundary, contour, eccentric, and peripheral vertices. It further addresses the previously unexplored problem of characterizing boundary-type vertex sets—including boundary, contour, and center sets—in coronal product graphs, along with their distance properties. Method: The approach integrates shortest-path analysis in directed graphs, modeling of the sum-distance function, coronal product construction, and vertex classification techniques. Contribution/Results: First, it establishes the first unified framework for boundary vertices in directed graphs under the sum metric. Second, it provides explicit, closed-form characterizations of all key boundary-type sets and the center in coronal product graphs. Third, it derives precise algebraic expressions linking these sets to the boundary and center structures of the factor graphs, thereby revealing fundamental structural dependencies governed by the underlying factor graph topology.
This study addresses the challenge of accurately inverting articulatory movements (tongue/lip) from acoustic signals in speech production. We propose a stacked BiLSTM-CNN architecture with fixed-weight initialization, trained on multi-speaker electromagnetic articulography (EMA) data. The model leverages bidirectional LSTMs to capture temporal dynamics and 1D CNNs to enhance local articulatory feature representation. Evaluation employs a comprehensive multi-paradigm framework—speaker-dependent (SD), speaker-independent (SI), cross-dataset (CD), and cross-corpus (CC). Our key contribution is the novel fixed-weight initialization strategy, which drastically mitigates overfitting within very few training epochs and substantially improves generalization across speakers and corpora. Experiments on multi-source EMA datasets demonstrate faster convergence, superior robustness, and consistently higher accuracy than adaptive-weight baselines. This work establishes a new, interpretable, and highly generalizable paradigm for modeling speech production mechanisms and enabling high-fidelity articulatory-to-acoustic synthesis.
This study addresses the acoustic-to-articulatory inversion (AAI) problem—mapping acoustic signals to articulatory motion trajectories. We systematically review data-driven AAI approaches from 2011 to 2021, covering speaker-dependent and speaker-independent modeling, multimodal articulatory corpora (EMA, EPG, rtMRI), and cross-task applications including automatic speech recognition (ASR), language learning, and speech rehabilitation. Methodologically, we propose a unified evaluation framework using correlation coefficient (CC), root-mean-square error (RMSE), and mean frame error (MFE), enabling the first quantitative performance comparison across state-of-the-art models. Our analysis identifies key bottlenecks in joint modeling of medical imaging and speech, clarifying translational pathways to clinical practice. Leveraging synchronized multi-source acoustic–articulatory data, we develop an interpretable trajectory feedback framework that significantly improves dynamic tongue visualization accuracy (CC ↑12.3%, RMSE ↓18.7%), thereby advancing computer-assisted language training and pathological speech intervention.
本文针对火星滑坡分割问题,提出了一种结合上下文渐进层扩展特征提取和基于Transformer上下文推理的U形网络TransCPLES,以提高分割精度和计算效率。
Traditional Online Public Access Catalogs (OPACs) rely on keyword-based and Boolean queries, which struggle to support efficient knowledge discovery in vast collections of scholarly literature and often lead to information overload. This work proposes an intelligent OPAC framework that integrates artificial intelligence and knowledge graphs, incorporating semantic embeddings and multi-source open scholarly data into the OPAC system for the first time. By leveraging knowledge graph construction, context-aware semantic search, dynamic topic navigation, and interactive visualizations, the framework transcends the limitations of linear querying. The approach significantly enhances retrieval efficiency and result relevance while effectively mitigating information overload, offering a viable pathway toward modernizing digital libraries and supporting next-generation scholarly workflows.
This study systematically investigates boundary vertex theory for strongly connected directed graphs under the sum metric, clarifying inclusion relations and structural properties among boundary, contour, eccentric, and peripheral vertices. It further addresses the previously unexplored problem of characterizing boundary-type vertex sets—including boundary, contour, and center sets—in coronal product graphs, along with their distance properties. Method: The approach integrates shortest-path analysis in directed graphs, modeling of the sum-distance function, coronal product construction, and vertex classification techniques. Contribution/Results: First, it establishes the first unified framework for boundary vertices in directed graphs under the sum metric. Second, it provides explicit, closed-form characterizations of all key boundary-type sets and the center in coronal product graphs. Third, it derives precise algebraic expressions linking these sets to the boundary and center structures of the factor graphs, thereby revealing fundamental structural dependencies governed by the underlying factor graph topology.
This study addresses the challenge of accurately inverting articulatory movements (tongue/lip) from acoustic signals in speech production. We propose a stacked BiLSTM-CNN architecture with fixed-weight initialization, trained on multi-speaker electromagnetic articulography (EMA) data. The model leverages bidirectional LSTMs to capture temporal dynamics and 1D CNNs to enhance local articulatory feature representation. Evaluation employs a comprehensive multi-paradigm framework—speaker-dependent (SD), speaker-independent (SI), cross-dataset (CD), and cross-corpus (CC). Our key contribution is the novel fixed-weight initialization strategy, which drastically mitigates overfitting within very few training epochs and substantially improves generalization across speakers and corpora. Experiments on multi-source EMA datasets demonstrate faster convergence, superior robustness, and consistently higher accuracy than adaptive-weight baselines. This work establishes a new, interpretable, and highly generalizable paradigm for modeling speech production mechanisms and enabling high-fidelity articulatory-to-acoustic synthesis.
This study addresses the acoustic-to-articulatory inversion (AAI) problem—mapping acoustic signals to articulatory motion trajectories. We systematically review data-driven AAI approaches from 2011 to 2021, covering speaker-dependent and speaker-independent modeling, multimodal articulatory corpora (EMA, EPG, rtMRI), and cross-task applications including automatic speech recognition (ASR), language learning, and speech rehabilitation. Methodologically, we propose a unified evaluation framework using correlation coefficient (CC), root-mean-square error (RMSE), and mean frame error (MFE), enabling the first quantitative performance comparison across state-of-the-art models. Our analysis identifies key bottlenecks in joint modeling of medical imaging and speech, clarifying translational pathways to clinical practice. Leveraging synchronized multi-source acoustic–articulatory data, we develop an interpretable trajectory feedback framework that significantly improves dynamic tongue visualization accuracy (CC ↑12.3%, RMSE ↓18.7%), thereby advancing computer-assisted language training and pathological speech intervention.