SALUTE: Benchmarking and Adapting LLMs for the Defense Domain
本文提出SALUTE框架,通过整合专门语料库和多阶段训练方法,解决国防领域语言模型适应性问题,提高其在该领域的性能。
本文提出SALUTE框架,通过整合专门语料库和多阶段训练方法,解决国防领域语言模型适应性问题,提高其在该领域的性能。
Existing optical flow methods suffer significant performance degradation under realistic image degradations such as motion blur, noise, and compression artifacts. This work proposes a hybrid architecture that integrates intermediate features from diffusion models with convolutional representations, leveraging—for the first time—the inherent degradation-aware capabilities of diffusion models for optical flow estimation. By introducing a cross-frame spatiotemporal attention mechanism, the method enables zero-shot correspondence modeling without requiring retraining or fine-tuning under diverse degradations. The resulting framework establishes a new paradigm for degradation-robust optical flow estimation, consistently outperforming current state-of-the-art approaches across multiple benchmarks under various severe degradation conditions.
To address three key challenges in audio-visual multimodal RAG—narrow modality coverage of knowledge graphs, weak multi-hop connectivity, and imprecise retrieval—this paper proposes a query-aligned Multi-hop Multimodal Knowledge Graph (M³KG) construction and retrieval framework. We introduce a novel lightweight multi-agent construction method that significantly expands the modality granularity and cross-modal path depth of multimodal knowledge graphs (MMKGs). Furthermore, we design the GRASP mechanism—comprising query-driven entity anchoring, supportiveness assessment, and redundant context pruning—to enhance retrieval precision. By integrating modality-aware retrieval, query grounding, relevance scoring, and embedding alignment, our approach improves fact consistency and cross-modal localization accuracy for multimodal large language models (MLLMs) in multi-hop reasoning. Extensive evaluation across multiple multimodal benchmarks demonstrates substantial gains in answer faithfulness and reasoning depth.
Traditional recurrent video super-resolution (VSR) methods suffer from gradient vanishing and poor parallelism, while causal Mamba-based models are inherently limited in modeling fine-grained spatial dependencies. To address these issues, we propose an efficient hybrid spatiotemporal modeling architecture. Our approach features: (1) a Gather-Scatter Mamba mechanism that aligns neighboring frame features to the central frame within a temporal window before aggregation and scattering, mitigating occlusion artifacts and enhancing feature redistribution; and (2) integration of shifted-window self-attention to explicitly capture local spatial dependencies, compensating for Mamba’s structural constraints. The architecture retains linear time complexity while enabling precise spatiotemporal feature propagation. Extensive experiments demonstrate state-of-the-art performance on multiple VSR benchmarks, along with significantly accelerated inference—effectively balancing accuracy and efficiency.
本文提出SALUTE框架,通过整合专门语料库和多阶段训练方法,解决国防领域语言模型适应性问题,提高其在该领域的性能。
Existing optical flow methods suffer significant performance degradation under realistic image degradations such as motion blur, noise, and compression artifacts. This work proposes a hybrid architecture that integrates intermediate features from diffusion models with convolutional representations, leveraging—for the first time—the inherent degradation-aware capabilities of diffusion models for optical flow estimation. By introducing a cross-frame spatiotemporal attention mechanism, the method enables zero-shot correspondence modeling without requiring retraining or fine-tuning under diverse degradations. The resulting framework establishes a new paradigm for degradation-robust optical flow estimation, consistently outperforming current state-of-the-art approaches across multiple benchmarks under various severe degradation conditions.
To address three key challenges in audio-visual multimodal RAG—narrow modality coverage of knowledge graphs, weak multi-hop connectivity, and imprecise retrieval—this paper proposes a query-aligned Multi-hop Multimodal Knowledge Graph (M³KG) construction and retrieval framework. We introduce a novel lightweight multi-agent construction method that significantly expands the modality granularity and cross-modal path depth of multimodal knowledge graphs (MMKGs). Furthermore, we design the GRASP mechanism—comprising query-driven entity anchoring, supportiveness assessment, and redundant context pruning—to enhance retrieval precision. By integrating modality-aware retrieval, query grounding, relevance scoring, and embedding alignment, our approach improves fact consistency and cross-modal localization accuracy for multimodal large language models (MLLMs) in multi-hop reasoning. Extensive evaluation across multiple multimodal benchmarks demonstrates substantial gains in answer faithfulness and reasoning depth.
Traditional recurrent video super-resolution (VSR) methods suffer from gradient vanishing and poor parallelism, while causal Mamba-based models are inherently limited in modeling fine-grained spatial dependencies. To address these issues, we propose an efficient hybrid spatiotemporal modeling architecture. Our approach features: (1) a Gather-Scatter Mamba mechanism that aligns neighboring frame features to the central frame within a temporal window before aggregation and scattering, mitigating occlusion artifacts and enhancing feature redistribution; and (2) integration of shifted-window self-attention to explicitly capture local spatial dependencies, compensating for Mamba’s structural constraints. The architecture retains linear time complexity while enabling precise spatiotemporal feature propagation. Extensive experiments demonstrate state-of-the-art performance on multiple VSR benchmarks, along with significantly accelerated inference—effectively balancing accuracy and efficiency.