Generalization Analysis of Distributed Kernel-based Robust Gradient Descent Algorithms
研究了在鲁棒损失函数下分布式核基梯度下降算法的泛化性能,通过选择合适的参数σ优化学习率并提高统计鲁棒性。
研究了在鲁棒损失函数下分布式核基梯度下降算法的泛化性能,通过选择合适的参数σ优化学习率并提高统计鲁棒性。
为解决扩散模型难以对齐特定目标且训练成本高问题,提出SwiftExplorer方法,通过探索机制和质量效率仲裁机制提高生成多样性与质量。
This work addresses the scarcity of training-free, open-vocabulary methods for semantic segmentation in remote sensing imagery, a task typically hindered by the high cost of pixel-level annotations. The authors propose DinoSplat-OV, a novel framework that, for the first time, directly leverages DINOv3 for open-vocabulary segmentation of remote sensing images without requiring fine-tuning or additional pretraining. To tackle the challenges posed by the dense, multi-scale, and large-size nature of such imagery, the method integrates text-guided denoising Laplacian propagation, RGB-guided anisotropic feature aggregation, Gaussian lattice upsampling, and a global anchor sliding-window mechanism. Evaluated on UDD5, DOTA, and LoveDA benchmarks, DinoSplat-OV achieves performance on par with or superior to existing zero-shot approaches, thereby filling a critical gap in the application of DINO-based models to this domain.
This work addresses the "understanding-action gap" in large language models (LLMs) when applied to recommender systems by proposing a feedback-driven agent framework. The approach first infers task-oriented user intent and then discovers effective recommendation strategies based on incremental utility and outcome feedback—rather than linguistic plausibility. It innovatively decouples the modeling of intent and policy knowledge, and compresses both into a lightweight semantic ID generator via dual-space relational distillation, enabling efficient LLM-free online inference. Evaluated on public benchmarks, the method significantly outperforms existing baselines, and large-scale online A/B tests demonstrate a 4.506% increase in revenue and a 4.621% improvement in ADVV (Average Daily Value per Visitor).
This work addresses the limitation of current large language models in multimodal sentiment analysis, which struggle to capture deep emotional semantics arising from structural dynamics and contextual interactions, often relying only on superficial features. To overcome this, the authors propose SentiLLM, a novel framework that first converts continuous non-verbal signals into compact, semantically grounded text-like tokens through structure-aware semantic abstraction. It then introduces a dual-stream salience–context calibration mechanism that decouples focal and ambient streams, leveraging textual priors to guide the detection of sentiment shifts and achieve cross-modal semantic alignment. Finally, a lightweight, plug-and-play adapter module is integrated for efficient adaptation. Evaluated on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2 benchmarks, SentiLLM achieves state-of-the-art performance with significantly improved discriminative capability while requiring only a minimal number of trainable parameters.
研究了在鲁棒损失函数下分布式核基梯度下降算法的泛化性能,通过选择合适的参数σ优化学习率并提高统计鲁棒性。
为解决扩散模型难以对齐特定目标且训练成本高问题,提出SwiftExplorer方法,通过探索机制和质量效率仲裁机制提高生成多样性与质量。
This work addresses the scarcity of training-free, open-vocabulary methods for semantic segmentation in remote sensing imagery, a task typically hindered by the high cost of pixel-level annotations. The authors propose DinoSplat-OV, a novel framework that, for the first time, directly leverages DINOv3 for open-vocabulary segmentation of remote sensing images without requiring fine-tuning or additional pretraining. To tackle the challenges posed by the dense, multi-scale, and large-size nature of such imagery, the method integrates text-guided denoising Laplacian propagation, RGB-guided anisotropic feature aggregation, Gaussian lattice upsampling, and a global anchor sliding-window mechanism. Evaluated on UDD5, DOTA, and LoveDA benchmarks, DinoSplat-OV achieves performance on par with or superior to existing zero-shot approaches, thereby filling a critical gap in the application of DINO-based models to this domain.
This work addresses the "understanding-action gap" in large language models (LLMs) when applied to recommender systems by proposing a feedback-driven agent framework. The approach first infers task-oriented user intent and then discovers effective recommendation strategies based on incremental utility and outcome feedback—rather than linguistic plausibility. It innovatively decouples the modeling of intent and policy knowledge, and compresses both into a lightweight semantic ID generator via dual-space relational distillation, enabling efficient LLM-free online inference. Evaluated on public benchmarks, the method significantly outperforms existing baselines, and large-scale online A/B tests demonstrate a 4.506% increase in revenue and a 4.621% improvement in ADVV (Average Daily Value per Visitor).
This work addresses the limitation of current large language models in multimodal sentiment analysis, which struggle to capture deep emotional semantics arising from structural dynamics and contextual interactions, often relying only on superficial features. To overcome this, the authors propose SentiLLM, a novel framework that first converts continuous non-verbal signals into compact, semantically grounded text-like tokens through structure-aware semantic abstraction. It then introduces a dual-stream salience–context calibration mechanism that decouples focal and ambient streams, leveraging textual priors to guide the detection of sentiment shifts and achieve cross-modal semantic alignment. Finally, a lightweight, plug-and-play adapter module is integrated for efficient adaptation. Evaluated on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2 benchmarks, SentiLLM achieves state-of-the-art performance with significantly improved discriminative capability while requiring only a minimal number of trainable parameters.