Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention
该研究通过引入Lévy Attention方法,解决了不规则采样时间序列模型在任意连续时间戳查询时缺乏预测可信度的问题。
该研究通过引入Lévy Attention方法,解决了不规则采样时间序列模型在任意连续时间戳查询时缺乏预测可信度的问题。
This study addresses the problem of predicting both the timing and location of link formation in complex networks. It proposes a closed-form, non-Markovian model that integrates latent hyperbolic geometry with long-range memory of historical interactions, thereby unifying geometric structure and memory effects within a single framework for the first time. The resulting approach features few parameters and strong interpretability, offering a principled method for temporal link prediction. By modeling network dynamics through a non-Markovian process and deriving probabilistic predictions, the model achieves excellent agreement with empirical connection probabilities across multiple large-scale real-world networks. These results reveal that network evolution is fundamentally governed by the interplay between geometric constraints and memory-driven mechanisms.
This work proposes a lightweight multimodal approach to accurately predict user gaze direction in virtual reality scenarios where eye-tracking hardware is unavailable or restricted by privacy constraints—a critical capability for techniques such as foveated rendering. The method uniquely integrates head-mounted display (HMD) motion signals with visual saliency cues from video frames by leveraging UniSal for visual feature extraction and combining TSMixer with LSTM to construct a temporal prediction module. Experiments on the EHTask dataset and commercial VR devices demonstrate that the proposed approach significantly outperforms baseline methods such as Center-of-HMD and Mean Gaze, achieving high prediction accuracy, low latency, and practical deployability without requiring eye-tracking data, thereby enhancing the naturalness and efficiency of VR interactions.
Urban sidewalks are frequently obstructed by hazards that compromise pedestrian safety, yet real-time detection is hindered by the absence of high-quality, multi-class egocentric visual datasets. To address this gap, we introduce the first large-scale egocentric video dataset specifically designed for sidewalk obstacle detection—comprising 340 real-world smartphone-recorded videos spanning 29 common obstacle categories. We systematically define and publicly release a high-fidelity, fine-grained annotation benchmark, the first of its kind, thereby filling a critical void in open pedestrian safety resources. Leveraging this dataset, we conduct a comprehensive evaluation of state-of-the-art object detectors—including YOLOv8 and Mask R-CNN—establishing fully reproducible baselines. Our best-performing model achieves a mean average precision (mAP@0.5) of 68.3%. This work provides both an essential data foundation and an authoritative performance benchmark for developing robust pedestrian safety warning systems.
This study presents the first systematic investigation of left-wing extremist content dissemination on the decentralized social media platform Lemmy (specifically the Lemmygrad.ml instance), focusing on subcommunities such as r/GenZedong and r/GenZhou that migrated from mainstream platforms. Employing Transformer-based topic modeling, time-series analysis, and multi-dimensional hate speech detection, we quantitatively assess user engagement, content toxicity, and thematic evolution. Results indicate a significant post-migration increase in user activity and extremist content—including authoritarian rhetoric, pro-Russian invasion stances, and antisemitic discourse. Our contributions are threefold: (1) addressing a critical empirical gap in research on left-wing extremism within decentralized platforms; (2) uncovering dynamic cross-platform migration patterns of political extremism; and (3) advocating balanced scholarly and regulatory attention to extremism across the full political spectrum—thereby advancing platform governance frameworks and digital public sphere scholarship.
该研究通过引入Lévy Attention方法,解决了不规则采样时间序列模型在任意连续时间戳查询时缺乏预测可信度的问题。
This study addresses the problem of predicting both the timing and location of link formation in complex networks. It proposes a closed-form, non-Markovian model that integrates latent hyperbolic geometry with long-range memory of historical interactions, thereby unifying geometric structure and memory effects within a single framework for the first time. The resulting approach features few parameters and strong interpretability, offering a principled method for temporal link prediction. By modeling network dynamics through a non-Markovian process and deriving probabilistic predictions, the model achieves excellent agreement with empirical connection probabilities across multiple large-scale real-world networks. These results reveal that network evolution is fundamentally governed by the interplay between geometric constraints and memory-driven mechanisms.
This work proposes a lightweight multimodal approach to accurately predict user gaze direction in virtual reality scenarios where eye-tracking hardware is unavailable or restricted by privacy constraints—a critical capability for techniques such as foveated rendering. The method uniquely integrates head-mounted display (HMD) motion signals with visual saliency cues from video frames by leveraging UniSal for visual feature extraction and combining TSMixer with LSTM to construct a temporal prediction module. Experiments on the EHTask dataset and commercial VR devices demonstrate that the proposed approach significantly outperforms baseline methods such as Center-of-HMD and Mean Gaze, achieving high prediction accuracy, low latency, and practical deployability without requiring eye-tracking data, thereby enhancing the naturalness and efficiency of VR interactions.
Urban sidewalks are frequently obstructed by hazards that compromise pedestrian safety, yet real-time detection is hindered by the absence of high-quality, multi-class egocentric visual datasets. To address this gap, we introduce the first large-scale egocentric video dataset specifically designed for sidewalk obstacle detection—comprising 340 real-world smartphone-recorded videos spanning 29 common obstacle categories. We systematically define and publicly release a high-fidelity, fine-grained annotation benchmark, the first of its kind, thereby filling a critical void in open pedestrian safety resources. Leveraging this dataset, we conduct a comprehensive evaluation of state-of-the-art object detectors—including YOLOv8 and Mask R-CNN—establishing fully reproducible baselines. Our best-performing model achieves a mean average precision (mAP@0.5) of 68.3%. This work provides both an essential data foundation and an authoritative performance benchmark for developing robust pedestrian safety warning systems.
This study presents the first systematic investigation of left-wing extremist content dissemination on the decentralized social media platform Lemmy (specifically the Lemmygrad.ml instance), focusing on subcommunities such as r/GenZedong and r/GenZhou that migrated from mainstream platforms. Employing Transformer-based topic modeling, time-series analysis, and multi-dimensional hate speech detection, we quantitatively assess user engagement, content toxicity, and thematic evolution. Results indicate a significant post-migration increase in user activity and extremist content—including authoritarian rhetoric, pro-Russian invasion stances, and antisemitic discourse. Our contributions are threefold: (1) addressing a critical empirical gap in research on left-wing extremism within decentralized platforms; (2) uncovering dynamic cross-platform migration patterns of political extremism; and (3) advocating balanced scholarly and regulatory attention to extremism across the full political spectrum—thereby advancing platform governance frameworks and digital public sphere scholarship.