GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering
本文提出GRACE框架,通过适配器组合和证据感知校准解决教育视觉问答中的问题,提高了模型在多模态情境下的准确率。
本文提出GRACE框架,通过适配器组合和证据感知校准解决教育视觉问答中的问题,提高了模型在多模态情境下的准确率。
This work addresses the computational and memory bottlenecks inherent in deploying real-time, high-density electromyography (HD-EMG)–based gesture recognition on resource-constrained embedded devices. The authors propose an edge-based neuromuscular interface system that, for the first time, enables end-to-end real-time gesture recognition using 192-channel HD-EMG signals directly on microcontrollers (ESP32 and Sony Spresense). The system integrates a custom HD-EMG StreamBridge wireless interface, a lightweight EdgeDL inference framework, and a one-dimensional convolutional neural network, with co-optimized DMA and SPI burst communication to establish an efficient data streaming and inference pipeline. Evaluated on seven distinct gestures, the system achieves a classification accuracy of 90% with an average end-to-end latency of only 83 milliseconds.
This study addresses the substantial overestimation of the “unsolvability ceiling” in multi-LLM routing observed in prior work, which stems from evaluation artifacts such as judge bias, generation truncation, and format mismatches, thereby misleading router design. The authors propose the first systematic decomposition framework that leverages dual-judge verification, exact-match anchoring, and randomized shuffling controls. Applying this framework to 206,000 query–model pairs across six benchmarks using Gemma 4 and Llama 3.1 series models, they rigorously quantify how these artifacts inflate perceived unsolvability and distort training signals for routers. Their analysis markedly reduces the measured unsolvable fractions across tasks and reveals that standard routers often degenerate into majority-class predictors, incurring opportunity cost losses of 13–17 percentage points.
This work addresses the cognitive overload experienced by visually impaired users due to continuous, undifferentiated feedback in existing navigation aids, which often impedes effective communication of dynamic environmental information. To mitigate this, we propose a motion-aware adaptive video-to-audio conversion framework that intelligently alternates between spoken descriptions and non-verbal auditory cues based on scene dynamics. The system incorporates prompt caching and category-based rate limiting to minimize auditory interference while maintaining low latency. It integrates a lightweight AI classifier, a decoder-only Transformer-based vision-language model enhanced with mixture-of-experts and cross-modal attention mechanisms, neural text-to-speech synthesis, and naturalistic sound generation. Real-world navigation experiments demonstrate that, compared to using a white cane alone, our system significantly improves users’ environmental awareness, subjective sense of safety, and confidence in navigation.
This study addresses the challenges of data privacy leakage, security risks, and regulatory compliance inherent in traditional centralized machine learning within cloud-edge environments. To overcome these issues, the authors propose a novel cloud-edge collaborative architecture that integrates federated learning with blockchain technology. They introduce the first four-dimensional taxonomy—encompassing coordination mechanisms, consensus algorithms, data storage, and trust models—to systematically evaluate existing blockchain-enabled federated learning (BCFL) frameworks. The work provides an in-depth comparative analysis of two representative approaches, MORFLB and FBCI-SHS, elucidating their respective strengths and limitations. Building on this assessment, the paper identifies key open challenges and outlines a forward-looking research agenda centered on adaptability, resilience, and standardization.
本文提出GRACE框架,通过适配器组合和证据感知校准解决教育视觉问答中的问题,提高了模型在多模态情境下的准确率。
This work addresses the computational and memory bottlenecks inherent in deploying real-time, high-density electromyography (HD-EMG)–based gesture recognition on resource-constrained embedded devices. The authors propose an edge-based neuromuscular interface system that, for the first time, enables end-to-end real-time gesture recognition using 192-channel HD-EMG signals directly on microcontrollers (ESP32 and Sony Spresense). The system integrates a custom HD-EMG StreamBridge wireless interface, a lightweight EdgeDL inference framework, and a one-dimensional convolutional neural network, with co-optimized DMA and SPI burst communication to establish an efficient data streaming and inference pipeline. Evaluated on seven distinct gestures, the system achieves a classification accuracy of 90% with an average end-to-end latency of only 83 milliseconds.
This study addresses the substantial overestimation of the “unsolvability ceiling” in multi-LLM routing observed in prior work, which stems from evaluation artifacts such as judge bias, generation truncation, and format mismatches, thereby misleading router design. The authors propose the first systematic decomposition framework that leverages dual-judge verification, exact-match anchoring, and randomized shuffling controls. Applying this framework to 206,000 query–model pairs across six benchmarks using Gemma 4 and Llama 3.1 series models, they rigorously quantify how these artifacts inflate perceived unsolvability and distort training signals for routers. Their analysis markedly reduces the measured unsolvable fractions across tasks and reveals that standard routers often degenerate into majority-class predictors, incurring opportunity cost losses of 13–17 percentage points.
This work addresses the cognitive overload experienced by visually impaired users due to continuous, undifferentiated feedback in existing navigation aids, which often impedes effective communication of dynamic environmental information. To mitigate this, we propose a motion-aware adaptive video-to-audio conversion framework that intelligently alternates between spoken descriptions and non-verbal auditory cues based on scene dynamics. The system incorporates prompt caching and category-based rate limiting to minimize auditory interference while maintaining low latency. It integrates a lightweight AI classifier, a decoder-only Transformer-based vision-language model enhanced with mixture-of-experts and cross-modal attention mechanisms, neural text-to-speech synthesis, and naturalistic sound generation. Real-world navigation experiments demonstrate that, compared to using a white cane alone, our system significantly improves users’ environmental awareness, subjective sense of safety, and confidence in navigation.
This study addresses the challenges of data privacy leakage, security risks, and regulatory compliance inherent in traditional centralized machine learning within cloud-edge environments. To overcome these issues, the authors propose a novel cloud-edge collaborative architecture that integrates federated learning with blockchain technology. They introduce the first four-dimensional taxonomy—encompassing coordination mechanisms, consensus algorithms, data storage, and trust models—to systematically evaluate existing blockchain-enabled federated learning (BCFL) frameworks. The work provides an in-depth comparative analysis of two representative approaches, MORFLB and FBCI-SHS, elucidating their respective strengths and limitations. Building on this assessment, the paper identifies key open challenges and outlines a forward-looking research agenda centered on adaptability, resilience, and standardization.