SkillNet: Create, Evaluate, and Connect AI Skills
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
People with visual impairments face significant challenges in independent navigation and spatial orientation. Method: This study introduces a 3D-printed tactile map and icon system designed to support on-site orientation and mobility training. Employing an iterative, user-centered design process, we conducted field-based human factors evaluations—including tactile recognition and spatial cognition assessments—in real-world public environments, marking the first empirical validation of 3D tactile maps in authentic settings. Contribution/Results: Results demonstrate that realistic 3D icons are accurately recognized without legends, significantly enhancing mental map construction. The system improves users’ spatial cognition, autonomous navigation proficiency, and sense of belonging. Furthermore, the study distills empirically grounded guidelines and reusable design principles for inclusive design, offering both methodological frameworks and technical pathways to advance accessible built environments.
This study investigates whether current AI-based assistive device research aligns with the authentic needs of blind and low-vision (BLV) individuals. Method: We conducted a systematic literature review of 646 papers and in-depth interviews with 24 BLV users, integrating bibliometric analysis, user-need prioritization, and Spearman rank correlation testing. Contribution/Results: Our analysis reveals, for the first time, only a weak correlation between prevalent academic task formulations—such as object detection and image captioning—and actual user preferences. Instead, the top five most frequently cited needs center on real-time scene understanding and natural language–based conversational interaction. Users strongly prefer head-mounted, lightweight, and minimally intrusive devices. These findings challenge the dominant vision-centric paradigm in assistive AI research and provide empirical grounding for a user-centered design shift—emphasizing contextual awareness, multimodal interaction, and ergonomic form factors—thereby informing more effective, human-centered assistive technology development.
This study critically examines AI’s dual impact in healthcare: its transformative potential in genomics and public health, alongside profound ethical and institutional risks—including privacy breaches, algorithmic bias, physician deskilling, and imbalanced human–machine decision authority. Moving beyond technocentric paradigms, it introduces two foundational conceptual contributions: the reconfiguration of care as “datafied caregiving” and the normative calibration of “machine recommendation weight,” both grounded in philosophy of technology and bioethics. Employing an interdisciplinary analytical framework integrating medical ethics, philosophy of science, health policy, and big-data governance, the study uncovers structurally embedded risks overlooked in prevailing discourse. Its key contribution lies in reframing global regulatory and ethics review frameworks to center transparency, redistribution of epistemic and decisional authority, and preservation of clinical agency as core evaluative criteria.
Existing VideoQA datasets lack fine-grained modeling of professional sports actions, hindering effective reasoning for descriptive, temporal, causal, and counterfactual questions. To address this, we introduce Sports-QA—the first video question answering benchmark tailored to professional sports scenarios—covering multiple sports disciplines and four categories of complex reasoning tasks. Methodologically, we propose the Auto-Focus Transformer (AFT), which employs an attention-driven dynamic focusing mechanism to adaptively model multi-scale temporal information and integrates joint video–language representation learning. Extensive experiments demonstrate that AFT achieves state-of-the-art performance on Sports-QA, substantially outperforming general-purpose VideoQA models. This work constitutes the first systematic validation of an architecture explicitly designed for fine-grained sports action understanding and dynamic logical reasoning, establishing a new foundation for domain-specific VideoQA research.
本文提出REPVIS2设计空间,通过八个维度描述复制研究与参考研究的关系,以解决复制研究设计难以描述和比较的问题。
为解决复杂代理推理中结构化证据与语义密度的限制,提出Wiki基础模型WFM,通过新的Wiki图模式和优化的基础设施协议提高性能并加速训练。
该研究提出RiskWorld框架,通过融合空间风险场、时间演员上下文和视觉鸟瞰图特征来预测交通风险,并选择性地替换轨迹以实现安全的自动驾驶规划。
本文针对3D重建模型在特定条件下的失效问题,提出了GeoCond方法,通过预测几何结构的不确定性及精修门控来提高模型的可靠性。
研究了图的强制着色函数,通过分析其性质及导数意义,并证明了在特定区间内计算其值是#P-hard问题。