SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes
为解决现有3D场景理解基准的局限性,本文通过构建包含966个光真实3D场景的SceneBench并定义三个评估任务来提升视觉-语言模型在3D空间推理方面的能力。
为解决现有3D场景理解基准的局限性,本文通过构建包含966个光真实3D场景的SceneBench并定义三个评估任务来提升视觉-语言模型在3D空间推理方面的能力。
研究利用大规模预训练策略改进基于深度学习的扩散加权成像几何失真校正,通过自监督和生成式预训练模型提升校正效果。
Existing 3D scene generation methods struggle to reliably satisfy task-critical functional constraints such as navigability and reachability, limiting the practical utility of synthetic data. This work proposes an iterative agent-based reinforcement learning framework that first enhances physical plausibility and layout quality through pretraining with generic rewards, then leverages a large language model (LLM) to generate executable, task-specific reward programs. These LLM-generated rewards are integrated into a feedback-driven reinforcement learning loop for iterative refinement. By uniquely combining LLM-synthesized reward functions with iterative reinforcement learning, the approach significantly improves adherence to functional constraints while preserving scene diversity, thereby enhancing downstream task performance.
Existing AI literacy frameworks struggle to address the systemic disruption generative AI poses to high-skill white-collar work and lack a robust framework for human–AI collaboration. This study proposes the “AI Pyramid” model, introducing the novel concept of “AI-native competencies” and reclassifying human capabilities into three tiers: AI-native, AI-foundational, and AI-advanced. Moving beyond traditional occupational hierarchies, this model reconceptualizes the societal distribution of skills at a systemic level. Through conceptual modeling, competency ontology design, scenario-based problem-based learning (PBL), and competency-oriented assessment, the study establishes a scalable pathway for cultivating an AI-ready workforce. The framework offers actionable strategies for educational institutions, enterprises, and governments to enhance societal productivity and resilience while mitigating technology-driven inequality.
In resource-limited settings, early screening for ophthalmic and otologic diseases is hindered by critical shortages of specialists, inadequate diagnostic equipment, and the difficulty of transitioning paper-based workflows to AI-ready digital systems. Method: This study proposes an end-to-end, iterative co-design methodology that tightly integrates AI model development—including transfer learning and automated image quality assessment—with digital health workflow reengineering. Field-based prototyping, shadow deployment, and continuous feedback loops were employed to rigorously evaluate system usability and operational feasibility. Contribution/Results: We introduce a novel “AI–Workflow–Human Factors” triadic framework for localized adaptation, distilling reusable deployment insights and evidence-based strategies for overcoming key implementation barriers—such as workflow misalignment, clinician trust deficits, and infrastructure constraints. The resulting dual-dimensional (technical and managerial) guidance framework enables sustainable, scalable deployment of AI-assisted screening programs in low-resource environments.
为解决现有3D场景理解基准的局限性,本文通过构建包含966个光真实3D场景的SceneBench并定义三个评估任务来提升视觉-语言模型在3D空间推理方面的能力。
研究利用大规模预训练策略改进基于深度学习的扩散加权成像几何失真校正,通过自监督和生成式预训练模型提升校正效果。
Existing 3D scene generation methods struggle to reliably satisfy task-critical functional constraints such as navigability and reachability, limiting the practical utility of synthetic data. This work proposes an iterative agent-based reinforcement learning framework that first enhances physical plausibility and layout quality through pretraining with generic rewards, then leverages a large language model (LLM) to generate executable, task-specific reward programs. These LLM-generated rewards are integrated into a feedback-driven reinforcement learning loop for iterative refinement. By uniquely combining LLM-synthesized reward functions with iterative reinforcement learning, the approach significantly improves adherence to functional constraints while preserving scene diversity, thereby enhancing downstream task performance.
Existing AI literacy frameworks struggle to address the systemic disruption generative AI poses to high-skill white-collar work and lack a robust framework for human–AI collaboration. This study proposes the “AI Pyramid” model, introducing the novel concept of “AI-native competencies” and reclassifying human capabilities into three tiers: AI-native, AI-foundational, and AI-advanced. Moving beyond traditional occupational hierarchies, this model reconceptualizes the societal distribution of skills at a systemic level. Through conceptual modeling, competency ontology design, scenario-based problem-based learning (PBL), and competency-oriented assessment, the study establishes a scalable pathway for cultivating an AI-ready workforce. The framework offers actionable strategies for educational institutions, enterprises, and governments to enhance societal productivity and resilience while mitigating technology-driven inequality.
In resource-limited settings, early screening for ophthalmic and otologic diseases is hindered by critical shortages of specialists, inadequate diagnostic equipment, and the difficulty of transitioning paper-based workflows to AI-ready digital systems. Method: This study proposes an end-to-end, iterative co-design methodology that tightly integrates AI model development—including transfer learning and automated image quality assessment—with digital health workflow reengineering. Field-based prototyping, shadow deployment, and continuous feedback loops were employed to rigorously evaluate system usability and operational feasibility. Contribution/Results: We introduce a novel “AI–Workflow–Human Factors” triadic framework for localized adaptation, distilling reusable deployment insights and evidence-based strategies for overcoming key implementation barriers—such as workflow misalignment, clinician trust deficits, and infrastructure constraints. The resulting dual-dimensional (technical and managerial) guidance framework enables sustainable, scalable deployment of AI-assisted screening programs in low-resource environments.