Search-to-World: Evaluation of 3D World Delivery from User Request through Web Search
本文解决了从用户请求到3D世界生成的问题,通过引入Search-to-World任务及WorldSearcher工具,系统评估了从网络搜索内容转换成3D世界的全过程。
本文解决了从用户请求到3D世界生成的问题,通过引入Search-to-World任务及WorldSearcher工具,系统评估了从网络搜索内容转换成3D世界的全过程。
This work addresses the lack of a unified, objective, and human-aligned evaluation standard for scientific idea generation by large language models. To this end, the authors propose LigBench—an automated, fine-grained evaluation benchmark—and introduce the PAIR-IQ dataset to train pairwise idea judgment models, enabling consistent and interpretable assessment across diverse generation distributions. The framework pioneers a shift from subjective scoring to structured comparative learning, substantially improving alignment with expert judgments. Experimental results demonstrate that LigBench outperforms existing methods in evaluation stability and interpretability, with models trained on PAIR-IQ achieving superior performance in ranking accuracy and robustness.
本文解决了从用户请求到3D世界生成的问题,通过引入Search-to-World任务及WorldSearcher工具,系统评估了从网络搜索内容转换成3D世界的全过程。
This work addresses the lack of a unified, objective, and human-aligned evaluation standard for scientific idea generation by large language models. To this end, the authors propose LigBench—an automated, fine-grained evaluation benchmark—and introduce the PAIR-IQ dataset to train pairwise idea judgment models, enabling consistent and interpretable assessment across diverse generation distributions. The framework pioneers a shift from subjective scoring to structured comparative learning, substantially improving alignment with expert judgments. Experimental results demonstrate that LigBench outperforms existing methods in evaluation stability and interpretability, with models trained on PAIR-IQ achieving superior performance in ranking accuracy and robustness.