Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions
本文探索使用视觉语言模型加速分类模型验证和确认过程中的人工检查,提出了一种基于VLM的错误切片检测方法来识别系统性错误。
本文探索使用视觉语言模型加速分类模型验证和确认过程中的人工检查,提出了一种基于VLM的错误切片检测方法来识别系统性错误。
研究通过使用生成式AI图像编辑方法,为训练数据添加伪装,以提高物体检测器在域迁移情况下的鲁棒性,特别是在军事车辆检测中。
研究通过使用生成模型(如GANs、ControlNet扩散模型等)将RGB图像转换为IR图像,以增强红外车辆检测在未知UAV领域的性能,有效缓解了IR数据稀缺问题。
This study addresses the limitations of restricted search spaces and non-editable outputs in traditional AutoML by proposing LACE, a novel code-based tabular AutoML framework. Leveraging large language models as mutation operators within an evolutionary algorithm, LACE optimizes populations of scikit-learn-compatible pipelines at the code level. Unlike conventional approaches constrained by predefined structures, this paradigm generates transparent, editable code rather than opaque black-box models, thereby enhancing reusability and interpretability. Extensive evaluations across 68 OpenML tasks demonstrate that LACE significantly outperforms auto-sklearn and achieves performance comparable to AutoGluon. By unifying comprehensive search coverage with high transparency, LACE effectively overcomes the rigidity of existing AutoML systems while delivering human-readable solutions that facilitate downstream refinement and domain-specific adaptation.
This study addresses the challenge of low initial coordination efficiency in urban search and rescue scenarios, where robots struggle to rapidly adapt to dynamic human-robot collaborative environments. To overcome this limitation, the work introduces a reusable episodic memory mechanism that encodes historical collaboration patterns into a knowledge graph. By integrating graph representation learning, the system automatically retrieves and initializes optimal behavioral policies prior to new tasks, enabling effective cross-task knowledge transfer. Experimental evaluations on the MATRX simulation platform demonstrate significant improvements in collaborative performance: rescue success rates increase from 25.7% to 41.3%, and average task completion time is reduced by 283 seconds, with particularly pronounced gains during the early phases of missions.
本文探索使用视觉语言模型加速分类模型验证和确认过程中的人工检查,提出了一种基于VLM的错误切片检测方法来识别系统性错误。
研究通过使用生成式AI图像编辑方法,为训练数据添加伪装,以提高物体检测器在域迁移情况下的鲁棒性,特别是在军事车辆检测中。
研究通过使用生成模型(如GANs、ControlNet扩散模型等)将RGB图像转换为IR图像,以增强红外车辆检测在未知UAV领域的性能,有效缓解了IR数据稀缺问题。
This study addresses the limitations of restricted search spaces and non-editable outputs in traditional AutoML by proposing LACE, a novel code-based tabular AutoML framework. Leveraging large language models as mutation operators within an evolutionary algorithm, LACE optimizes populations of scikit-learn-compatible pipelines at the code level. Unlike conventional approaches constrained by predefined structures, this paradigm generates transparent, editable code rather than opaque black-box models, thereby enhancing reusability and interpretability. Extensive evaluations across 68 OpenML tasks demonstrate that LACE significantly outperforms auto-sklearn and achieves performance comparable to AutoGluon. By unifying comprehensive search coverage with high transparency, LACE effectively overcomes the rigidity of existing AutoML systems while delivering human-readable solutions that facilitate downstream refinement and domain-specific adaptation.
This study addresses the challenge of low initial coordination efficiency in urban search and rescue scenarios, where robots struggle to rapidly adapt to dynamic human-robot collaborative environments. To overcome this limitation, the work introduces a reusable episodic memory mechanism that encodes historical collaboration patterns into a knowledge graph. By integrating graph representation learning, the system automatically retrieves and initializes optimal behavioral policies prior to new tasks, enabling effective cross-task knowledge transfer. Experimental evaluations on the MATRX simulation platform demonstrate significant improvements in collaborative performance: rescue success rates increase from 25.7% to 41.3%, and average task completion time is reduced by 283 seconds, with particularly pronounced gains during the early phases of missions.