SpatialTrust: A Benchmark for Environmental Risk Recognition in Secure Authentication

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SpatialTrust基准,评估多模态大语言模型在安全认证中识别环境风险的能力,并通过SpatialTrustGuard提升模型性能。
📝 Abstract
Visual environmental risk recognition plays an important role in secure authentication, where a user's surroundings may reveal sensitive information or introduce potential security risks. However, existing evaluations of multimodal large language models (MLLMs) rarely examine whether models can reliably recognize, localize, and explain such risks in spatially grounded authentication scenarios. We present SpatialTrust, a question-answering benchmark for evaluating environmental risk recognition in secure authentication. SpatialTrust assesses five complementary abilities: sensitive factor detection, direct factor identification, indirect factor identification, direct factor explanation, and indirect factor explanation. We evaluate both proprietary and open-source MLLMs and find that current models show limited performance, especially in understanding and explaining indirect risks, indicating that spatial risk awareness remains a challenging capability for MLLMs. In addition, we introduce SpatialTrustGuard, a structured QA-and-audit pipeline that improves Qwen3-VL-30B-A3B-Instruct from 36.78% to 41.12% overall. Our findings highlight the need for dedicated benchmarks and structured inference methods to improve the trustworthiness of MLLMs in secure authentication.
Problem

Research questions and friction points this paper is trying to address.

environmental risk recognition
secure authentication
multimodal large language models
indirect risks
spatially grounded
Innovation

Methods, ideas, or system contributions that make the work stand out.

Environmental Risk Recognition
Question-Answering Benchmark
SpatialTrust
Secure Authentication
SpatialTrustGuard
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junbin Lu
University of Washington, United States
Hsiang-Wei Huang
Hsiang-Wei Huang
University of Washington
Computer VisionDeep Learning3D Vision
S
Saesha Wadhwa
University of Washington, United States
Y
Yu Ting Hsu
University of Washington, United States
Jenq-Neng Hwang
Jenq-Neng Hwang
University of Washington
Signal ProcessingPattern RecognitionWirelessVideo Analysis