VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出VeriScene,通过整合法医照片和证词重建犯罪现场,并利用世界模型验证动态合理性,解决了原始证据直接输入模型导致的问题。
📝 Abstract
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance with the laws of physics, thus opening a compelling application: fusing multimodal legal evidence to re-create a crime scene and re-enact how an offence could have been committed. However, feeding the raw, unorganized evidence into a world model fails in forensic use: it silently drops evidence, glosses over contradictory testimony, and produces motion that violates the evidentiary record. This paper presents VeriScene, an agent that orchestrates the world model: it reconstructs crime scenes from forensic photographs and witness statements of varying reliability, keeping every claim traceable to evidence and every motion physically plausible. VeriScene iteratively fuses the evidence into a cited narrative under an auditing loop, verifies the hypothesized dynamics via probe rollouts in the world model with corrective constraint injection, and renders the offence as a re-enactment video from a fused keyframe. On a benchmark of 25 crime scenarios across 7 physically-driven case types (139 forensic-style photographs and 65 statements with planted unreliability), VeriScene attains 0.9014 evidence coverage and 0.7217 factual consistency (0-1 scale) on the 20 test scenes, outperforming an end-to-end multimodal-LLM baseline by 20.35% in factual consistency and 34.88% in temporal coherence, while generalizing across four LLM orchestration backends at USD 1.82 per scene.
Problem

Research questions and friction points this paper is trying to address.

Crime Scene Reconstruction
Multimodal Evidence Fusion
Forensic Analysis
Physical Plausibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Model Agent
Evidence Fusion
Physical Plausibility
Crime Scene Reconstruction
Multimodal Inputs
🔎 Similar Papers
No similar papers found.