Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical vulnerability in multimodal large language model (MLLM) peer-review panels: their susceptibility to fabricated peer judgments due to an inability to discern the provenance of input assessments—a phenomenon termed “source-blind anchoring.” The work presents the first identification of this text-level attack surface and introduces a novel panel consensus mechanism based on independent blind cross-validation. By employing leave-one-out revalidation coupled with statistical significance testing, the proposed method rigorously evaluates the reliability of review inputs. Experimental results demonstrate that this approach effectively blocks 84.9% of forged attacks and reduces net harm by 97.5%, while preserving the constructive value of authentic peer judgments.
📝 Abstract
Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-level attack surface in vision-language model (VLM) panels. Quoting independent visual judgments creates large anchoring gaps (19--26 percentage points) under both self and peer framing. A matched-content, label-only control changes the broken rate by only $-0.17$pp (95\% CI $[-0.68,0.35]$), showing that the self/peer label itself does not explain the effect. Under our tested construction, deliberately generated, concise wrong quotes overturn originally-correct verdicts 1.5--2.7$\times$ more often than naturally occurring wrong peer statements, with bootstrap 95\% CIs excluding parity across two datasets and seven VLM judges. Because the two statement populations differ in selection and form, this ratio measures differential damage under the tested attack rather than a provenance-only causal effect. We then introduce panel-consensus verification, which cross-checks a quote against independently collected blind votes. It blocks 84.9\% of fabricated attacks, cuts their net harm by 97.5\%, and preserves the positive but statistically inconclusive point estimate for genuine peer information under leave-one-out re-verification. These results identify a low-cost attack surface and a concrete defense for safer multimodal collaborative evaluation.
Problem

Research questions and friction points this paper is trying to address.

forged peer judgments
source-blind anchoring
multimodal LLM judge panels
vision-language models
panel-consensus verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

source-blind anchoring
multimodal LLM judge panels
forged peer judgments
panel-consensus verification
vision-language models
🔎 Similar Papers
No similar papers found.