Institution profile

INFLY TECH

Industry research
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Visual-Semantic Knowledge Conflicts in Operating Rooms: Synthetic Data Curation for Surgical Risk Perception in Multimodal Large Language Models

Jun 25, 2025

In operating room (OR) risk identification, multimodal large language models (MLLMs) suffer from visual-semantic knowledge conflict (VS-KC)—i.e., strong textual rule comprehension but poor detection of safety violations in images. To address this, we propose a diffusion-based method to synthesize rule-violating scenarios, generating 34,000 high-fidelity synthetic images containing diverse safety violations, augmented with 214 human-annotated ground-truth images, forming the first open-source OR-VSKC dataset and benchmark. This dataset systematically exposes MLLMs’ inconsistency in knowledge alignment at the violation-entity level—a previously uncharacterized deficiency. Fine-tuned models achieve significant gains in detecting known violation types and demonstrate viewpoint generalization; however, performance degrades markedly on unseen entities, underscoring the critical need for cross-entity coverage in training. Our work establishes a foundational resource and diagnostic framework for advancing robust, entity-aware safety reasoning in surgical AI.

0 citationsRead paper
Recent publications

Latest Papers

Visual-Semantic Knowledge Conflicts in Operating Rooms: Synthetic Data Curation for Surgical Risk Perception in Multimodal Large Language Models

Jun 25, 2025

In operating room (OR) risk identification, multimodal large language models (MLLMs) suffer from visual-semantic knowledge conflict (VS-KC)—i.e., strong textual rule comprehension but poor detection of safety violations in images. To address this, we propose a diffusion-based method to synthesize rule-violating scenarios, generating 34,000 high-fidelity synthetic images containing diverse safety violations, augmented with 214 human-annotated ground-truth images, forming the first open-source OR-VSKC dataset and benchmark. This dataset systematically exposes MLLMs’ inconsistency in knowledge alignment at the violation-entity level—a previously uncharacterized deficiency. Fine-tuned models achieve significant gains in detecting known violation types and demonstrate viewpoint generalization; however, performance degrades markedly on unseen entities, underscoring the critical need for cross-entity coverage in training. Our work establishes a foundational resource and diagnostic framework for advancing robust, entity-aware safety reasoning in surgical AI.

0 citationsRead paper