TransPhy: Visual In-Context Learning for Physically Grounded Image Editing

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决物理基础图像变换问题,提出TransPhy框架,通过物理规则归纳和转换对齐渲染方法提高现有视觉上下文学习的规则遵循性、查询一致性和泛化能力。
📝 Abstract
Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically grounded VICL requires a model to infer the demonstrated transformation, adapt its effects to the query-specific scene context, and preserve rule-irrelevant content. We introduce PhysVICL-74, comprising 74 physically grounded transformation rules and 5,240 source--target image pairs that form nearly 75K training and evaluation contexts. Its benchmark split separately evaluates novel-instance transfer and unseen-rule generalization. We further propose TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering. TransPhy first predicts the demonstrated rule and an explicit query-specific target-state description, and then synthesizes the target image through token-wise mixture-of-experts adaptation, with expert routing guided by localized transition cues. Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization over existing visual in-context editing methods.
Problem

Research questions and friction points this paper is trying to address.

physically grounded transformations
visual in-context learning
material properties
geometry
object interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physically grounded transformations
Visual in-context learning (VICL)
TransPhy framework
Physical-rule induction
Transition-aligned rendering
S
Siyi Xie
Peking University
X
Xuanke Shi
SenseTime Research
J
Jinsheng Quan
Zhejiang University
Haoran Tang
Haoran Tang
Peking University
Text-video retrievalVideo-LLM
Z
Zukai Chen
SenseTime Research
Lei Yang
Lei Yang
SenseTime Research
Machine LearningComputer Vision
Quan Wang
Quan Wang
sensetime