TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

πŸ“… 2026-08-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of balancing accuracy, fidelity, and editability in cross-border e-commerce image translation by proposing a Structured Visual Code framework that reframes translation as generating renderable HTML patches. By decoupling semantics from pixel rendering, the approach integrates Vision-Language Model comprehension and diffusion-based inpainting with a three-stage post-training strategy comprising supervised fine-tuning, self-distillation, and reinforcement learning. Additionally, a specialized multilingual dataset and benchmark are constructed to facilitate evaluation. Experimental results demonstrate that this method significantly outperforms cascaded pipelines and mainstream systems, achieving precise, controllable, and easily editable cross-lingual translation. Consequently, this work provides an efficient solution for e-commerce scenarios requiring high-quality visual text translation.
πŸ“ Abstract
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously. To address these challenges, we introduce TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages. Our framework decouples semantic generation from pixel rendering: a vision-language model (VLM) handles visual understanding, cross-lingual translation, and structured visual generation, while a diffusion model performs background inpainting and pixel-level refinement, followed by deterministic rendering to synthesize the final image. Based on this formulation, we develop a three-stage post-training framework, where supervised fine-tuning (SFT) establishes the image-to-code mapping, privilege-gap weighted self-distillation (PWSD) improves the learning of style and layout tokens, and reinforcement learning with verifiable rewards (RLVR) further optimizes task-level performance. We further introduce TransAnyDataset and TransAnyBench, a multilingual dataset and benchmark for e-commerce image translation. Extensive experiments demonstrate competitive performance against cascaded pipelines, open-source end-to-end models, and closed-source image editing systems, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.
Problem

Research questions and friction points this paper is trying to address.

E-commerce Image Translation
Visual Identity Preservation
Editable Output
Cross-border E-commerce
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Visual Generation
HTML Patch Generation
Privilege-Gap Weighted Self-Distillation
Reinforcement Learning with Verifiable Rewards
Decoupled Rendering
πŸ”Ž Similar Papers
No similar papers found.
X
Xiaoan Liu
Wuhan University
L
Lichen Ma
JD.com
Z
Zipeng Guo
JD.com
Y
Yu He
JD.com
X
Xiaoyan Su
The Hong Kong University of Science and Technology (Guangzhou)
S
Shaojie Guo
JD.com
H
Hao Yang
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
J
Jingling Fu
JD.com
X
Xiaolong Fu
JD.com
Z
Zhen Chen
JD.com
Yu Guo
Yu Guo
Xi’an Jiaotong University
6D pose estimationtime series predictiongraph learning
F
Fei Wang
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
Xinyi Liu
Xinyi Liu
Wuhan University
3D ReconstructionPoint Cloud and Image IntegrationComputational Origami
Yongjun Zhang
Yongjun Zhang
Wuhan University
PhotogrammetryRemote SensingComputer Vision
K
Ke Zhang
JD.com
Junshi Huang
Junshi Huang
Meituan
Computer VisionNLPMachine Learning