Institution profile

Taiyuan University of Science and Technology

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views

Aug 09, 2026

This work addresses the limitation of standard 3D Gaussian splatting, which assumes pristine input images and struggles with multi-view datasets containing mixed-quality imagery affected by JPEG compression artifacts. To overcome this, the authors propose the first method that explicitly incorporates JPEG compression information into the training pipeline by leveraging quantization tables to construct view-specific observation operators, enabling consistent supervision directly in the compressed domain. They further introduce a DCT-band-weighted loss and a Gaussian regularization strategy based on block-wise inconsistency to guide the optimization toward more robust scene representations. Evaluated across seven scenes under three mixed-quality configurations, the proposed approach consistently achieves the lowest average LPIPS and highest average SSIM while maintaining a rendering speed of approximately 150 FPS.

0 citationsRead paper

Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams

Aug 23, 2025

Current large vision-language models (LVLMs) exhibit severe robustness deficiencies in understanding scientific charts under realistic visual perturbations—such as noise, blur, and occlusion—while existing evaluation benchmarks lack systematic assessment of such perturbations. Method: We propose the first programmatic perturbation-robustness evaluation framework for scientific chart understanding: (1) constructing a large-scale scientific chart QA dataset with diverse, controllable perturbations; (2) designing an adaptive multi-view reasoning mechanism with consistency-based self-correction; and (3) introducing two novel quantitative metrics—Perturbation Robustness Consistency (PRC) and Visual Degradation Sensitivity (VDC). Results: Experiments demonstrate substantial improvements in LVLM reasoning performance across perturbation types. Notably, mainstream models—including GPT-4V—suffer over 13 percentage-point accuracy drops under perturbation, revealing critical robustness gaps. Our framework establishes a new paradigm for evaluating and advancing robustness in scientific multimodal models.

0 citationsRead paper
Recent publications

Latest Papers

JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views

Aug 09, 2026

This work addresses the limitation of standard 3D Gaussian splatting, which assumes pristine input images and struggles with multi-view datasets containing mixed-quality imagery affected by JPEG compression artifacts. To overcome this, the authors propose the first method that explicitly incorporates JPEG compression information into the training pipeline by leveraging quantization tables to construct view-specific observation operators, enabling consistent supervision directly in the compressed domain. They further introduce a DCT-band-weighted loss and a Gaussian regularization strategy based on block-wise inconsistency to guide the optimization toward more robust scene representations. Evaluated across seven scenes under three mixed-quality configurations, the proposed approach consistently achieves the lowest average LPIPS and highest average SSIM while maintaining a rendering speed of approximately 150 FPS.

0 citationsRead paper

Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams

Aug 23, 2025

Current large vision-language models (LVLMs) exhibit severe robustness deficiencies in understanding scientific charts under realistic visual perturbations—such as noise, blur, and occlusion—while existing evaluation benchmarks lack systematic assessment of such perturbations. Method: We propose the first programmatic perturbation-robustness evaluation framework for scientific chart understanding: (1) constructing a large-scale scientific chart QA dataset with diverse, controllable perturbations; (2) designing an adaptive multi-view reasoning mechanism with consistency-based self-correction; and (3) introducing two novel quantitative metrics—Perturbation Robustness Consistency (PRC) and Visual Degradation Sensitivity (VDC). Results: Experiments demonstrate substantial improvements in LVLM reasoning performance across perturbation types. Notably, mainstream models—including GPT-4V—suffer over 13 percentage-point accuracy drops under perturbation, revealing critical robustness gaps. Our framework establishes a new paradigm for evaluating and advancing robustness in scientific multimodal models.

0 citationsRead paper