Institution profile

Guangxi University

Academic institutionasia · cn
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

Aug 02, 2026

This work addresses the challenge in zero-shot image captioning where synthetic training data generated by text-to-image models often suffers from fine-grained entity misalignment—such as missing objects or mislocalized attributes—leading to distorted supervision signals. To mitigate this, the authors propose ReCap, a framework that explicitly detects image entities and guides caption rewriting to achieve fine-grained image-text alignment. ReCap further incorporates an adaptive dynamic weighting strategy to downweight unreliable synthetic samples during training. By shifting data refinement from implicit global matching to explicit entity-level realignment, the method introduces a plug-and-play mechanism for fine-grained correction. Experiments demonstrate that ReCap achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks, significantly improving caption consistency and supervision fidelity.

0 citationsRead paper

VCT: A Verifiable Transcript System for LLM Conversations

Jun 22, 2026

This work addresses the challenge that nonlinear operations in large language model (LLM) conversations—such as reprompting, response regeneration, message deletion, and concurrent multi-device interactions—cannot be faithfully captured by traditional linear logs, thereby undermining the reliability of digital forensics and compliance auditing. To resolve this, the paper introduces Verifiable Conversation Transcript (VCT), a novel system that models nonlinear dialogues as serialized state transitions with deletion barriers. VCT ensures integrity and accountable verifiability through a three-layer hash chain structure (spanning QA pairs, sessions, and account-level Merkle roots), joint user–server signatures, and an asynchronous view-fork detection mechanism. Prototype evaluation demonstrates sub-millisecond to low-millisecond cryptographic overhead for core operations and only 0.9% metadata overhead for 21KB transcripts, confirming VCT’s feasibility for high-assurance auditing in production-grade LLM platforms.

0 citationsRead paper
Recent publications

Latest Papers

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

Aug 02, 2026

This work addresses the challenge in zero-shot image captioning where synthetic training data generated by text-to-image models often suffers from fine-grained entity misalignment—such as missing objects or mislocalized attributes—leading to distorted supervision signals. To mitigate this, the authors propose ReCap, a framework that explicitly detects image entities and guides caption rewriting to achieve fine-grained image-text alignment. ReCap further incorporates an adaptive dynamic weighting strategy to downweight unreliable synthetic samples during training. By shifting data refinement from implicit global matching to explicit entity-level realignment, the method introduces a plug-and-play mechanism for fine-grained correction. Experiments demonstrate that ReCap achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks, significantly improving caption consistency and supervision fidelity.

0 citationsRead paper

VCT: A Verifiable Transcript System for LLM Conversations

Jun 22, 2026

This work addresses the challenge that nonlinear operations in large language model (LLM) conversations—such as reprompting, response regeneration, message deletion, and concurrent multi-device interactions—cannot be faithfully captured by traditional linear logs, thereby undermining the reliability of digital forensics and compliance auditing. To resolve this, the paper introduces Verifiable Conversation Transcript (VCT), a novel system that models nonlinear dialogues as serialized state transitions with deletion barriers. VCT ensures integrity and accountable verifiability through a three-layer hash chain structure (spanning QA pairs, sessions, and account-level Merkle roots), joint user–server signatures, and an asynchronous view-fork detection mechanism. Prototype evaluation demonstrates sub-millisecond to low-millisecond cryptographic overhead for core operations and only 0.9% metadata overhead for 21KB transcripts, confirming VCT’s feasibility for high-assurance auditing in production-grade LLM platforms.

0 citationsRead paper