Institution profile

Xiaobing.AI

Industry researchasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

Jul 29, 2026

This work addresses the limitations of existing closed-source large language models in efficiently acquiring, updating, and reusing domain-specific skills, which often rely on costly human annotations or unreliable model-based evaluations. The authors propose an unsupervised self-evolving agent framework that, for the first time, adapts the destroy-and-reconstruct mechanism from diffusion models to skill learning. By contrasting agent-reconstructed text with original high-quality human-written text, the framework generates self-supervised signals to refine an external skill library—comprising skills that are readable, transferable, and composable—without updating the model’s internal weights. Integrating self-supervised contrastive learning with a semantic diffusion mechanism, the approach enables autonomous skill distillation and continuous iteration. Evaluated on short-form screenplay generation, it significantly improves output quality, demonstrating strong scalability and generalization in autonomously learning complex human artifacts.

0 citationsRead paper

Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning

Mar 22, 2025

Current large language models (LLMs) suffer from weak persona consistency and insufficient fine-grained perception of emotions and character-specific traits in role-playing tasks. Moreover, prevailing human-alignment approaches—such as high-quality supervised annotation or reinforcement learning from human feedback (RLHF)—are costly and ill-suited to the inherent diversity of character behaviors. To address these challenges, we propose Persona-aware Contrastive Learning (PCL), an unsupervised, annotation-free framework that introduces a novel role-chain self-questioning and iterative contrastive learning paradigm for persona alignment. PCL jointly optimizes black-box and white-box LLMs to enable fine-grained persona modeling without human labels. Extensive evaluations—including CharEval benchmarking, GPT-4 automated assessment, and expert human evaluation—demonstrate that PCL significantly improves persona consistency and dialogue personalization, surpassing conventional supervised and RL-based paradigms.

0 citationsRead paper
Recent publications

Latest Papers

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

Jul 29, 2026

This work addresses the limitations of existing closed-source large language models in efficiently acquiring, updating, and reusing domain-specific skills, which often rely on costly human annotations or unreliable model-based evaluations. The authors propose an unsupervised self-evolving agent framework that, for the first time, adapts the destroy-and-reconstruct mechanism from diffusion models to skill learning. By contrasting agent-reconstructed text with original high-quality human-written text, the framework generates self-supervised signals to refine an external skill library—comprising skills that are readable, transferable, and composable—without updating the model’s internal weights. Integrating self-supervised contrastive learning with a semantic diffusion mechanism, the approach enables autonomous skill distillation and continuous iteration. Evaluated on short-form screenplay generation, it significantly improves output quality, demonstrating strong scalability and generalization in autonomously learning complex human artifacts.

0 citationsRead paper

Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning

Mar 22, 2025

Current large language models (LLMs) suffer from weak persona consistency and insufficient fine-grained perception of emotions and character-specific traits in role-playing tasks. Moreover, prevailing human-alignment approaches—such as high-quality supervised annotation or reinforcement learning from human feedback (RLHF)—are costly and ill-suited to the inherent diversity of character behaviors. To address these challenges, we propose Persona-aware Contrastive Learning (PCL), an unsupervised, annotation-free framework that introduces a novel role-chain self-questioning and iterative contrastive learning paradigm for persona alignment. PCL jointly optimizes black-box and white-box LLMs to enable fine-grained persona modeling without human labels. Extensive evaluations—including CharEval benchmarking, GPT-4 automated assessment, and expert human evaluation—demonstrate that PCL significantly improves persona consistency and dialogue personalization, surpassing conventional supervised and RL-based paradigms.

0 citationsRead paper