Institution profile

Shanda Group

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning

Jan 05, 2026arXiv.org

This work addresses the challenge that large language models, constrained by finite context windows, struggle to maintain coherent behavior over long-term interactions due to existing memory systems predominantly storing isolated records without effectively modeling user state evolution or resolving conflicts. To overcome this limitation, the paper proposes a neuroscience-inspired self-organizing memory operating system that enables structured long-term reasoning through a three-stage process: episodic memory unit generation, semantic integration, and reconstructive recall. Key innovations include the MemCell and MemScene architectures, time-bounded Foresight signals, and a theme-driven mechanism for organizing memory scenes, collectively supporting dynamic user profile updates and conflict resolution. The system achieves state-of-the-art performance on LoCoMo and LongMemEval benchmarks and demonstrates superior capabilities in user modeling and proactive dialogue, as validated by evaluations on PersonaMem v2 and case studies.

3 citationsRead paper

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity

Jul 15, 2026

This work addresses how to enhance the performance of large language models on real-world tasks without updating their weights, while avoiding spurious improvements caused by self-generated feedback noise or overfitting. The authors propose a self-evolving agent framework in which the language model diagnoses failures and generates patches, while sampling, evaluation, and significance testing are handled by deterministic code to ensure credible improvements. A key innovation is the introduction of a Gated Quality-Diversity archive based on Pathological Semantics (WHERE×WHY), which categorizes patches according to root causes, enabling generalization across tasks and model families. Evaluated on seven domains using frozen open-source models, the method achieves performance gains of 9–15.5 percentage points on held-out test sets, preserving 86%–147% of the training-time gains, thereby demonstrating both efficacy and transferability.

0 citationsRead paper

MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors

Apr 08, 2026

This study addresses the significant decline in diagnostic accuracy of large language models (LLMs) when interacting with non-cooperative patients, a challenge inadequately captured by existing evaluations lacking systematic characterization of multidimensional, graded adversarial behaviors. The authors propose MedDialBench, a novel benchmark that introduces a five-dimensional model of adversarial behavior—encompassing logical consistency, health literacy, expressive style, information disclosure, and attitude—paired with case-specific dialogue scripts. Through a factorial experimental design enabling dose–response analysis and cross-dimensional interaction detection across 7,225 dialogues, they find that symptom fabrication reduces diagnostic accuracy by 38.8–54.1 percentage points, exerting 1.7–3.4 times greater harm than information omission. Crucially, fabrication not only severely impairs all models individually but also induces strong super-additive failure effects when combined with other dimensions, revealing that information pollution is substantially more detrimental than information scarcity.

0 citationsRead paper
Recent publications

Latest Papers

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity

Jul 15, 2026

This work addresses how to enhance the performance of large language models on real-world tasks without updating their weights, while avoiding spurious improvements caused by self-generated feedback noise or overfitting. The authors propose a self-evolving agent framework in which the language model diagnoses failures and generates patches, while sampling, evaluation, and significance testing are handled by deterministic code to ensure credible improvements. A key innovation is the introduction of a Gated Quality-Diversity archive based on Pathological Semantics (WHERE×WHY), which categorizes patches according to root causes, enabling generalization across tasks and model families. Evaluated on seven domains using frozen open-source models, the method achieves performance gains of 9–15.5 percentage points on held-out test sets, preserving 86%–147% of the training-time gains, thereby demonstrating both efficacy and transferability.

0 citationsRead paper

MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors

Apr 08, 2026

This study addresses the significant decline in diagnostic accuracy of large language models (LLMs) when interacting with non-cooperative patients, a challenge inadequately captured by existing evaluations lacking systematic characterization of multidimensional, graded adversarial behaviors. The authors propose MedDialBench, a novel benchmark that introduces a five-dimensional model of adversarial behavior—encompassing logical consistency, health literacy, expressive style, information disclosure, and attitude—paired with case-specific dialogue scripts. Through a factorial experimental design enabling dose–response analysis and cross-dimensional interaction detection across 7,225 dialogues, they find that symptom fabrication reduces diagnostic accuracy by 38.8–54.1 percentage points, exerting 1.7–3.4 times greater harm than information omission. Crucially, fabrication not only severely impairs all models individually but also induces strong super-additive failure effects when combined with other dimensions, revealing that information pollution is substantially more detrimental than information scarcity.

0 citationsRead paper

EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning

Jan 05, 2026arXiv.org

This work addresses the challenge that large language models, constrained by finite context windows, struggle to maintain coherent behavior over long-term interactions due to existing memory systems predominantly storing isolated records without effectively modeling user state evolution or resolving conflicts. To overcome this limitation, the paper proposes a neuroscience-inspired self-organizing memory operating system that enables structured long-term reasoning through a three-stage process: episodic memory unit generation, semantic integration, and reconstructive recall. Key innovations include the MemCell and MemScene architectures, time-bounded Foresight signals, and a theme-driven mechanism for organizing memory scenes, collectively supporting dynamic user profile updates and conflict resolution. The system achieves state-of-the-art performance on LoCoMo and LongMemEval benchmarks and demonstrates superior capabilities in user modeling and proactive dialogue, as validated by evaluations on PersonaMem v2 and case studies.

3 citationsRead paper