Institution profile

Beijing Wenge Technology Co., Ltd

Industry researchasia · cn
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

Jul 29, 2026

This work addresses a critical limitation in existing benchmarks for scientific equation discovery, which often fail to distinguish whether models genuinely infer underlying laws or merely reproduce known formulas. To this end, the authors propose the LSR-Synth evaluation framework, which introduces novel synthetic terms into established mechanisms and incorporates strategies such as semantic blinding, operator library weakening, and exclusion of matching operator families to construct a semantics-free baseline. This setup rigorously isolates and quantifies the marginal contribution of language model priors. Experiments reveal that, under current task sets and search budgets, fixed operator libraries already cover most problems; only when lexical coverage is selectively disrupted do language model–generated candidate solutions substantially increase the number of solvable instances, thereby providing the first quantitative evidence of their practical value in out-of-distribution symbolic regression.

0 citationsRead paper

MatMind: A Structure-Activity Knowledge-Driven Generative Foundation Model for Materials Science

Jun 05, 2026

This work proposes the first generative foundation model for crystalline materials, addressing the limitations of conventional task-specific AI models that struggle to jointly handle structural representation, property prediction, and structure–activity reasoning. By integrating structure–activity knowledge injection, a dual-head joint training architecture, and multi-objective physics-informed reinforcement learning, the model achieves synergistic optimization of symbolic reasoning and numerical regression within a unified framework. It attains state-of-the-art accuracy in predicting formation energy above the convex hull, bulk modulus, and bandgap, while achieving a 65.3% S.U.N. rate for unconditional crystal generation. Furthermore, it significantly improves success rates in conditional generation tasks involving rare magnetization densities, outperforming existing specialized models across multiple benchmarks.

0 citationsRead paper

Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics

May 16, 2026

This work addresses the frequent neglect of logical rigor in scientific reasoning by current large language models, which undermines the reliability of their conclusions. The study introduces scientific logicality as a core evaluation dimension and systematically develops a logic-oriented assessment framework alongside a principled data sampling methodology. A high-logicality question dataset is constructed from physics literature to support this approach. Through a logic-guided training paradigm, the authors demonstrate significant improvements in both logical faithfulness and task performance across three mainstream large language models. These results establish that enhancing logical coherence plays a critical role in advancing the models’ capacity for solving scientific problems.

0 citationsRead paper

DongYuan: An LLM-Based Framework for Integrative Chinese and Western Medicine Spleen-Stomach Disorders Diagnosis

Mar 30, 2026

This study addresses three key challenges in integrative Chinese and Western medicine for gastrointestinal disorders: the scarcity of high-quality data, the disconnect between traditional Chinese medicine (TCM) pattern differentiation and Western diagnostic logic, and the absence of standardized evaluation benchmarks. To tackle these issues, the authors propose the DongYuan framework, which introduces SSDF-Bench—a dedicated high-quality dataset and evaluation benchmark—and SSDF-Core, a large language model trained via a two-stage paradigm that integrates TCM syndrome differentiation with Western diagnostic reasoning through supervised fine-tuning (SFT) and direct preference optimization (DPO). Additionally, the framework incorporates SSDF-Navigator, a plug-and-play inquiry navigation module to refine clinical questioning strategies. Experimental results demonstrate that SSDF-Core significantly outperforms twelve mainstream baseline models on SSDF-Bench, establishing a methodological and technical foundation for intelligent integrative diagnosis.

0 citationsRead paper

APEX-Searcher: Augmenting LLMs' Search Capabilities through Agentic Planning and Execution

Mar 14, 2026

This work addresses the challenges in complex multi-hop question answering, where single-round retrieval often fails to support accurate reasoning, and existing end-to-end approaches suffer from ambiguous task decomposition and sparse rewards in reinforcement learning, leading to inaccurate retrieval and degraded performance. To overcome these limitations, the authors propose APEX-Searcher, a novel framework that decouples retrieval into planning and execution phases. It first optimizes task planning via a decomposition-aware reward mechanism in reinforcement learning and then enhances iterative subtask execution through supervised fine-tuning on high-quality multi-hop trajectories. This approach effectively mitigates path ambiguity and reward sparsity, significantly improving performance in multi-hop retrieval-augmented generation and task planning across multiple benchmarks, thereby demonstrating its effectiveness and robustness in complex retrieval tasks.

0 citationsRead paper
Recent publications

Latest Papers

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

Jul 29, 2026

This work addresses a critical limitation in existing benchmarks for scientific equation discovery, which often fail to distinguish whether models genuinely infer underlying laws or merely reproduce known formulas. To this end, the authors propose the LSR-Synth evaluation framework, which introduces novel synthetic terms into established mechanisms and incorporates strategies such as semantic blinding, operator library weakening, and exclusion of matching operator families to construct a semantics-free baseline. This setup rigorously isolates and quantifies the marginal contribution of language model priors. Experiments reveal that, under current task sets and search budgets, fixed operator libraries already cover most problems; only when lexical coverage is selectively disrupted do language model–generated candidate solutions substantially increase the number of solvable instances, thereby providing the first quantitative evidence of their practical value in out-of-distribution symbolic regression.

0 citationsRead paper

MatMind: A Structure-Activity Knowledge-Driven Generative Foundation Model for Materials Science

Jun 05, 2026

This work proposes the first generative foundation model for crystalline materials, addressing the limitations of conventional task-specific AI models that struggle to jointly handle structural representation, property prediction, and structure–activity reasoning. By integrating structure–activity knowledge injection, a dual-head joint training architecture, and multi-objective physics-informed reinforcement learning, the model achieves synergistic optimization of symbolic reasoning and numerical regression within a unified framework. It attains state-of-the-art accuracy in predicting formation energy above the convex hull, bulk modulus, and bandgap, while achieving a 65.3% S.U.N. rate for unconditional crystal generation. Furthermore, it significantly improves success rates in conditional generation tasks involving rare magnetization densities, outperforming existing specialized models across multiple benchmarks.

0 citationsRead paper

Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics

May 16, 2026

This work addresses the frequent neglect of logical rigor in scientific reasoning by current large language models, which undermines the reliability of their conclusions. The study introduces scientific logicality as a core evaluation dimension and systematically develops a logic-oriented assessment framework alongside a principled data sampling methodology. A high-logicality question dataset is constructed from physics literature to support this approach. Through a logic-guided training paradigm, the authors demonstrate significant improvements in both logical faithfulness and task performance across three mainstream large language models. These results establish that enhancing logical coherence plays a critical role in advancing the models’ capacity for solving scientific problems.

0 citationsRead paper

DongYuan: An LLM-Based Framework for Integrative Chinese and Western Medicine Spleen-Stomach Disorders Diagnosis

Mar 30, 2026

This study addresses three key challenges in integrative Chinese and Western medicine for gastrointestinal disorders: the scarcity of high-quality data, the disconnect between traditional Chinese medicine (TCM) pattern differentiation and Western diagnostic logic, and the absence of standardized evaluation benchmarks. To tackle these issues, the authors propose the DongYuan framework, which introduces SSDF-Bench—a dedicated high-quality dataset and evaluation benchmark—and SSDF-Core, a large language model trained via a two-stage paradigm that integrates TCM syndrome differentiation with Western diagnostic reasoning through supervised fine-tuning (SFT) and direct preference optimization (DPO). Additionally, the framework incorporates SSDF-Navigator, a plug-and-play inquiry navigation module to refine clinical questioning strategies. Experimental results demonstrate that SSDF-Core significantly outperforms twelve mainstream baseline models on SSDF-Bench, establishing a methodological and technical foundation for intelligent integrative diagnosis.

0 citationsRead paper

APEX-Searcher: Augmenting LLMs' Search Capabilities through Agentic Planning and Execution

Mar 14, 2026

This work addresses the challenges in complex multi-hop question answering, where single-round retrieval often fails to support accurate reasoning, and existing end-to-end approaches suffer from ambiguous task decomposition and sparse rewards in reinforcement learning, leading to inaccurate retrieval and degraded performance. To overcome these limitations, the authors propose APEX-Searcher, a novel framework that decouples retrieval into planning and execution phases. It first optimizes task planning via a decomposition-aware reward mechanism in reinforcement learning and then enhances iterative subtask execution through supervised fine-tuning on high-quality multi-hop trajectories. This approach effectively mitigates path ambiguity and reward sparsity, significantly improving performance in multi-hop retrieval-augmented generation and task planning across multiple benchmarks, thereby demonstrating its effectiveness and robustness in complex retrieval tasks.

0 citationsRead paper