Institution profile

NNAISENSE

Industry researcheurope · ch
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

Jun 15, 2026

This work addresses the vulnerability of large language model (LLM) search agents to adversarial manipulation, wherein attacker-controlled web content may be erroneously treated as credible evidence, leading LLMs to endorse harmful claims. To systematically evaluate this risk, the authors propose SearchGEO, a novel evaluation framework that introduces recommendation reliability as a core dimension of LLM backend safety. SearchGEO establishes a controlled and reproducible paradigm for assessing endorsement vulnerabilities through an integrated pipeline comprising web evidence manipulation, five adversarial attack patterns, multi-level output metrics, and auxiliary skill probes—such as command conversion. Evaluations across 13 mainstream LLMs on 308 cases reveal attack success rates ranging from 0.0% (Claude-Sonnet-4.6) to 31.4% (Gemini-3-Flash), with substantial response variation even among models of similar architecture; auxiliary probes further indicate that Claude tends toward excessive refusal, whereas GPT models exhibit undue trust.

0 citationsRead paper

PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors

Jul 21, 2025

Existing benchmarks lack the capability to rigorously evaluate large language model (LLM)-based agents’ scientific reasoning—particularly their dependence on prior knowledge and adaptability to environmental complexity. Method: We introduce the first benchmark for scientific discovery in interactive physical environments, enabling fine-grained, controllable modulation of agents’ prior knowledge levels and multidimensional disentanglement of performance across hypothesis generation, active exploration, and constrained data acquisition. The benchmark integrates high-fidelity physics simulation, structured data collection, and standardized evaluation protocols to ensure reproducibility and quantifiability. Contribution/Results: Experiments demonstrate significant performance divergence across LLMs under varying prior knowledge availability and task complexity, validating the benchmark’s strong discriminative power and scalability. It establishes a principled, extensible framework for evaluating and advancing LLM agents in scientific reasoning tasks.

0 citationsRead paper

Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models

Mar 11, 2025

Existing long-form writing agents rely on predefined outlines and rigid pipelines, resulting in inflexible coordination among information retrieval, reasoning, and generation—compromising adaptability and fidelity to human writing styles. This paper proposes an agent framework for long-text generation that abandons static planning in favor of a novel heterogeneous recursive planning mechanism. It enables dynamic, real-time re-decomposition and seamless integration across retrieval, reasoning, and writing tasks. Leveraging language-model-driven recursive task decomposition, adaptive workflow scheduling, and cross-task joint modeling, the framework achieves end-to-end adaptability. Evaluated on novel writing and technical report generation, it consistently surpasses state-of-the-art methods across all automated metrics—demonstrating superior effectiveness, generalizability, and stylistic consistency with human-authored texts.

0 citationsRead paper

On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers

Feb 08, 2025

This work establishes, for the first time, a unified theoretical framework for convergence and noise robustness of three emerging sequence-based RL paradigms—Episodic Upside-Down RL, Goal-Conditioned Supervised Learning, and Online Decision Transformers—under Markovian dynamics. Methodologically, it introduces episodic state-space modeling, quotient-topological continuity analysis, and dynamical systems fixed-point theory, integrated with transition kernel perturbation analysis to derive explicit kernel-dependent bounds on policy performance, value functions, and goal-reaching capability. Theoretically, it proves that algorithms converge to near-optimal policies as the transition kernel approaches determinism, and that solutions are continuous and asymptotically stable in the quotient topology. Numerical experiments empirically validate these theoretical guarantees. This work provides the first rigorous, unifying foundation for supervised and sequence-based RL, bridging theoretical analysis with practical algorithm design.

0 citationsRead paper

Upside Down Reinforcement Learning with Policy Generators

Jan 27, 2025

This work addresses the low policy-generation efficiency and poor generalization in command-driven reinforcement learning. We propose a critic-free, end-to-end framework. Our key contributions are: (1) a Hypernetworks-based, command-conditioned policy generator that directly synthesizes neural network policies satisfying specified return targets; (2) a decoupled sampling mechanism that separates buffer sampling probability from policy count, incorporating weighted sampling to enhance training stability; and (3) zero-shot generalization to unseen return targets within the Universal Dynamic Reward Learning (UDRL) framework. Experiments on multiple benchmark tasks demonstrate substantial improvements in sample efficiency and high-return policy performance, alongside strong cross-target generalization capability—achieving effective policy synthesis for novel return specifications without additional fine-tuning.

0 citationsRead paper
Recent publications

Latest Papers

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

Jun 15, 2026

This work addresses the vulnerability of large language model (LLM) search agents to adversarial manipulation, wherein attacker-controlled web content may be erroneously treated as credible evidence, leading LLMs to endorse harmful claims. To systematically evaluate this risk, the authors propose SearchGEO, a novel evaluation framework that introduces recommendation reliability as a core dimension of LLM backend safety. SearchGEO establishes a controlled and reproducible paradigm for assessing endorsement vulnerabilities through an integrated pipeline comprising web evidence manipulation, five adversarial attack patterns, multi-level output metrics, and auxiliary skill probes—such as command conversion. Evaluations across 13 mainstream LLMs on 308 cases reveal attack success rates ranging from 0.0% (Claude-Sonnet-4.6) to 31.4% (Gemini-3-Flash), with substantial response variation even among models of similar architecture; auxiliary probes further indicate that Claude tends toward excessive refusal, whereas GPT models exhibit undue trust.

0 citationsRead paper

PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors

Jul 21, 2025

Existing benchmarks lack the capability to rigorously evaluate large language model (LLM)-based agents’ scientific reasoning—particularly their dependence on prior knowledge and adaptability to environmental complexity. Method: We introduce the first benchmark for scientific discovery in interactive physical environments, enabling fine-grained, controllable modulation of agents’ prior knowledge levels and multidimensional disentanglement of performance across hypothesis generation, active exploration, and constrained data acquisition. The benchmark integrates high-fidelity physics simulation, structured data collection, and standardized evaluation protocols to ensure reproducibility and quantifiability. Contribution/Results: Experiments demonstrate significant performance divergence across LLMs under varying prior knowledge availability and task complexity, validating the benchmark’s strong discriminative power and scalability. It establishes a principled, extensible framework for evaluating and advancing LLM agents in scientific reasoning tasks.

0 citationsRead paper

Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models

Mar 11, 2025

Existing long-form writing agents rely on predefined outlines and rigid pipelines, resulting in inflexible coordination among information retrieval, reasoning, and generation—compromising adaptability and fidelity to human writing styles. This paper proposes an agent framework for long-text generation that abandons static planning in favor of a novel heterogeneous recursive planning mechanism. It enables dynamic, real-time re-decomposition and seamless integration across retrieval, reasoning, and writing tasks. Leveraging language-model-driven recursive task decomposition, adaptive workflow scheduling, and cross-task joint modeling, the framework achieves end-to-end adaptability. Evaluated on novel writing and technical report generation, it consistently surpasses state-of-the-art methods across all automated metrics—demonstrating superior effectiveness, generalizability, and stylistic consistency with human-authored texts.

0 citationsRead paper

On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers

Feb 08, 2025

This work establishes, for the first time, a unified theoretical framework for convergence and noise robustness of three emerging sequence-based RL paradigms—Episodic Upside-Down RL, Goal-Conditioned Supervised Learning, and Online Decision Transformers—under Markovian dynamics. Methodologically, it introduces episodic state-space modeling, quotient-topological continuity analysis, and dynamical systems fixed-point theory, integrated with transition kernel perturbation analysis to derive explicit kernel-dependent bounds on policy performance, value functions, and goal-reaching capability. Theoretically, it proves that algorithms converge to near-optimal policies as the transition kernel approaches determinism, and that solutions are continuous and asymptotically stable in the quotient topology. Numerical experiments empirically validate these theoretical guarantees. This work provides the first rigorous, unifying foundation for supervised and sequence-based RL, bridging theoretical analysis with practical algorithm design.

0 citationsRead paper

Upside Down Reinforcement Learning with Policy Generators

Jan 27, 2025

This work addresses the low policy-generation efficiency and poor generalization in command-driven reinforcement learning. We propose a critic-free, end-to-end framework. Our key contributions are: (1) a Hypernetworks-based, command-conditioned policy generator that directly synthesizes neural network policies satisfying specified return targets; (2) a decoupled sampling mechanism that separates buffer sampling probability from policy count, incorporating weighted sampling to enhance training stability; and (3) zero-shot generalization to unseen return targets within the Universal Dynamic Reward Learning (UDRL) framework. Experiments on multiple benchmark tasks demonstrate substantial improvements in sample efficiency and high-return policy performance, alongside strong cross-target generalization capability—achieving effective policy synthesis for novel return specifications without additional fine-tuning.

0 citationsRead paper