Institution profile

Skyfall AI

Industry researchnorthamerica · us
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

Aug 17, 2026

This study addresses the exploration-exploitation imbalance caused by semantic priors in large language model decision-making by proposing a Semantic Multi-Armed Bandit framework. This approach formally quantifies how the alignment between linguistic labels and reward structures influences exploration strategies. Through in-context learning and inductive bias analysis, we reveal the bias effects inherent in label semantics and reward signals. Empirical results demonstrate that semantically consistent labels significantly enhance decision performance, whereas mismatches cause severe degradation. Furthermore, negative rewards elicit greater exploration than positive ones, validating scale biases present in pretraining data. Collectively, this work offers novel insights into understanding and optimizing the decision-making behaviors of large language models.

0 citationsRead paper

World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems

Jan 29, 2026

This work addresses the challenge that state-of-the-art large language models struggle to model latent enterprise workflows and their cascading side effects, often violating implicit constraints due to limited observability. To bridge this gap, we construct World of Workflows (WoW), a realistic enterprise simulation environment based on ServiceNow, encompassing over 4,000 business rules and 55 active workflows. We introduce the WoW-bench benchmark to evaluate agents’ capabilities in constrained task completion and dynamic system modeling. Our study reveals, for the first time, a “dynamic blind spot” of large models in enterprise settings and advocates for embodied world models that explicitly learn latent system states. Experiments demonstrate that agents equipped with latent state simulation significantly improve both task success rates and compliance with implicit constraints, establishing a new paradigm for reliable enterprise AI agents.

0 citationsRead paper
Recent publications

Latest Papers

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

Aug 17, 2026

This study addresses the exploration-exploitation imbalance caused by semantic priors in large language model decision-making by proposing a Semantic Multi-Armed Bandit framework. This approach formally quantifies how the alignment between linguistic labels and reward structures influences exploration strategies. Through in-context learning and inductive bias analysis, we reveal the bias effects inherent in label semantics and reward signals. Empirical results demonstrate that semantically consistent labels significantly enhance decision performance, whereas mismatches cause severe degradation. Furthermore, negative rewards elicit greater exploration than positive ones, validating scale biases present in pretraining data. Collectively, this work offers novel insights into understanding and optimizing the decision-making behaviors of large language models.

0 citationsRead paper

World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems

Jan 29, 2026

This work addresses the challenge that state-of-the-art large language models struggle to model latent enterprise workflows and their cascading side effects, often violating implicit constraints due to limited observability. To bridge this gap, we construct World of Workflows (WoW), a realistic enterprise simulation environment based on ServiceNow, encompassing over 4,000 business rules and 55 active workflows. We introduce the WoW-bench benchmark to evaluate agents’ capabilities in constrained task completion and dynamic system modeling. Our study reveals, for the first time, a “dynamic blind spot” of large models in enterprise settings and advocates for embodied world models that explicitly learn latent system states. Experiments demonstrate that agents equipped with latent state simulation significantly improve both task success rates and compliance with implicit constraints, establishing a new paradigm for reliable enterprise AI agents.

0 citationsRead paper