Institution profile

Upstage AI Research

Industry researchasia · kr
Official website
Research library25linked papers
Opportunities0open roles
Selected work

Representative Papers

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

Jul 05, 2026

This work addresses the challenge in reinforcement learning for instruction following where static prompts mismatch the policy’s evolving capabilities, resulting in poorly discriminative reward signals. To resolve this, the authors propose the LLM-as-a-Tutor framework, which treats prompt adaptation as a policy-aware process. Leveraging a large language model as both examiner and tutor, the method identifies non-challenging prompts through pairwise comparisons and monotonically increases task difficulty by appending atomic constraints—ensuring training signals remain calibrated to the policy’s current proficiency. Notably, this approach requires no external scheduling mechanism and significantly outperforms policy-agnostic baselines and existing adaptive methods across three complex instruction-following benchmarks.

0 citationsRead paper

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Jun 30, 2026

This work addresses a critical limitation in existing post-training methods for medical multimodal large language models, which overly prioritize final answer correctness while neglecting the optimization of intermediate reasoning steps, thereby suffering from cascading failures triggered by early errors. To mitigate this, we propose Medical Reasoning-aware Policy Optimization (MRPO), the first reinforcement learning framework tailored for medical multimodal reasoning that incorporates step-aware feedback. MRPO introduces step-level process rewards, applying exponentially stronger penalties to ineffective early reasoning even when the final answer is incorrect, thus effectively interrupting error propagation. Experimental results demonstrate that MRPO consistently outperforms standard GRPO and state-of-the-art reinforcement learning baselines across three multimodal LLM backbones. Notably, on Qwen3-VL-8B-Instruct, it surpasses HuatuoGPT-Vision-34B by 2.79 points and reduces early reasoning failure rates from 64.0% to 13.0%.

0 citationsRead paper

Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents

Jun 25, 2026

This study addresses the lack of effective evaluation for breadth-first search capabilities—specifically, exhaustive enumeration of members within a closed set and their structured attributes—in non-English contexts. It presents the first Korean-language benchmark for breadth-oriented agent assessment, leveraging an automated synthesis and validation pipeline to generate tasks that require agents to fully enumerate members of a given parent entity and populate their structured attribute tables. The work introduces a novel structured difficulty modulation mechanism, controlling table width and two-dimensional composite keys, alongside a unified normalized matcher and a multidimensional scoring framework (Item-, Column-, and Row-F1). Experiments reveal that while current agents achieve strong member identification (Item-F1: 92.8), they struggle significantly with complete row completion (Row-F1: 53.7), with performance degrading as task difficulty increases, thereby underscoring the benchmark’s necessity and challenge.

0 citationsRead paper

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

Jun 20, 2026

This work addresses the critical issue that existing biomedical agents often generate responses inconsistent with cited literature, a flaw overlooked by conventional evaluations reliant on fixed-answer benchmarks. To tackle this, the authors introduce the first retrieval-based agent evaluation benchmark for open-ended biomedical questions, comprising 12,553 unresolved queries spanning 12 domains. Agents must perform multi-turn tool-augmented reasoning and autonomously decide when to abstain from answering, enabling assessment of both factual faithfulness and refusal capability. Key innovations include an open-question paradigm without predefined answers, validation of question openness via subsequent real-world literature, objective difficulty calibration based on reference model failure rates, and a frozen checklist that significantly improves annotation consistency (Spearman correlation rising from 0.35 to 0.82). Experiments reveal that even state-of-the-art agents achieve only 29%–60% success on the hardest subset and exhibit pronounced tool-use degradation under high difficulty.

0 citationsRead paper
Recent publications

Latest Papers

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

Jul 05, 2026

This work addresses the challenge in reinforcement learning for instruction following where static prompts mismatch the policy’s evolving capabilities, resulting in poorly discriminative reward signals. To resolve this, the authors propose the LLM-as-a-Tutor framework, which treats prompt adaptation as a policy-aware process. Leveraging a large language model as both examiner and tutor, the method identifies non-challenging prompts through pairwise comparisons and monotonically increases task difficulty by appending atomic constraints—ensuring training signals remain calibrated to the policy’s current proficiency. Notably, this approach requires no external scheduling mechanism and significantly outperforms policy-agnostic baselines and existing adaptive methods across three complex instruction-following benchmarks.

0 citationsRead paper

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Jun 30, 2026

This work addresses a critical limitation in existing post-training methods for medical multimodal large language models, which overly prioritize final answer correctness while neglecting the optimization of intermediate reasoning steps, thereby suffering from cascading failures triggered by early errors. To mitigate this, we propose Medical Reasoning-aware Policy Optimization (MRPO), the first reinforcement learning framework tailored for medical multimodal reasoning that incorporates step-aware feedback. MRPO introduces step-level process rewards, applying exponentially stronger penalties to ineffective early reasoning even when the final answer is incorrect, thus effectively interrupting error propagation. Experimental results demonstrate that MRPO consistently outperforms standard GRPO and state-of-the-art reinforcement learning baselines across three multimodal LLM backbones. Notably, on Qwen3-VL-8B-Instruct, it surpasses HuatuoGPT-Vision-34B by 2.79 points and reduces early reasoning failure rates from 64.0% to 13.0%.

0 citationsRead paper

Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents

Jun 25, 2026

This study addresses the lack of effective evaluation for breadth-first search capabilities—specifically, exhaustive enumeration of members within a closed set and their structured attributes—in non-English contexts. It presents the first Korean-language benchmark for breadth-oriented agent assessment, leveraging an automated synthesis and validation pipeline to generate tasks that require agents to fully enumerate members of a given parent entity and populate their structured attribute tables. The work introduces a novel structured difficulty modulation mechanism, controlling table width and two-dimensional composite keys, alongside a unified normalized matcher and a multidimensional scoring framework (Item-, Column-, and Row-F1). Experiments reveal that while current agents achieve strong member identification (Item-F1: 92.8), they struggle significantly with complete row completion (Row-F1: 53.7), with performance degrading as task difficulty increases, thereby underscoring the benchmark’s necessity and challenge.

0 citationsRead paper

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

Jun 20, 2026

This work addresses the critical issue that existing biomedical agents often generate responses inconsistent with cited literature, a flaw overlooked by conventional evaluations reliant on fixed-answer benchmarks. To tackle this, the authors introduce the first retrieval-based agent evaluation benchmark for open-ended biomedical questions, comprising 12,553 unresolved queries spanning 12 domains. Agents must perform multi-turn tool-augmented reasoning and autonomously decide when to abstain from answering, enabling assessment of both factual faithfulness and refusal capability. Key innovations include an open-question paradigm without predefined answers, validation of question openness via subsequent real-world literature, objective difficulty calibration based on reference model failure rates, and a frozen checklist that significantly improves annotation consistency (Spearman correlation rising from 0.35 to 0.82). Experiments reveal that even state-of-the-art agents achieve only 29%–60% success on the hardest subset and exhibit pronounced tool-use degradation under high difficulty.

0 citationsRead paper