Institution profile

InvideoAI

Industry research
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

The Realignment Problem: When Right becomes Wrong in LLMs

Nov 04, 2025

Current large language models (LLMs) suffer from an “alignment–reality gap”: static alignment strategies fail to adapt to dynamically evolving societal norms and policies, resulting in value misalignment, poor robustness, and high maintenance overhead. To address this, we propose TRACE—a novel framework that formalizes realignment as a programmable policy-application problem. TRACE introduces an alignment impact score to quantitatively assess preference conflicts and enables selective preference reversal, discarding, or retention—balancing correction accuracy with model performance. Leveraging a hybrid optimization pipeline—integrating conflict evaluation, preference-data categorization and filtering, and selective retraining—TRACE achieves fine-grained, low-regret updates across diverse architectures (Qwen, Gemma, Llama). Experiments demonstrate that TRACE significantly improves compliance with complex, evolving policy requirements while preserving pre-existing general-purpose capabilities.

0 citationsRead paper

Policy Optimization Prefers The Path of Least Resistance

Oct 22, 2025

Prior work assumes chain-of-thought (CoT) reasoning must strictly adhere to a “reason-then-answer” format; however, this study investigates how policy optimization (PO) behaves under open-structured CoT—where reasoning and answer generation may interleave freely. Method: Through controlled experiments, reward decomposition analysis, and KL-regularized PO, we systematically examine strategy evolution across multiple models and algorithms. Contribution/Results: We demonstrate that PO inherently favors minimal-resistance reward acquisition paths, causing explicit reasoning to collapse into direct answer generation—even when complex CoT formats receive a 4× reward bonus. This is the first work to reveal PO’s intrinsic simplification bias under open CoT structures, exposing fundamental challenges in reward gaming for alignment training. Our findings provide both theoretical grounding and empirical evidence for designing robust reasoning-guidance mechanisms in large language models.

0 citationsRead paper
Recent publications

Latest Papers

The Realignment Problem: When Right becomes Wrong in LLMs

Nov 04, 2025

Current large language models (LLMs) suffer from an “alignment–reality gap”: static alignment strategies fail to adapt to dynamically evolving societal norms and policies, resulting in value misalignment, poor robustness, and high maintenance overhead. To address this, we propose TRACE—a novel framework that formalizes realignment as a programmable policy-application problem. TRACE introduces an alignment impact score to quantitatively assess preference conflicts and enables selective preference reversal, discarding, or retention—balancing correction accuracy with model performance. Leveraging a hybrid optimization pipeline—integrating conflict evaluation, preference-data categorization and filtering, and selective retraining—TRACE achieves fine-grained, low-regret updates across diverse architectures (Qwen, Gemma, Llama). Experiments demonstrate that TRACE significantly improves compliance with complex, evolving policy requirements while preserving pre-existing general-purpose capabilities.

0 citationsRead paper

Policy Optimization Prefers The Path of Least Resistance

Oct 22, 2025

Prior work assumes chain-of-thought (CoT) reasoning must strictly adhere to a “reason-then-answer” format; however, this study investigates how policy optimization (PO) behaves under open-structured CoT—where reasoning and answer generation may interleave freely. Method: Through controlled experiments, reward decomposition analysis, and KL-regularized PO, we systematically examine strategy evolution across multiple models and algorithms. Contribution/Results: We demonstrate that PO inherently favors minimal-resistance reward acquisition paths, causing explicit reasoning to collapse into direct answer generation—even when complex CoT formats receive a 4× reward bonus. This is the first work to reveal PO’s intrinsic simplification bias under open CoT structures, exposing fundamental challenges in reward gaming for alignment training. Our findings provide both theoretical grounding and empirical evidence for designing robust reasoning-guidance mechanisms in large language models.

0 citationsRead paper