VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
为解决SVG生成、编辑等缺乏专业级基准问题,提出VectorGym,采用多任务强化学习方法优化,并提供人类标注数据和评估指标。
为解决SVG生成、编辑等缺乏专业级基准问题,提出VectorGym,采用多任务强化学习方法优化,并提供人类标注数据和评估指标。
本文针对LLM生成的论文复现代码科学不忠实的问题,提出SA-Bench方法,通过定义和评估语义对齐单元来诊断和量化代码与原论文之间的偏差。
研究解决了论文到代码复现难题,通过引入ReproAgent,一个基于持续实现合同的四阶段管道,利用实现需求和参考证据两个通道来生成和修复代码。
This study addresses the dependence of regret bounds on peak backlog in capacity-constrained multi-armed bandits with delayed feedback. We propose a scheduler-side conditional energy interface that decouples rate adaptation from disturbance filtering to optimize delay complexity. By revealing how temporal geometry influences regret, we demonstrate that minimax regret can differ polynomially under identical delay statistics, thereby overcoming the limitations of aggregated metrics. Leveraging semi-transparent oracles and strong convexity analysis, we derive a parameter-free regret bound with a delay term of o(√(e_c d_tot)) and establish a lower bound for capacity-scarce regimes. These contributions significantly refine the theoretical granularity of regret analysis in delayed feedback settings beyond conventional aggregate measures.
This study addresses the challenge of stance detection in real-world social media, where targets are often undefined and dynamically evolving, rendering traditional methods ineffective. The work introduces, for the first time, an open-domain zero-shot stance detection task that leverages large language models (LLMs) to dynamically generate stance targets and adapt to multiple targets without requiring prior target knowledge. Key contributions include the construction of the first Chinese social media stance dataset with multidimensional evaluation metrics and the design of both integrated and two-stage fine-tuning frameworks. Experimental results demonstrate that the two-stage fine-tuned Qwen2.5-7B achieves a composite score of 66.99% in target identification, while the integrated fine-tuned DeepSeek-R1-Distill-Qwen-7B attains an F1 score of 79.26% in stance detection.
本文针对LLM生成的论文复现代码科学不忠实的问题,提出SA-Bench方法,通过定义和评估语义对齐单元来诊断和量化代码与原论文之间的偏差。
研究解决了论文到代码复现难题,通过引入ReproAgent,一个基于持续实现合同的四阶段管道,利用实现需求和参考证据两个通道来生成和修复代码。
This study addresses the dependence of regret bounds on peak backlog in capacity-constrained multi-armed bandits with delayed feedback. We propose a scheduler-side conditional energy interface that decouples rate adaptation from disturbance filtering to optimize delay complexity. By revealing how temporal geometry influences regret, we demonstrate that minimax regret can differ polynomially under identical delay statistics, thereby overcoming the limitations of aggregated metrics. Leveraging semi-transparent oracles and strong convexity analysis, we derive a parameter-free regret bound with a delay term of o(√(e_c d_tot)) and establish a lower bound for capacity-scarce regimes. These contributions significantly refine the theoretical granularity of regret analysis in delayed feedback settings beyond conventional aggregate measures.
为解决SVG生成、编辑等缺乏专业级基准问题,提出VectorGym,采用多任务强化学习方法优化,并提供人类标注数据和评估指标。
This study addresses the challenge of stance detection in real-world social media, where targets are often undefined and dynamically evolving, rendering traditional methods ineffective. The work introduces, for the first time, an open-domain zero-shot stance detection task that leverages large language models (LLMs) to dynamically generate stance targets and adapt to multiple targets without requiring prior target knowledge. Key contributions include the construction of the first Chinese social media stance dataset with multidimensional evaluation metrics and the design of both integrated and two-stage fine-tuning frameworks. Experimental results demonstrate that the two-stage fine-tuned Qwen2.5-7B achieves a composite score of 66.99% in target identification, while the integrated fine-tuned DeepSeek-R1-Distill-Qwen-7B attains an F1 score of 79.26% in stance detection.