TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages
为了解决低资源语言Kinyarwanda缺乏表格表示学习资源的问题,通过引入新的预训练目标和结构化嵌入方法,提升了该语言在表格问答任务上的表现。
为了解决低资源语言Kinyarwanda缺乏表格表示学习资源的问题,通过引入新的预训练目标和结构化嵌入方法,提升了该语言在表格问答任务上的表现。
为了解决基尼扬达语在现有嵌入模型中的不足,研究通过四阶段课程训练和多种负例排序损失方法构建了KinyaEmbed模型。
This study addresses the residual climate hedging valuation adjustment (HVA) that persists for trading desks despite existing hedging and overlay strategies—a gap not inferable from standalone stress losses. To tackle this, the authors propose an optimized modeling paradigm that pairs simulations of climate-activated and baseline worlds, reformulating hedge instrument discovery as an optimization objective aimed at minimizing residual HVA. The approach integrates an analytical linear-Gaussian Riccati solution with Climate-Dyna model-based reinforcement learning for nonlinear refinement, augmented by a gating mechanism to regulate policy updates. Empirical evaluation in a semi-synthetic EU ETS environment demonstrates that the method reduces average climate costs from 0.906 to 0.831—approaching the theoretical lower bound of 0.821—while achieving a 93% reduction in regret using only one-quarter of the typical trajectory count and capturing 60.7% of the theoretical gain within just 25 target transitions.
This work addresses the computational expense, noise sensitivity, and iterative McKean–Vlasov fixed-point solving inherent in calibrating local stochastic volatility (LSV) models by introducing the first end-to-end neural operator framework. The proposed method jointly outputs arbitrage-free implied volatility surfaces, Dupire local volatilities, LSV leverage functions, and projection-consistent conditional moments. By shifting fixed-point computation to an offline phase, online calibration requires only a single operator evaluation. The approach innovatively incorporates a division-free Dupire residual and a quotient-form Fokker–Planck equation, combining DeepONet with Fourier neural operators to enforce joint constraints—on market quote fitting, static no-arbitrage conditions, dynamic consistency, and projection identities—in log-implied-variance coordinates. Experiments demonstrate a reduction in calibration latency from 98.5 ms to 0.6 ms, a 36% decrease in local volatility RMSE, 7–16% lower leverage function RMSE, and forward-start and cliquet option pricing errors of merely 0.1% and 0.2%, respectively.
This work addresses the challenge of diagnosing tool-use failures in AI agents operating within high-stakes enterprise workflows, where early errors can cascade into significant downstream risks. To enhance the internal observability of agent behavior, the authors propose a mechanistic interpretability approach that predicts both tool invocation requirements and action impacts by inspecting the agent’s internal states prior to execution. The method innovatively integrates sparse autoencoders (SAEs) with linear probes, complemented by feature ablation analysis, to identify critical internal layers and features associated with tool usage. Evaluated on the NVIDIA Nemotron dataset using GPT-OSS 20B and Gemma 3 27B models, the approach successfully isolates interpretable mechanisms predictive of tool-related behavior, thereby enabling early error detection and supporting robust risk monitoring of autonomous agents.
为了解决低资源语言Kinyarwanda缺乏表格表示学习资源的问题,通过引入新的预训练目标和结构化嵌入方法,提升了该语言在表格问答任务上的表现。
为了解决基尼扬达语在现有嵌入模型中的不足,研究通过四阶段课程训练和多种负例排序损失方法构建了KinyaEmbed模型。
This study addresses the residual climate hedging valuation adjustment (HVA) that persists for trading desks despite existing hedging and overlay strategies—a gap not inferable from standalone stress losses. To tackle this, the authors propose an optimized modeling paradigm that pairs simulations of climate-activated and baseline worlds, reformulating hedge instrument discovery as an optimization objective aimed at minimizing residual HVA. The approach integrates an analytical linear-Gaussian Riccati solution with Climate-Dyna model-based reinforcement learning for nonlinear refinement, augmented by a gating mechanism to regulate policy updates. Empirical evaluation in a semi-synthetic EU ETS environment demonstrates that the method reduces average climate costs from 0.906 to 0.831—approaching the theoretical lower bound of 0.821—while achieving a 93% reduction in regret using only one-quarter of the typical trajectory count and capturing 60.7% of the theoretical gain within just 25 target transitions.
This work addresses the computational expense, noise sensitivity, and iterative McKean–Vlasov fixed-point solving inherent in calibrating local stochastic volatility (LSV) models by introducing the first end-to-end neural operator framework. The proposed method jointly outputs arbitrage-free implied volatility surfaces, Dupire local volatilities, LSV leverage functions, and projection-consistent conditional moments. By shifting fixed-point computation to an offline phase, online calibration requires only a single operator evaluation. The approach innovatively incorporates a division-free Dupire residual and a quotient-form Fokker–Planck equation, combining DeepONet with Fourier neural operators to enforce joint constraints—on market quote fitting, static no-arbitrage conditions, dynamic consistency, and projection identities—in log-implied-variance coordinates. Experiments demonstrate a reduction in calibration latency from 98.5 ms to 0.6 ms, a 36% decrease in local volatility RMSE, 7–16% lower leverage function RMSE, and forward-start and cliquet option pricing errors of merely 0.1% and 0.2%, respectively.
This work addresses the challenge of diagnosing tool-use failures in AI agents operating within high-stakes enterprise workflows, where early errors can cascade into significant downstream risks. To enhance the internal observability of agent behavior, the authors propose a mechanistic interpretability approach that predicts both tool invocation requirements and action impacts by inspecting the agent’s internal states prior to execution. The method innovatively integrates sparse autoencoders (SAEs) with linear probes, complemented by feature ablation analysis, to identify critical internal layers and features associated with tool usage. Evaluated on the NVIDIA Nemotron dataset using GPT-OSS 20B and Gemma 3 27B models, the approach successfully isolates interpretable mechanisms predictive of tool-related behavior, thereby enabling early error detection and supporting robust risk monitoring of autonomous agents.