Learned Look-Ahead Splitting Rule for CART
提出了一种前瞻性的CART树构建方法,通过预测误差减少评估每个候选分裂点,并使用节点特征学习下游分裂值以提高计算效率。
提出了一种前瞻性的CART树构建方法,通过预测误差减少评估每个候选分裂点,并使用节点特征学习下游分裂值以提高计算效率。
本文通过引入StateSight基准测试,评估视觉-语言模型从单张图片重建潜在空间结构的能力,采用特定任务挑战现有模型,并分析了图像状态重建中的常见错误。
This work addresses the challenge of ensuring strict regulatory compliance for large language model agents in financial settings, where inherent non-determinism conflicts with the certainty demanded by regulators. To this end, the authors propose the Lean-Agent protocol, which automatically formalizes institutional compliance policies as theorems in Lean 4 and treats agent actions as conjectures requiring formal proof. An action is executed only if it is formally verified against predefined regulatory axioms within the Lean 4 proof assistant. This approach establishes, for the first time, a formal-verification-based compliance guardrail for AI systems with cryptographic-grade determinism, enabling interpretable compliance decisions at microsecond latency. The framework directly satisfies key regulatory mandates—including SEC Rule 15c3-5, FINRA Rule 3110, OCC Bulletin 2011-12, and CFPB requirements—and outlines a practical three-stage deployment pathway from shadow validation to production.
In sparse-reward environments, standard DQN suffers from unstable convergence and low learning efficiency due to its fixed learning rate and uniform update mechanism. This work proposes DISRC, a novel approach that integrates a biologically inspired surprise mechanism into DQN. DISRC employs a LayerNorm encoder to construct state representations and computes a discrepancy-based intrinsic surprise signal derived from a moving latent variable setpoint. This surprise signal is combined with the temporal difference (TD) error to dynamically modulate the intensity of Q-value updates, enabling expectation-violation-driven adaptive learning. Evaluated on MiniGrid-DoorKey, DISRC achieves a 33% faster training speed, lower reward variance, and higher AUC. On LavaCrossing, it attains the best-reported AUC (957.04) and final reward, demonstrating significantly improved exploration efficiency and training stability.
To address the high cost of manual household water consumption surveys in rapidly urbanizing regions, this paper proposes a low-cost alternative leveraging publicly available multi-source geospatial data. We develop a joint model integrating convolutional neural network (CNN) embeddings with ordinal classification, trained on satellite imagery, Google Street View (GSV) semantic segmentation maps, nighttime light intensity, and population density as covariates to predict household water consumption in Hubballi-Dharwad, India. This work is the first to synergistically combine GSV semantic segmentation with remote sensing features for water use estimation, demonstrating the efficacy of visual semantic information in socio-infrastructure modeling—particularly in capturing distributional extremes. The model achieves an accuracy of 0.55, comparable to conventional survey-based models (0.59), while substantially reducing data collection costs. It offers a scalable, reproducible technical pathway for water monitoring in resource-constrained settings.
提出了一种前瞻性的CART树构建方法,通过预测误差减少评估每个候选分裂点,并使用节点特征学习下游分裂值以提高计算效率。
本文通过引入StateSight基准测试,评估视觉-语言模型从单张图片重建潜在空间结构的能力,采用特定任务挑战现有模型,并分析了图像状态重建中的常见错误。
This work addresses the challenge of ensuring strict regulatory compliance for large language model agents in financial settings, where inherent non-determinism conflicts with the certainty demanded by regulators. To this end, the authors propose the Lean-Agent protocol, which automatically formalizes institutional compliance policies as theorems in Lean 4 and treats agent actions as conjectures requiring formal proof. An action is executed only if it is formally verified against predefined regulatory axioms within the Lean 4 proof assistant. This approach establishes, for the first time, a formal-verification-based compliance guardrail for AI systems with cryptographic-grade determinism, enabling interpretable compliance decisions at microsecond latency. The framework directly satisfies key regulatory mandates—including SEC Rule 15c3-5, FINRA Rule 3110, OCC Bulletin 2011-12, and CFPB requirements—and outlines a practical three-stage deployment pathway from shadow validation to production.
In sparse-reward environments, standard DQN suffers from unstable convergence and low learning efficiency due to its fixed learning rate and uniform update mechanism. This work proposes DISRC, a novel approach that integrates a biologically inspired surprise mechanism into DQN. DISRC employs a LayerNorm encoder to construct state representations and computes a discrepancy-based intrinsic surprise signal derived from a moving latent variable setpoint. This surprise signal is combined with the temporal difference (TD) error to dynamically modulate the intensity of Q-value updates, enabling expectation-violation-driven adaptive learning. Evaluated on MiniGrid-DoorKey, DISRC achieves a 33% faster training speed, lower reward variance, and higher AUC. On LavaCrossing, it attains the best-reported AUC (957.04) and final reward, demonstrating significantly improved exploration efficiency and training stability.
To address the high cost of manual household water consumption surveys in rapidly urbanizing regions, this paper proposes a low-cost alternative leveraging publicly available multi-source geospatial data. We develop a joint model integrating convolutional neural network (CNN) embeddings with ordinal classification, trained on satellite imagery, Google Street View (GSV) semantic segmentation maps, nighttime light intensity, and population density as covariates to predict household water consumption in Hubballi-Dharwad, India. This work is the first to synergistically combine GSV semantic segmentation with remote sensing features for water use estimation, demonstrating the efficacy of visual semantic information in socio-infrastructure modeling—particularly in capturing distributional extremes. The model achieves an accuracy of 0.55, comparable to conventional survey-based models (0.59), while substantially reducing data collection costs. It offers a scalable, reproducible technical pathway for water monitoring in resource-constrained settings.