The Veto Variable: Human Override as a Goal-Independent Cost Term
本文探讨了AI系统中人类监督作为独立成本项的问题,提出了一种新的方法来量化这种成本,并讨论了如何保持人类对AI的控制权。
本文探讨了AI系统中人类监督作为独立成本项的问题,提出了一种新的方法来量化这种成本,并讨论了如何保持人类对AI的控制权。
This study addresses whether multidimensional predictors of ant mortality risk are shared across traits. Through paired field and laboratory survival experiments combined with Cox regression and AIC-based model selection, we demonstrate that body size predicts only lifespan duration, while senescence trajectories are driven by circadian rhythms. Furthermore, thermal vulnerability exhibits phylogenetic specificity and plateaus above 20°C. These findings reveal a decoupling mechanism among distinct mortality risk dimensions, challenging the assumption that a single metric can uniformly predict survival. By disentangling these factors, this work provides novel insights into insect life-history evolution and establishes a foundation for constructing multidimensional risk assessment frameworks in ecological and evolutionary research.
This study addresses the challenge of isolating linguistic variables in evaluating large language models, which has obscured whether non-English comprehension genuinely lags behind English performance. The authors introduce and quantify the “Cross-Linguistic Comprehension Gap” (CLCG) through a rigorously controlled parallel evaluation framework where content is held constant across languages. Leveraging ParallelQA-18—a human-translated dataset spanning 18 languages—and combining token-level F1 micro-averaging, passage clustering with bootstrapping, and blind human preference trials, they find an overall CLCG of 0.078, corresponding to an approximate 17% performance drop. Crucially, CLCG exhibits a significant negative correlation with language resource availability: responses in high-resource languages are consistently preferred by human evaluators, revealing a systematic overestimation of model capabilities for low-resource language users under current English-centric evaluation paradigms.
This study addresses the lack of a unified evaluation benchmark and insufficient attention to temporal interpretability and calibration in existing dropout prediction research within learning analytics. The authors construct the first multidimensional benchmark tailored for survival analysis, systematically comparing diverse models—including random survival forests, piecewise exponential additive models, parametric survival models, and neural survival models—under both dynamic weekly-granularity and continuous-time representations. Leveraging person-period data formatting and refit-free bootstrapping, they conduct a comprehensive assessment through a four-dimensional framework encompassing predictive performance, ablation, interpretability, and calibration. Results reveal that temporal behavioral features dominate predictive signals, whereas static background factors exert limited influence; random survival forests perform best under continuous-time settings, while piecewise exponential models show slight advantages in dynamic settings. Notably, models with high discriminative ability generally exhibit strong calibration.
This study proposes an integrated framework combining time-series modeling with counterfactual policy simulation to predict weekly-granularity dropout risk among higher education students and evaluate intervention efficacy. Leveraging learning management system logs and administrative withdrawal records, the authors construct a person-period model using discrete-time survival analysis and penalized class-balanced logistic regression, achieving a test-set AUC of 0.8405. A counterfactual policy layer incorporating trigger mechanisms and scheduling contracts enables structured scenario comparisons. Bootstrap subgroup analyses reveal that only shock-type interventions significantly improve student survival rates (ΔS = 0.0819), whereas mechanism-aware interventions exhibit negative effects. Although gender-based survival gaps remain directionally consistent, their magnitude is minimal, highlighting substantial heterogeneity in intervention effectiveness across subpopulations.
本文探讨了AI系统中人类监督作为独立成本项的问题,提出了一种新的方法来量化这种成本,并讨论了如何保持人类对AI的控制权。
This study addresses whether multidimensional predictors of ant mortality risk are shared across traits. Through paired field and laboratory survival experiments combined with Cox regression and AIC-based model selection, we demonstrate that body size predicts only lifespan duration, while senescence trajectories are driven by circadian rhythms. Furthermore, thermal vulnerability exhibits phylogenetic specificity and plateaus above 20°C. These findings reveal a decoupling mechanism among distinct mortality risk dimensions, challenging the assumption that a single metric can uniformly predict survival. By disentangling these factors, this work provides novel insights into insect life-history evolution and establishes a foundation for constructing multidimensional risk assessment frameworks in ecological and evolutionary research.
This study addresses the challenge of isolating linguistic variables in evaluating large language models, which has obscured whether non-English comprehension genuinely lags behind English performance. The authors introduce and quantify the “Cross-Linguistic Comprehension Gap” (CLCG) through a rigorously controlled parallel evaluation framework where content is held constant across languages. Leveraging ParallelQA-18—a human-translated dataset spanning 18 languages—and combining token-level F1 micro-averaging, passage clustering with bootstrapping, and blind human preference trials, they find an overall CLCG of 0.078, corresponding to an approximate 17% performance drop. Crucially, CLCG exhibits a significant negative correlation with language resource availability: responses in high-resource languages are consistently preferred by human evaluators, revealing a systematic overestimation of model capabilities for low-resource language users under current English-centric evaluation paradigms.
This study addresses the lack of a unified evaluation benchmark and insufficient attention to temporal interpretability and calibration in existing dropout prediction research within learning analytics. The authors construct the first multidimensional benchmark tailored for survival analysis, systematically comparing diverse models—including random survival forests, piecewise exponential additive models, parametric survival models, and neural survival models—under both dynamic weekly-granularity and continuous-time representations. Leveraging person-period data formatting and refit-free bootstrapping, they conduct a comprehensive assessment through a four-dimensional framework encompassing predictive performance, ablation, interpretability, and calibration. Results reveal that temporal behavioral features dominate predictive signals, whereas static background factors exert limited influence; random survival forests perform best under continuous-time settings, while piecewise exponential models show slight advantages in dynamic settings. Notably, models with high discriminative ability generally exhibit strong calibration.
This study proposes an integrated framework combining time-series modeling with counterfactual policy simulation to predict weekly-granularity dropout risk among higher education students and evaluate intervention efficacy. Leveraging learning management system logs and administrative withdrawal records, the authors construct a person-period model using discrete-time survival analysis and penalized class-balanced logistic regression, achieving a test-set AUC of 0.8405. A counterfactual policy layer incorporating trigger mechanisms and scheduling contracts enables structured scenario comparisons. Bootstrap subgroup analyses reveal that only shock-type interventions significantly improve student survival rates (ΔS = 0.0819), whereas mechanism-aware interventions exhibit negative effects. Although gender-based survival gaps remain directionally consistent, their magnitude is minimal, highlighting substantial heterogeneity in intervention effectiveness across subpopulations.