Just add noise: Debiasing tree-based variable importance in mixed data

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了树基方法中变量重要性评分偏向连续预测变量的问题,通过向分类预测变量添加少量噪声的方法进行修正。
📝 Abstract
Variable importance scores from tree-based methods such as random forests favor continuous predictors over categorical ones. We present a theoretical analysis of this bias and propose a simple remedy: add a small amount of noise to each categorical predictor. The correction is demonstrated on a variety of simulated and real-world datasets and combined with integrated path stability selection to perform variable selection with false discovery control for mixed data.
Problem

Research questions and friction points this paper is trying to address.

variable importance
tree-based methods
bias
mixed data
Innovation

Methods, ideas, or system contributions that make the work stand out.

noise
variable importance
tree-based methods
bias correction
mixed data
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jiahe Li
Department of Statistical Science, Duke University
Omar Melikechi
Omar Melikechi
Duke University
statisticsbiostatisticsmachine learning