sklearn的feature_importances_含义是什么？

最新推荐文章于 2024-07-09 20:51:47 发布

菜鸡的自我拯救

最新推荐文章于 2024-07-09 20:51:47 发布

阅读量1.6w

点赞数 4

分类专栏：机器学习/深度学习算法

本文链接：https://blog.csdn.net/weixin_37659245/article/details/100171656

版权

本文介绍了scikit-learn库中特征重要性`feature_importances_`的含义，它基于特征对节点纯度的总减少（平均在所有树上）。另外，还提到了另一种评估方法——平均准确率下降，这种方法直接衡量特征对模型准确性的影响。文中包含了两种方法的代码示例，并讨论了它们之间的差异。

摘要由CSDN通过智能技术生成

增添：这篇博文讲的也特别好

正文：

Sk-learn作者的答案：

There are indeed several ways to get feature “importances”. As often, there is no strict consensus about what this word means.
In scikit-learn, we implement the importance as described in [1] (often cited, but unfortunately rarely read…). It is sometimes called “gini importance” or “mean decrease impurity” and is defined as the total decrease in node impurity (weighted by the probability of reaching that node (which is approximated by the proportion of samples reaching that node)) averaged over all trees of the ensemble.
In the literature or in some other packages, you can also find feature importances implemented as the “mean decrease accuracy”. Basically, the idea is to measure the decrease in accuracy on OOB data when you randomly permute the values for that feature. If the decrease is low, then the feature is not important, and vice-versa.
(Note that both algorithms are available in the randomForest R package.)
[1]: Breiman, Friedman, “Classification and regression trees”, 1984.

所以，共有两种比较流行的特征重要性评估方法：
这一篇文中有两种方法的代码：Selecting good features – Part III: random forests

Mean decrease impurity
这个方法的原理其实就是Tree-Model进行分类、回归的原理：特征越重要，对节点的纯度增加的效果越好。而纯度的判别标准有很多，如GINI、信息熵、信息熵增益。
这也是Sklearn的feature_importances_的意义。
代码：

from sklearn.datasets import load_boston
from sklearn.ensemble imp

最低0.47元/天解锁文章

菜鸡的自我拯救

关注

4
点赞
踩
29

收藏

觉得还不错? 一键收藏
1
评论
复制链接

分享到 QQ

分享到新浪微博

扫一扫

专栏目录