二分类模型-分布式SPARK效果评估实现代码+混淆矩阵

最新推荐文章于 2023-02-07 20:51:45 发布

泰格数据

最新推荐文章于 2023-02-07 20:51:45 发布

阅读量1.5k

点赞数

分类专栏：模型评估算法机器学习

本文链接：https://blog.csdn.net/xiefu5hh/article/details/106135738

版权

最近在做一个平台级的项目，为了保证分布式的可扩展性，评估最终用sparkmlib进行模型的评估，sparkmlib里面封装好了二分类、
多分类、聚类的通用的评估指标，通用指标实现起来都比较简单。

关键点：
 val metrics=new BinaryClassificationMetrics(scoreAndLable,100)
  获取到预测列和标签列，并转化为RDD[double,double]。

BinaryClassificationMetrics第二个参数解释：这个一个分箱参数，可能你们预测值的不同值会有几百万个，这样会导致计算量巨大，计算的结果也巨大，所以引入分箱，对预测列进行分箱降采样后进行计算。
官方：param: scoreAndLabels an RDD of (score, label) pairs. param: numBins if greater than 0, then the curves (ROC curve, PR curve) computed internally will be down-sampled to this many "bins". If 0, no down-sampling will occur. This is useful because the curve contains a point for each distinct score in the input, and this could be as large as the input itself -- millions of points or more, when thousands may be entirely sufficient to summarize the curve. After down-sampling, the curves will instead be made of approximately numBins points instead. Points are made from bins of equal numbers of consecutive points. The size of each bin is floor(scoreAndLabels.count() / numBins), which means the resulting number of bins may not exactly equal numBins. The last bin in each partition may be smaller as a result, meaning

最低0.47元/天解锁文章

泰格数据

关注

0
点赞
踩
3

收藏

觉得还不错? 一键收藏
0
评论
二分类模型-分布式SPARK效果评估实现代码+混淆矩阵

最近在做一个平台级的项目，为了保证分布式的可扩展性，评估最终用sparkmlib进行模型的评估，sparkmlib里面封装好了二分类、多分类、聚类的通用的评估指标，通用指标实现起来都比较简单。关键点： val metrics=new BinaryClassificationMetrics(scoreAndLable,100) 获取到预测列和标签列，并转化为RDD[double,double]。BinaryClassificationMetrics第二个参数解释：这个一个分箱参数，可能你...
复制链接

扫一扫