Python机器学习中的异常数据剔除

最新推荐文章于 2024-06-11 10:24:50 发布

civilpy

最新推荐文章于 2024-06-11 10:24:50 发布

阅读量642

点赞数 10

文章标签： python 机器学习开发语言

本文链接：https://blog.csdn.net/baidu_22713341/article/details/138437625

版权

机器学习中的异常数据剔除

在机器学习中，异常数据可能会对模型的训练和预测产生负面影响。为了提高模型的性能，我们需要在数据预处理阶段剔除异常数据。以下是使用Python剔除异常数据的一些方法：

1. 使用箱线图（Boxplot）进行异常值检测

箱线图是一种常用的数据可视化方法，可以帮助我们识别异常值。以下是使用matplotlib库绘制箱线图的示例：

import numpy as np
import matplotlib.pyplot as plt

data = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 100])
plt.boxplot(data)
plt.show()

2. 使用Z-score进行异常值检测

Z-score是一种常用的异常值检测方法，它计算数据点与均值之间的标准差数。以下是使用scipy库计算Z-score的示例：

from scipy import stats

data = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 100])
z_scores = np.abs(stats.zscore(data))
threshold = 2
outliers = np.where(z_scores > threshold)
print("异常值索引：", outliers)
print("异常值：", data[outliers])

3. 使用IQR（四分位距）进行异常值检测

IQR是一种基于分位数的异常值检测方法。以下是使用numpy库计算IQR的示例：

import numpy as np

data = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 100])
q1 = np.percentile(data, 25)
q3 = np.percentile(data, 75)
iqr = q3 - q1
lower_bound = q1 - 1.5 * iqr
upper_bound = q3 + 1.5 * iqr
outliers = np.where((data< lower_bound) | (data > upper_bound))
print("异常值索引：", outliers)
print("异常值：", data[outliers])

4. 使用DBSCAN（密度聚类）进行异常值检测

DBSCAN是一种基于密度的聚类算法，可以用来检测异常值。以下是使用sklearn库进行DBSCAN的示例：

from sklearn.cluster import DBSCAN

data = np.array([[1], [2], [3], [4], [5], [6], [7], [8], [9], [10], [11], [12], [13], [14], [15], [100]])
dbscan = DBSCAN(eps=2, min_samples=2)
clusters = dbscan.fit_predict(data)
outliers = np.where(clusters == -1)
print("异常值索引：", outliers)
print("异常值：", data[outliers])

5. 使用隔离森林（Isolation Forest）进行异常值检测

隔离森林是一种基于树结构的异常值检测算法。以下是使用sklearn库进行隔离森林的示例：

from sklearn.ensemble import IsolationForest

data = np.array([[1], [2], [3], [4], [5], [6], [7], [8], [9], [10], [11], [12], [13], [14], [15], [100]])
isolation_forest = IsolationForest(contamination=0.1)
outliers = isolation_forest.fit_predict(data)
outlier_index = np.where(outliers == -1)
print("异常值索引：", outlier_index)
print("异常值：", data[outlier_index])