[优达机器学习入门]课程4：决策树

最新推荐文章于 2021-05-28 14:42:41 发布

热心市民Daisy

最新推荐文章于 2021-05-28 14:42:41 发布

阅读量256

点赞数

分类专栏：机器学习

本文链接：https://blog.csdn.net/Daisy_fight/article/details/80638824

版权

机器学习专栏收录该内容

13 篇文章 0 订阅

订阅专栏

from sklearn import tree
clf = tree.DecisionTreeClassifier()
clf.fit(features_train, labels_train)
pred = clf.predict(features_test)
accuracy = clf.score(features_test, labels_test)

min_samples_split :
The minimum number of samples required to split an internal node:
当min_samples_split设为50时，可以一定程度减少过拟合

##决策树编码

def classify(features_train, labels_train):
    from sklearn import tree
    clf = tree.DecisionTreeClassifier()
    clf = clf.fit(features_train, labels_train)
    return clf

##决策树准确性

from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier()
clf.fit(features_train, labels_train)
pred = clf.predict(features_test)
acc = clf.score(features_test, labels_test)

##决策树准确性

from sklearn.ensemble import RandomForestClassifier
##min_samples_split=2
clf = RandomForestClassifier(min_samples_split=2)
clf.fit(features_train, labels_train)
pred = clf.predict(features_test)
acc_min_samples_split_2 = clf.score(features_test, labels_test)
##min_samples_split=50
clf = RandomForestClassifier(min_samples_split=50)
clf.fit(features_train, labels_train)
pred = clf.predict(features_test)
acc_min_samples_split_50 = clf.score(features_test, labels_test)

##熵公式

##信息增益

##第一个邮件 DT：准确率

from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier(min_samples_split = 40)
clf.fit(features_train, labels_train)
pred = clf.predict(features_test)
accuracy = clf.score(features_test, labels_test)
print(accuracy)

##通过特征选择加速

print(len(features_train[0]))

##更改特征数量

#email_preprocess.py
selector = SelectPercentile(f_classif, percentile=1)  #percentile=1即1%可用特征
#dt_author_id.py
print(len(features_train[0]))

热心市民Daisy

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
[优达机器学习入门]课程4：决策树

决策树编码#classifyDT.pydef classify(features_train, labels_train): ### your code goes here--should return a trained decision tree classifer from sklearn import tree clf = tree.DecisionTr...
复制链接

扫一扫