xgboost入门与实战（实战调参篇）

最新推荐文章于 2024-10-11 11:20:04 发布

我曾经被山河大海跨过

最新推荐文章于 2024-10-11 11:20:04 发布

阅读量6.8w

点赞数 37

本文链接：https://blog.csdn.net/sb19931201/article/details/52577592

版权

本文介绍了使用xgboost进行手写数字识别的实战过程，基于kaggle上的MNIST数据集。内容涵盖数据集介绍、xgboost模型的训练与调参、预测结果评价，并分享了不同参数配置下的运行时间和精度表现。

摘要由CSDN通过智能技术生成

xgboost入门与实战（实战调参篇）

前言

前面几篇博文都在学习原理知识，是时候上数据上模型跑一跑了。本文用的数据来自kaggle，相信搞机器学习的同学们都知道它，kaggle上有几个老题目一直开放，适合给新手练级，上面还有很多老司机的方案共享以及讨论，非常方便新手入门。这次用的数据是Classify handwritten digits using the famous MNIST data—手写数字识别，每个样本相当于一个图片像素矩阵（28x28）,每个像元就是一个特征啦。用这个数据的好处就是不用做特征工程了，对于上手模型很方便。
xgboost安装看这里：http://blog.csdn.net/sb19931201/article/details/52236020

拓展一下：XGBoost Plotting API以及GBDT组合特征实践

数据集

1.数据介绍：数据采用的是广泛应用于机器学习社区的MNIST数据集：

The data for this competition were taken from the MNIST dataset. The MNIST (“Modified National Institute of Standards and Technology”) dataset is a classic within the Machine Learning community that has been extensively studied. More detail about the dataset, including Machine Learning algorithms that have been tried on it and their levels of success, can be found at http://yann.lecun.com/exdb/mnist/index.html.

2.训练数据集（共42000个样本）：如下图所示，第一列是label标签，后面共28x28=784个像素特征，特征取值范围为0-255。

3.测试数据集（共28000条记录）：就是我们要预测的数据集了，没有标签值，通过train.csv训练出合适的模型，然后根据test.csv的特征数据集来预测这28000的类别（0-9的数字标签）。

4.结果数据样例：两列，一列为ID，一列为预测标签值。

xgboost模型调参、训练（python）

1.导入相关库，读取数据

import numpy as np
import pandas as pd
import xgboost as xgb
from sklearn.cross_validation import train_test_split

#记录程序运行时间
import time 
start_time = time.time()

#读入数据
train = pd.read_csv("Digit_Recognizer/train.csv")
tests = pd.read_csv("Digit_Recognizer/test.csv")