2020李宏毅机器学习课程作业——Homework2：classification（Logistic Regression）

最新推荐文章于 2022-10-17 10:46:27 发布

梅菜扣肉鱼丸粗面

最新推荐文章于 2022-10-17 10:46:27 发布

阅读量2.1k

点赞数 1

分类专栏：机器学习文章标签：机器学习

本文链接：https://blog.csdn.net/qushuo123/article/details/106987801

版权

一、作业获取途径

课程网址：http://speech.ee.ntu.edu.tw/~tlkagk/courses_ML20.html

B站视频地址：https://www.bilibili.com/video/BV1JE411g7XF?from=search&seid=18330864429491522852

如果不能访问课程网址，作业的压缩包如下，链接：https://pan.baidu.com/s/1PZZnWZKZONDCGTMznKINEg
提取码：5of8

案例提供的是在Colab上运行的，不能使用的同学可以直接使用其他Jupyter Notebook即可。

二、作业说明

本次作业给了两个模型，分别为判别式模型和生成式模型，具体两者的推到，可以看我上一篇文章。

逻辑回归的两种模型——判别式和生成式

2.1判别式模型

下载好的问价夹里包括所有需要的文件，打开ipython文件，运行如下代码，将数据集下载，对每个特征进行标准化，并将数据集切分为训练集和发展集（交叉验证集）。

import numpy as np

np.random.seed(0)
X_train_fpath = './data/X_train'
Y_train_fpath = './data/Y_train'
X_test_fpath = './data/X_test'
output_fpath = './output_{}.csv'

# Parse csv files to numpy array
with open(X_train_fpath) as f:
    next(f)
    X_train = np.array([line.strip('\n').split(',')[1:] for line in f], dtype = float)
with open(Y_train_fpath) as f:
    next(f)
    Y_train = np.array([line.strip('\n').split(',')[1] for line in f], dtype = float)
with open(X_test_fpath) as f:
    next(f)
    X_test = np.array([line.strip('\n').split(',')[1:] for line in f], dtype = float)

def _normalize(X, train = True, specified_column = None, X_mean = None, X_std = None):
    # This function normalizes specific columns of X.
    # The mean and standard variance of training data will be reused when processing testing data.
    #
    # Arguments:
    #     X: data to be processed
    #     train: 'True' when processing training data, 'False' for testing data
    #     specific_column: indexes of the columns that will be normalized. If 'None', all columns
    #         will be normalized.
    #     X_mean: mean value of training data, used when train = 'False'
    #     X_std: standard deviation of training data, used when train = 'False'
    # Outputs:
    #     X: normalized data
    #     X_mean: computed mean value of training data
    #     X_std: computed standard deviation of training data

    if specified_column == None:
        specified_column = np.arange(X.shape[1])
    if train:
        X_mean = np.mean(X[:, specified_column] ,0).reshape(1, -1)
        X_std  = np.std(X[:, specified_column], 0).reshape(1, -1)

    X[:,specified_column] = (X[:, specified_column] - X_mean) / (X_std + 1e-8)
     
    return X, X_mean, X_std

def _train_dev_split(X, Y, dev_ratio = 0.25):
    # This function spilts data into training set and development set.
    train_size = int(len(X) * (1 - dev_ratio))
    return X[:train_size], Y[:train_size], X[train_size:], Y[train_size:]#没有逗号，只有行分割，包含所有列

# Normalize training and testing data
X_train, X_mean, X_std = _normalize(X_train, train = True)
X_test, _, _= _normalize(X_test, train = False, specified_column = None, X_mean = X_mean, X_std = X_std)
    
# Split data into training set and development set
dev_ratio = 0.1
X_train, Y_train, X_dev, Y_dev = _train_dev_split(X_train, Y_train, dev_ratio = dev_ratio)

train_size = X_train.shape[0]
dev_size = X_dev.shape[0]
test_size = X_test.shape[0]
data_dim = X_train.shape[1]
print('Size of training set: {}'.format(train_size))
print('Size of development set: {}'.format(dev_size))
print('Size of testing set: {}'.format(test_size))
print('Dimension of data: {}'.format(data_dim))

输出结果为：

接下来把用到的一些操作定义为函数，包括打乱数据集（ _shuffle(X, Y)）、sigmoid函数（_sigmoid(z)）、激活函数，也就是将wx+b输入sigmoid函数（_f(X, w, b)）、预测函数，也就是计算预测值的函数（_predict(X, w, b)）、计算准确率的函数（）、_accuracy(Y_pred, Y_label)、交叉熵损失函数（_cross_entropy_loss(y_pred, Y_label)）和梯度函数，用来计算梯度以更新w和b（ _gradient(X, Y_label, w, b)）。

def _shuffle(X, Y):
    # This function shuffles two equal-length list/array, X and Y, together.
    randomize = np.arange(len(X))
    np.random.shuffle(randomize)
    return (X[randomize], Y[randomize])

def _sigmoid(z):
    # Sigmoid function can be used to calculate probability.
    # To avoid overflow, minimum/maximum output value is set.
    return np.clip(1 / (1.0 + np.exp(-z)), 1e-8, 1 - (1e-8))

def _f(X, w, b):
    # This is the logistic regression function, parameterized by w and b
    #
    # Arguements:
    #     X: input data, shape = [batch_size, data_dimension]
    #     w: weight vector, shape = [data_dimension, ]
    #     b: bias, scalar
    # Output:
    #     predicted probability of each row of X being positively labeled, shape = [batch_size, ]
    return _sigmoid(np.matmul(X, w) + b)

def _predict(X, w, b):
    # This function returns a truth value prediction for each row of X 
    # by rounding the result of logistic regression function.
    return np.round(_f(X, w, b)).astype(np.int)
    
def _accuracy(Y_pred, Y_label):
    # This function calculates prediction accuracy
    acc = 1 - np.mean(np.abs(Y_pred - Y_label))
    return acc
def _cross_entropy_loss(y_pred, Y_label):
    # This function computes the cross entropy.
    #
    # Arguements:
    #     y_pred: probabilistic predictions, float vector
    #

最低0.47元/天解锁文章

梅菜扣肉鱼丸粗面

关注

1
点赞
踩
21

收藏

觉得还不错? 一键收藏
7
评论
2020李宏毅机器学习课程作业——Homework2：classification（Logistic Regression）

一、作业获取途径课程网址：http://speech.ee.ntu.edu.tw/~tlkagk/courses_ML20.htmlB站视频地址：https://www.bilibili.com/video/BV1JE411g7XF?from=search&seid=18330864429491522852如果不能访问课程网址，作业的压缩包如下，链接：https://pan.baidu.com/s/1PZZnWZKZONDCGTMznKINEg提取码：5of8案例提供的是在Col.
复制链接

扫一扫

专栏目录