机器学习实战——ch8.2 回归之预测乐高玩具价格

591984826

于 2016-08-24 09:37:38 发布

阅读量2.8k

点赞数 3

分类专栏：机器学习实战

本文链接：https://blog.csdn.net/u011629133/article/details/52296892

版权

这部分由于书上提供的Google的购物API已经关闭，所以只能在实验楼上完成了这个实验(这一次，我只是代码的搬运工)
这里写图片描述

完整的代码及注释：

#-*- coding: utf-8 -*-
from numpy import *
from BeautifulSoup import BeautifulSoup

# 从页面读取数据，生成retX和retY列表
def scrapePage(retX, retY, inFile, yr, numPce, origPrc):

    # 打开并读取HTML文件
    fr = open(inFile);
    soup = BeautifulSoup(fr.read())
    i=1

    # 根据HTML页面结构进行解析
    currentRow = soup.findAll('table', r="%d" % i)
    while(len(currentRow)!=0):
        currentRow = soup.findAll('table', r="%d" % i)
        title = currentRow[0].findAll('a')[1].text
        lwrTitle = title.lower()

        # 查找是否有全新标签
        if (lwrTitle.find('new') > -1) or (lwrTitle.find('nisb') > -1):
            newFlag = 1.0
        else:
            newFlag = 0.0

        # 查找是否已经标志出售，我们只收集已出售的数据
        soldUnicde = currentRow[0].findAll('t