Machine Learning in Action
p52 Listing 2.2 :Text record to Numpy parsing code
使用python自带的stream和str类实现
def file2matrix(filename):
fr = open(filename)
numberOfLines = len(fr.readlines())
returnMat = np.zeros((numberOfLines,3))
classLabelVector = []
fr = open(filename)
index = 0
for line in fr.readlines():
line = line.strip()
listFromLine = line.split('\t')
returnMat[index,:] = listFromLine[0:3]
#把表示数值的字符串赋值给储存数值的np.array时会自动转换
classLabelVector.append(int(listFromLine[-1]))
index += 1
return returnMat,classLabelVector
(fun) file = open(filename)
Open file and return a corresponding stream.
https://docs.python.org/3/library/functions.html#open
(om) file.readlines()
Read and return a list of lines from the stream
返回的是一个1D list, 文件中的每行内容为list中的一个str
for line in file.readlines() 在py3中可以简写为 for line in file
https://docs.python.org/3/library/io.html?highlight=readlines#io.IOBase.readlines
相似的method有
(om) file.readline(size=-1)
Read and return one line from the stream. If size is specified, at most size bytes will be read.
https://docs.python.org/3/library/io.html?highlight=readlines#io.IOBase.readline
(om) str.strip([chars])
Return a copy of the string with the leading and trailing characters removed