coursera机器学习公开课笔记:11 machine-learning-system-design

最新推荐文章于 2022-06-03 09:32:53 发布

VIP文章 SnailDove

最新推荐文章于 2022-06-03 09:32:53 发布

阅读量366

点赞数

分类专栏：机器学习吴恩达《机器学习》公开课的完整笔记文章标签：机器学习

本文链接：https://blog.csdn.net/you1314520me/article/details/80101844

版权

Note

This personal note is written after studying the opening course on the coursera website, Machine Learning by Andrew NG . And images, audios of this note all comes from the opening course.

- Note
Table of Contents
01_building-a-spam-classifier
- Prioritizing What to Work On
- Error Analysis
02_handling-skewed-data
- 01_error-metrics-for-skewed-classes
- 02_trading-off-precision-and-recall

01_building-a-spam-classifier

Prioritizing What to Work On

example_of_spam-email_and_non-spam-email

System Design Example:

Given a data set of emails, we could construct a vector for each email. Each entry in this vector represents a word. The vector normally contains 10,000 to 50,000 entries gathered by finding the most frequently used words in our data set. If a word is to be found in the email, we would assign its respective entry a 1, else if it is not found, that entry would be a 0. Once we have all our x vectors ready, we train our algorithm and finally, we could use it to classify if an email is a spam or not.

Building_a_spam_classifier

So how could you spend your time to improve the accuracy of this classifier?

Collect lots of data (for example “honeypot” project but doesn’t always work)
Develop sophisticated features (for example: using email header data in spam emails)
Develop sophis3cated features for message body (for example: should“discount” and “discounts” be treated as the same word? How about “deal” and “Dealer”? Features about punctuation)?
Develop algorithms to process your input in different ways (recognizing misspellings in spam, for example, med1cine, m0rtgage, w4tches).

It is difficult to tell which of the options will be most helpful.

Error Analysis

The recommended approach to solving machine learning problems is to:

Start with a simple algorithm, implement it quickly, and test it early on your cross validation data.
Plot learning curves to decide if more data, more features, etc. are likely to help.
Manually examine the errors on examples in the cross validation set and try to spot a trend where most of the errors were made.

For example, assume that we have 500 emails and our algorithm misclassifies a 100 of them. We could manually analyze the 100 emails and categorize them based on what type of emails they are. We could then try to come up with new cues and features that would help us classify these 100 emails correctly. Hence, if most of our misclassified emails are those which try to steal passwords, then we could find some features that are particular to those emails and add them to our model.

error_analysis

We could also see how classifying each word according to its root changes our error rate:

The_importance_of_numerical_evaluation

It is very important to get error results as a single, numerical val

最低0.47元/天解锁文章

SnailDove

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
coursera机器学习公开课笔记:11 machine-learning-system-design

NoteThis personal note is written after studying the opening course on the coursera website, Machine Learning by Andrew NG . And images, audios of this note all comes from the opening course. Tabl...
复制链接

扫一扫