python模块之HTMLParser解析出URL链接

最新推荐文章于 2021-06-04 04:41:29 发布

weixin_34375233

最新推荐文章于 2021-06-04 04:41:29 发布

阅读量114

点赞数

文章标签： python

# -*- coding: utf-8 -*-
#python 27
#xiaodeng
#python模块之HTMLParser解析出URL链接
#http://www.cnblogs.com/mfryf/p/3691563.html



from HTMLParser import HTMLParser
class MyHTMLParser(HTMLParser):   
    def __init__(self):   
        HTMLParser.__init__(self) #继承  
        self.links = []#links 链接
    
    def handle_starttag(self, tag, attrs):   
        #print "Encountered the beginning of a %s tag" % tag
        
        if tag == "a":   
            if len(attrs) == 0:   
                pass   
            else:   
                for variable, value in attrs:
                    if variable == "href":   
                        self.links.append(value)   

                     
if __name__ == "__main__":
    #写入一个html长字符串
    html_code = """<a href="www.google.com"> google.com</a>
<A Href="www.pythonclub.org"> PythonClub </a>
<A HREF = "www.sina.com.cn"> Sina </a>
"""   
    hp = MyHTMLParser()
    hp.feed(html_code)
    hp.close()
    #print hp.handle_starttag('a', 'href')
    print hp.links #['www.google.com', 'www.pythonclub.org', 'www.sina.com.cn']

关注博主即可阅读全文

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

weixin_34375233

关注关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python模块之HTMLParser解析出URL链接

# -*- coding: utf-8 -*-#python 27#xiaodeng#python模块之HTMLParser解析出URL链接#http://www.cnblogs.com/mfryf/p/3691563.htmlfrom HTMLParser import HTMLParserclass MyHTMLParser(HTMLParser): ...
复制链接

扫一扫