scrapy获取a标签的连接_Scrapy：获取某个<a>标记后面的所有标记

最新推荐文章于 2021-06-07 18:01:00 发布

weixin_39611031

最新推荐文章于 2021-06-07 18:01:00 发布

阅读量334

点赞数

文章标签： scrapy获取a标签的连接

本文链接：https://blog.csdn.net/weixin_39611031/article/details/113998918

版权

更新：

您可以使用以sel.xpath('.//a[@name="summaries"]')开头的xpath。。。我在这台mac电脑上没什么问题，所以我用的是lxml，事实上，在lxml中，你可以使用getparent()，iterslibings等等。实际上，这里有一个例子：from lxml import html

s = '... your very long html page source ...'

tree = html.fromstring(s)

for a in tree.xpath('.//a[@name="summaries"]'):

td = a.getparent() # getparent() which returns td

# iterchildren() get all children nodes under td

for node in td.iterchildren():

print node.text

结果：

^{pr2}$

或者，使用itersiblings()获取周围的所有同级节点：for a in tree.xpath('.//a[@name="summaries"]'):

for node in t.itersiblings():

print node.text

。。。在

或者，如果您是在父文件td中实际包含的所有文本，则可以使用xpath //text()来获取所有文本：for a in tree.xpath('.//a[@name="summaries"]'):

print a.xpath('./..//text()')

非常长的结果：['\n\t', '\n', '\n', 'Jump to Section...', '\n', 'Aliases', '\n', 'Databases', '\n', 'Disorders / Diseases', '\n', 'Domains / Families', '\n', 'Drugs / Compounds', '\n', 'Expression', '\n', 'Function', '\n', 'Genomic Views', '\n', 'Intellectual Property', '\n', 'Localization', '\n', 'Orthologs', '\n', 'Paralogs', '\n', 'Pathways / Interactions', '\n', 'Products', '\n', 'Proteins', '\n', 'Publications', '\n', 'Search Box', '\n', 'Summaries', '\n', 'Transcripts', '\n', 'Variants', '\n', 'TOP', '\n', 'BOTTOM', '\n', '\n', '\n', 'Summaries', 'for B2M gene', '(According to ', 'Entrez Gene', ',\n\t\t', 'GeneCards', ',\n\t\t', 'Tocris Bioscience', ',\n\t\t', "Wikipedia's", ' \n\t\t', 'Gene Wiki', ',\n\t\t', 'PharmGKB', ',', '\n\t\t', 'UniProtKB/Swiss-Prot', ',\n\t\tand/or \n\t\t', 'UniProtKB/TrEMBL', ')\n\t\t', 'About This Section', 'Try', 'GeneCards Plus']

['Entrez Gene summary for ', 'B2M', ' Gene:', ' This gene encodes a serum protein found in association with the major histocompatibility complex (MHC) class I', 'heavy chain on the surface of nearly all nucleated cells. The protein has a predominantly beta-pleated sheet', 'structure that can form amyloid fibrils in some pathological conditions. A mutation in this gene has been shown', 'to result in hypercatabolic hypoproteinemia.(provided by RefSeq, Sep 2009) ', 'GeneCards Summary for B2M Gene:', ' B2M (beta-2-microglobulin) is a protein-coding gene. Diseases associated with B2M include ', 'balkan nephropathy', ', and ', 'plasmacytoma', '. GO annotations related to this gene include ', 'identical protein binding', '.', 'UniProtKB/Swiss-Prot: ', 'B2MG_HUMAN, P61769', 'Function', ': Component of the class I major histocompatibility complex (MHC). Involved in the presentation of peptide', 'antigens to the immune system', 'Gene Wiki entry for ', 'B2M', ' (Beta-2 microglobulin) Gene']

weixin_39611031

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
复制链接

分享到 QQ

分享到新浪微博

扫一扫