I need help with an lxml statement in python that extracts a meta-tag value

2023-02-06 06:51 问答作者：

I need help fixing this lxml statement to extract the: http://www.etc../1tru.jpg link in the head section of http://www.yfrog.com/9d1truj

#This doe开发者_C百科sn't work!

# <link rel="image_src" href="http://img337.yfrog.com/img337/5023/1tru.jpg" />
def extract_imageurl(self, doc):
    try:
        self.url, = doc.xpath('//head//link[@rel="image_src"][1]/@href')
    except ValueError:
        self.url = "Error"

thanks

In [32]: doc.xpath('//head/link[@rel="image_src"]/@href')[0]
Out[32]: 'http://img337.yfrog.com/img337/5023/1tru.jpg'

Notice xpath returns a list of nodes:

In [25]: doc.xpath('//head/link')
Out[25]: [<Element link at 9c94c5c>, <Element link at 9c94b6c>]

Once you've specified [@rel="image_src"] there is only one node in the list. You can pick off the node with [0] after the xpath call.

In [29]: doc.xpath('//head/link[@rel="image_src"]')[0]
Out[29]: <Element link at 9c94c5c>

import lxml.html as lh
import urllib2

url=r'http://www.yfrog.com/9d1truj'
doc=lh.parse(urllib2.urlopen(url))
link=doc.xpath('//head/link[@rel="image_src"]/@href')[0]
print(link)
# http://img337.yfrog.com/img337/5023/1tru.jpg

继续阅读：html-parsing lxml python

I need help with an lxml statement in python that extracts a meta-tag value

更多精彩内容

精彩评论

最新问答

央视是哪个频道？

请问买过的朋友，舒提啦旅行箱实际使用体验如何？？

检查不孕不育需要的费用？

海信ULED电视画质有什么不同的地方?？

钉子可以挂的住画框幕布吗？

问答排行榜

河神2九牛入海钓河妖是第几集河妖什么来历可活吞牛？

性激素六项检查的最佳时间是多久？多少钱？？

Easiest way to get words of one line from istream into a vector?

《梦在燃烧 (《三国演义》动画片主题曲)》MP3歌词-汤子星？

抽烟只抽炫赫门？