Parse xml file while a tag is missing

2022-12-15 12:09 问答作者：

I try to parse an xml file. The text which is in tags is parsed successfully (or it开发者_如何学运维 seems so) but I want to output as the text which is not contained in some tags and the following program just ignores it.

from xml.etree.ElementTree import XMLTreeBuilder

class HtmlLatex:                     # The target object of the parser
    out = ''
    var = ''
    def start(self, tag, attrib):   # Called for each opening tag.
        pass
    def end(self, tag):             # Called for each closing tag.
        if tag == 'i':
            self.out += self.var
        elif tag == 'sub':
            self.out += '_{' + self.var + '}'
        elif tag == 'sup':
            self.out += '^{' + self.var + '}'
        else:
            self.out += self.var
    def data(self, data):
        self.var = data
    def close(self):
        print(self.out)


if __name__ == '__main__':
    target = HtmlLatex()
    parser = XMLTreeBuilder(target=target)

    text = ''
    with open('input.txt') as f1:
        text = f1.read()

    print(text)

    parser.feed(text)
    parser.close()

A part of the input I want to parse: p0 = (m3+(2l2+l1) m2+(l22+2l1 l2+l12) m) /(m3+(3l2+2l1) ) }.

Have a look at BeautifulSoup, a python library for parsing, navigating and manipulating html and xml. It has a handy interface and might solve your problem ...

Here's a pyparsing version - I hope the comments are sufficiently explanatory.

src = """<p><i>p</i><sub>0</sub> = (<i>m</i><sup>3</sup>+(2<i>l</i><sub>2</sub>+<i>l</i><sub>1</sub>) """ \
      """<i>m</i><sup>2</sup>+(<i>l</i><sub>2</sub><sup>2</sup>+2<i>l</i><sub>1</sub> <i>l</i><sub>2</sub>+""" \
      """<i>l</i><sub>1</sub><sup>2</sup>) <i>m</i>) /(<i>m</i><sup>3</sup>+(3<i>l</i><sub>2</sub>+""" \
      """2<i>l</i><sub>1</sub>) ) }.</p>"""

from pyparsing import makeHTMLTags, anyOpenTag, anyCloseTag, Suppress, replaceWith

# set up tag matching for <sub> and <sup> tags
SUB,endSUB = makeHTMLTags("sub")
SUP,endSUP = makeHTMLTags("sup")

# all other tags will be suppressed from the output
ANY,endANY = map(Suppress,(anyOpenTag,anyCloseTag))

SUB.setParseAction(replaceWith("_{"))
SUP.setParseAction(replaceWith("^{"))
endSUB.setParseAction(replaceWith("}"))
endSUP.setParseAction(replaceWith("}"))

transformer = (SUB | endSUB | SUP | endSUP | ANY | endANY)

# now use the transformer to apply these transforms to the input string
print transformer.transformString(src)

Gives

p_{0} = (m^{3}+(2l_{2}+l_{1}) m^{2}+(l_{2}^{2}+2l_{1} l_{2}+l_{1}^{2}) m) /(m^{3}+(3l_{2}+2l_{1}) ) }.

继续阅读：parsing python xml

Parse xml file while a tag is missing

更多精彩内容

精彩评论

最新问答

央视是哪个频道？

请问买过的朋友，舒提啦旅行箱实际使用体验如何？？

检查不孕不育需要的费用？

海信ULED电视画质有什么不同的地方?？

钉子可以挂的住画框幕布吗？

问答排行榜

河神2九牛入海钓河妖是第几集河妖什么来历可活吞牛？

性激素六项检查的最佳时间是多久？多少钱？？

Easiest way to get words of one line from istream into a vector?

《梦在燃烧 (《三国演义》动画片主题曲)》MP3歌词-汤子星？

抽烟只抽炫赫门？