How do I extract all HTML tags from a webpage into an array?

2023-01-19 03:18 问答作者：

I need to extract all HTML tags from a webpage into an array without the data inside the tags. It would look something like...

I'm using PHP

Array 
{
   html =>
             Array 
             {
                 head =>
                          Array
                          {
                              title,
                              meta name='description' content='bla bla'
                              meta name='keyword' content='bla bla'
                              ....
                          },
                 body =>
                          Array
                          {
                              div id='header' =>
                                              Array
                                              {
                                                  div class='logo',
                                                  div class='nav'
                                              },
                              div id='content' =>
                                              Array
                                              {
                                                  h1,
                                                  p class='first-para',
                                                  p,
                                                  p,
                                                  div id='ad'
                                              },
                              div id='footer' =>
                                              Array
                                              {
                                                  ul =>
                                                        Array
                                                        {
                                                            li =>
                                                                  Array
                                                                  {
                                                                     a href='link.htm'
                                                                  },
                                                            li =>
                                                                  Array
                                                                  {
                                                                     a href='link.htm'
                                                                  },
                                                            li =>
                                                                  Array
     开发者_StackOverflow中文版                                                             {
                                                                     a href='link.htm'
                                                                  }
                                                        }
                                              }
                          }

             }
}

What you need is an HTML parser (an XML parser would probably not do because HTML often is invalid). Maybe: http://simplehtmldom.sourceforge.net/

You can also use the PHP DOM extension.

I think the simplest way is to use XPath.

//*::name()

Should give you the names of all nodes on all levels. Iam not sure wheather not hierarchy will be flattened though.

继续阅读：dom extract php xml

How do I extract all HTML tags from a webpage into an array?

更多精彩内容

精彩评论

最新问答

央视是哪个频道？

请问买过的朋友，舒提啦旅行箱实际使用体验如何？？

检查不孕不育需要的费用？

海信ULED电视画质有什么不同的地方?？

钉子可以挂的住画框幕布吗？

问答排行榜

河神2九牛入海钓河妖是第几集河妖什么来历可活吞牛？

性激素六项检查的最佳时间是多久？多少钱？？

Easiest way to get words of one line from istream into a vector?

《梦在燃烧 (《三国演义》动画片主题曲)》MP3歌词-汤子星？

抽烟只抽炫赫门？