利用Python抓取阿里云盘资源

2022-12-11 12:59 开发作者：派森酱

前阵子阿里云盘大火，送了好多的容量空间。而且阿里云盘下载是不限速，这点比百度网盘好太多了。这两天看到一个第三方网站可以搜索阿里云盘上的资源，但是它的资源顺序不是按时间排序的。这种情况会造成排在前面时间久远的资源是一个已经失效的资源。小编这里用 python 抓取后重新排序。

利用Python抓取阿里云盘资源

网页分析

这个网站有两个搜索路线：搜索线路一和搜索线路二，本文章使用的是搜索线路二。

利用Python抓取阿里云盘资源

打开控制面板下的网络，一眼就看到一个 seach.html 的 get 请求。

利用Python抓取阿里云盘资源

上面带了好几个参数，四个关键参数：

page：页数，
keyword：搜索的关键字
category：文件分类，all(全部)，video(视频)，image(图片)，doc(文档)，audio(音频)，zip(压缩文件)，others(其他)，脚本中默认写 all
search_model：搜索的线路

也是在控制面板中，看出这个网页跳转到阿里云盘获取真实的的链接是编程客栈在标题上面的。用 bs4 解析页面上的 div(class=resource-item border-dashed-eee) 标签下的 a 标签就能得到跳转网盘的地址，解析 div 下的 p 标签获取资源日期。

利用Python抓取阿里云盘资源

抓取与解析

首先安装需要的 bs4 第三方库用于解析页面。

pip3installbs4

下面是抓取解析网页的脚本代码，最后按日期降序排序。

importrequests
frombs4importBeautifulSoup
importstring


word=input('请输入要搜索的资源名称：')

headers={
'User-Agent':'Mozilla/5.0(WindowsNT10.0;Win64;x64)AppleWebKit/537.36(KHTML,likeGecko)Chrome/96.0.4664.45Safari/537.36'
}

result_list=[]
foriinrange(1,11):
print('正在搜索第{http://www.cppcns.com}页'.format(i))
params={
'page':i,
'keyword':word,
'search_folder_or_file':0,
'is_search_folder_content':0,
'is_search_path_title':0,
'category':'all',
'file_extension':'all',
'search_model':0
}
response_html=requests.get('https://www.alipanso.com/search.html',headers=headers,params=params)
response_data=response_html.content.decode()

soup=BeautifulSoup(response_data,"html.parser");
divs=soup.find_all('div',class_='resource-itemborder-dashed-eee')

iflen(divs)<=0:
break

fordivindivs[1:]:
p=div.find('p',class_='em')
ifp==None:
break

download_url='https://www.alipanso.com/'+div.a['href']
date=p.text.strip();
name=div.a.text.strip();
result_list.append({'date':date,'name':name,'url':download_url})

iflen(result_list)==0:
break

result_list.sort(key=lambdak:k.get('date'),reverse=True)

示例结果：

利用Python抓取阿里云盘资源

模板

上面抓取完内容后，还需要将内容一个个复制到 google 浏览器中访问，有点太麻烦了。要是直接点击一下能访问就好了。小编在这里就用 Python 的模板方式写一个 html 文件。

模板文件小编是用 elements-ui 做的，下面是关键的代码：

<body>
<divid="app">
<el-table:data="table"style="width:100%":row-class-name="tableRowClassName">
<el-table-columnprop="date"label="日期"width="180"></el-table-column>
<el-table-columnprop="name"label="名称"width="600"></el-table-column>
<el-table-columnlabel="链接">
<templateslot-scope="scope">
<a:href="'http://'+scope.row.url" rel="external nofollow" 
target="_blank"
class="buttonText">{{scope.row.url}}</a>
</template>
</el-table>
</div>

<script>
constApp={
data(){
return{
table:${elements}

};
}
};
constapp=vue.createApp(App);
app.use(ElementPlus);
app.mount("#app");
</script>
</body>

在 python 中读取这个模板文件，并将 ${elements} 关键词替换为上面的解析结果。最后生成一个 report.html 文件。

withopen("aliso.html",encoding='utf-8')ast:
template=string.Template(t.read())

final_output=template.substitute(elements=result_list)
withopen("report.html","w",encoding='utf-8')asoutput:
output.write(final_output)

示例结果：

利用Python抓取阿里云盘资源

跳转到阿里云盘界面

利用Python抓取阿里云盘资源

完整代码

aliso.html

<html>
  <VCkxJXIhead>
    <meta charset="UTF-8" />
    <meta name="viewport" content="width=device-width,initial-scale=1.0" />
    <script src="https://unpkg.com/vue@next"></script>
    <!-- import css -->
    <link rel="stylesheet" href="https://unpkg.com/element-plus/dist/iwww.cppcns.comndex.css">
    <!-- import javascript -->
    <script src="https://unpkg.com/element-plus"></script>
    <title>阿里云盘资源</title>
  </head>
  <body>
    <div id="app">

        <el-table :data="table" style="width: 100%" :row-class-name="tableRowClassName">
            <el-table-column prop="date" label="日期" width="180"> </el-table-column>
            <el-table-column prop="name" label="名称" width="600"> </el-table-column>
            <el-table-column label="链接">
              <template v-slot="scope">
              <a :href="scope.row.url"
                target="_blank"
                class="buttonText">{{scope.row.url}}</a>
            </template>
        </el-table>
    </div>

    <script>
      const App = {
        data() {
          return {
              table: ${elements}
            
          };
        }
      };
      const app = Vue.createApp(App);
      app.use(ElementPlus);
      app.mount("#app");
    </script>
  </body>
</html>

aliso.py

# -*- coding: UTF-8 -www.cppcns.com*-

import requests
from bs4 import BeautifulSoup
import string


word = input('请输入要搜索的资源名称：')
    
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36'
}

result_list = []
for i in range(1, 11):
    print('正在搜索第 {} 页'.format(i))
    params = {
        'page': i,
        'keyword': word,
        'search_folder_or_file': 0,
        'is_search_folder_content': 0,
        'is_search_path_title': 0,
        'category': 'all',
        'file_extension': 'all',
        'search_model': 2
    }
    response_html = requests.get('https://www.alipanso.com/search.html', headers = headers,params=params)
    response_data = response_html.content.decode()
   
    soup = BeautifulSoup(response_data, "html.parser");
    divs = soup.find_all('div', class_='resource-item border-dashed-eee')
    
    if len(divs) <= 0:
        break

    for div in divs[1:]:
        p = div.find('p',class_='em')
        if p == None:
            break

        download_url = 'https://www.alipanso.com/' + div.a['href']
        date = p.text.strip();
        name = div.a.text.strip();
        result_list.append({'date':date, 'name':name, 'url':download_url})
    
    if len(result_list) == 0:
        break
    
result_list.sort(key=lambda k: k.get('date'),reverse=True)
print(result_list)

with open("aliso.html", encoding='utf-8') as t:
    template = string.Template(t.read())

final_output = template.substitute(elements=result_list)
with open("report.html", "w", encoding='utf-8') as output:
    output.write(final_output)

总结

用 python 做一些小爬虫，不仅去掉网站上烦人的广告，也更加的便利了。

以上就是利用Python抓取阿里云盘资源的详细内容，更多关于Python抓取云盘资源的资料请关注我们其它相关文章！

继续阅读：Python 抓取云盘资源 Python 抓取阿里云盘资源 Python 阿里云盘

利用Python抓取阿里云盘资源

目录

网页分析

抓取与解析

模板

完整代码

总结

更多精彩内容

精彩评论

最新开发

Go语言中uintptr和unsafe.Pointer的区别的实现小结

Go语言中栈扩容和栈缩容的使用

Go 语言中的命令行参数操作详解

浅谈Go 语言中逃逸分析是怎么进行的

Go语言错误和异常实现

开发排行榜

springboot后端存储富文本内容的思路与步骤(含图片内容)

PyCharm运行python测试,报错“没有发现测试”/“空套件”的解决

return base64.b64encode(b).decode(

基于C语言实现钻石棋游戏的示例代码

Sublime Text 3解决中文乱码问题（实测可用）