python 爬取媒体文件（无防火墙）

#coding = utf-8

import requests

import pandas as pd

import os,time

root_path = './根目录/'

input_file = '码表.xlsx'

url = 'http://api.map.baidu.com/geocoder/v2/?id = %s&local=1'

fail_file = root_path +'fail.csv'

class Auto_down:

    def __init__(self):

        print("--start--")

    def read_excel(self):

        # pd.read_excel(converters = {u'列名':str})按照str类型读入，不会出现0被舍去的情况

        sheet = pd.read_excel(input_file,converters = {u'列名':str},sheetname = '子表名')

        cust_Id = sheet['cust_id']

        void_Id = sheet['void_id']

        for i in range(len(cust_Id)):

            self.create_file(cust_Id[i],void_Id[i])

    def download_voice(self,custid_filename,voiceid):

        print(voiceid)

        try:

            r = requests.get(url%voiceid)

            return_code = r.status_code

            if return_code == 200:

                voice_filename = '%s/%s.mp3'%(custid_filename,voiceid)

                with open(voice_filename, 'wb') as fd:

                    fd.write(r.content)

            else:

                with open(fail_file, 'a+') as ff:

                    ff.write(voiceid + '\n')

        except:

            print('request url is fail!!')

            with open(fail_file, 'a+') as ff:

                ff.write(voiceid + '\n')

    def create_file(self, custid, voiceid):

        custid_filename = root_path + custid

        if not os.path.exists(custid_filename):

            os.mkdir(custid_filename)

        else:

            self.download_voice(custid_filename,voiceid)

if __name__ == '__main__':

    tStart = time.clock()

    AD = Auto_down()

    AD.read_excel()

    tEnd = time.clock()

    print("%s s"%(tEnd - tStart))

#coding = utf-8

import requests

root_path = "./下载/"

url = ""

fail_file = root_path + 'fail.csv'

voiceid = ''

for i in range(3):

    try:

        r = requests.get(url)

        return_code = r.status_code

        if r.status_code == 200:

            voice_filename = root_path + 'dada.fdf'

            with open(voice_filename,'wb') as fd:

                fd.write(r.content)

        else:

            with open(fail_file,'a+') as ff:

                ff.write(voiceid + '\n')

    except:

        prin("fail")

        with open(fail_file,'a+') as ff:

            ff.write(voiceid + '\n')

r = request.get(url)
r.status_code 获取响应状态码
r.text 获取响应内容
r.headers 获取响应头
r.encoding 获取响应编码
r.content 获取二进制响应内容
r.json() 获取JSON响应内容

python 爬取媒体文件（无防火墙）的更多相关文章

python 爬取媒体文件（使用chrome代理，启动客户端，有防火墙）
#coding = utf-8 ''' 中文转经纬度 ''' import time,json import urllib.request from selenium import webdriver ...
scrapy --爬取媒体文件示例详解
scrapy 图片数据的爬取基于scrapy进行图片数据的爬取: 在爬虫文件中只需要解析提取出图片地址,然后将地址提交给管道配置文件中写入文件存储位置:IMAGES_STORE = './imgs ...
python爬取当当网的书籍信息并保存到csv文件
python爬取当当网的书籍信息并保存到csv文件依赖的库: requests #用来获取页面内容 BeautifulSoup #opython3不能安装BeautifulSoup,但可以安装Bea ...
python爬取网站数据
开学前接了一个任务,内容是从网上爬取特定属性的数据.正好之前学了python,练练手. 编码问题因为涉及到中文,所以必然地涉及到了编码的问题,这一次借这个机会算是彻底搞清楚了. 问题要从文字的编码讲 ...
python爬取网站数据保存使用的方法
这篇文章主要介绍了使用Python从网上爬取特定属性数据保存的方法,其中解决了编码问题和如何使用正则匹配数据的方法,详情看下文编码问题因为涉及到中文,所以必然地涉及到了编码的问题,这一次借这 ...
Python爬取中国天气网
Python爬取中国天气网基于requests库制作的爬虫. 使用方法:打开终端输入 “python3 weather.py 北京(或你所在的城市)" 程序正常运行需要在同文件夹下加入一个 ...
用Python爬取B站、腾讯视频、爱奇艺和芒果TV视频弹幕！
众所周知,弹幕,即在网络上观看视频时弹出的评论性字幕.不知道大家看视频的时候会不会点开弹幕,于我而言,弹幕是视频内容的良好补充,是一个组织良好的评论序列.通过分析弹幕,我们可以快速洞察广大观众对于视频 ...
毕设之Python爬取天气数据及可视化分析
写在前面的一些P话:(https://jq.qq.com/?_wv=1027&k=RFkfeU8j) 天气预报我们每天都会关注,我们可以根据未来的天气增减衣物.安排出行,每天的气温.风速风向. ...
Python 爬取途虎养车全系车型轮胎保养数据
Python 爬取途虎养车全系车型轮胎保养数据 2021.7.27 更新增加标题.发布时间参数 demo文末自行下载,需要完整数据私聊我 2021.2.19 更新增加大保养数据 2020. ...

随机推荐

[笔记] vs code 设置终端
设置文件: setting.json 1 设置自定义终端 cmd "terminal.integrated.shell.windows": "C:\\WINDOWS\\S ...
Java面试题：Java中的集合及其继承关系
关于集合的体系是每个人都应该烂熟于心的,尤其是对我们经常使用的List,Map的原理更该如此.这里我们看这张图即可: 1.List.Set.Map是否继承自Collection接口? List.Set ...
Vue.js项目实战-多语种网站（租车）
首先来看一下网站效果,想写这个项目的读者可以自行下载哦,地址:https://github.com/Stray-Kite/Car: 在这个项目中,我们主要是为了学习语种切换,也就是右上角的中文/En ...
ucoreOS_lab3 实验报告
所有的实验报告将会在 Github 同步更新,更多内容请移步至Github:https://github.com/AngelKitty/review_the_national_post-graduat ...
Android不显示开机向导和开机气泡
修改好的代码下载地址: https://github.com/Vico-H/Launcher 不显示开机向导修改Launcher2.java的代码 (文件位置: /alps/packages/app ...
flink WaterMark之TumblingEventWindow
1.WaterMark,翻译成水印或水位线,水印翻译更抽象,水位线翻译接地气. watermark是用于处理乱序事件的,通常用watermark机制结合window来实现. 流处理从事件产生,到流经s ...
flask uwsgi和nginx配置信息
1. 安装 pip3 install uwsgi 2. uwsgi配置信息创建一个uwsgi.ini文件 [uwsgi] socket=/opt/script/uwsgi.sock #启动程序时所使 ...
kolla部署openstack allinone，报错 ImportError: cannot import name decorate
使用 kolla-ansible 部署 opnenstack:stein,最后无法导入变量脚本,报错信息如下: [root@kolla ~]# . /etc/kolla/admin-openrc.sh ...
LoadRunner性能测试工具下载
LoadRunner性能测试工具 LoadRunner是前美科利(Mercury Interactive)公司著名的性能测试产品.Mercury公司曾经是全球业务优化科技领域的领导者.2006年由惠普 ...
JS高阶---对象创建模式（5种）
[前言] 函数高级部分先看到这里,接下里看下面向对象高级部分 .对象创建模式 .继承模式 [主体] (1)Object构造函数模式案例如下: 测试结果如右图所示 (2)对象字面量形式创建案例如下: ...

python 爬取媒体文件（无防火墙）

python 爬取媒体文件（无防火墙）的更多相关文章

随机推荐

热门专题