scrapy爬虫框架setting模块解析

2020-03-11 分类：Python 作者：admin 阅读（62）

平时写爬虫的时候并不需要设置setting里所有的参数，今天心血来潮，花了点时间查了一下setting模块创建后自动写入的所有参数的含义，记录一下。

模块相关说明信息

# -*- coding: utf-8 -*-

# Scrapy settings for new_center project
#
# For simplicity, this file contains only settings considered important or
# commonly used. You can find more settings consulting the documentation:
#
#     http://doc.scrapy.org/en/latest/topics/settings.html
#     http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html
#     http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html

# -*- coding: utf-8 -*-

# Scrapy settings for new_center project

# For simplicity, this file contains only settings considered important or

# commonly used. You can find more settings consulting the documentation:

# http://doc.scrapy.org/en/latest/topics/settings.html

# http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html

# http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html

项目名字和爬虫模块说明，引擎根据这个信息找到爬虫

BOT_NAME = 'new_center'  # 项目名字

SPIDER_MODULES = ['new_center.spiders']
NEWSPIDER_MODULE = 'new_center.spiders'

BOT_NAME = 'new_center' # 项目名字

SPIDER_MODULES = ['new_center.spiders']

NEWSPIDER_MODULE = 'new_center.spiders'

浏览器的USER_AGENT，可以自定义伪装。

# Crawl responsibly by identifying yourself (and your website) on the user-agent
#USER_AGENT = 'new_center (+http://www.yourdomain.com)'

1 2	# Crawl responsibly by identifying yourself (and your website) on the user-agent #USER_AGENT = 'new_center (+http://www.yourdomain.com)'

是否遵守robots协议，默认是遵守的，可以改成False或将其注释

# Obey robots.txt rules
ROBOTSTXT_OBEY = True

1 2	# Obey robots.txt rules ROBOTSTXT_OBEY = True

设置scrapy爬虫最大的并发请求数量，默认是16

# Configure maximum concurrent requests performed by Scrapy (default: 16)
#CONCURRENT_REQUESTS = 32

1 2	# Configure maximum concurrent requests performed by Scrapy (default: 16) #CONCURRENT_REQUESTS = 32

设置对访问同一个网站进行请求的延时时间，默认情况下，Scrapy在两个请求间不等待一个固定的值，而是使用0.5到1.5之间的一个随机值。

# Configure a delay for requests for the same website (default: 0)
# See http://scrapy.readthedocs.org/en/latest/topics/settings.html#download-delay
# See also autothrottle settings and docs
#DOWNLOAD_DELAY = 3
# The download delay setting will honor only one of:

# Configure a delay for requests for the same website (default: 0)

# See http://scrapy.readthedocs.org/en/latest/topics/settings.html#download-delay

# See also autothrottle settings and docs

#DOWNLOAD_DELAY = 3

# The download delay setting will honor only one of:

设置对每个网站和每个IP的最大并发请求数量，两个最好只设置一个，如果都设置，则按照限制IP生效。

#CONCURRENT_REQUESTS_PER_DOMAIN = 16
#CONCURRENT_REQUESTS_PER_IP = 16

1 2	#CONCURRENT_REQUESTS_PER_DOMAIN = 16 #CONCURRENT_REQUESTS_PER_IP = 16

设置是否禁用cookie，目前默认是可用的，去掉注释则禁用

# Disable cookies (enabled by default)
#COOKIES_ENABLED = False

1 2	# Disable cookies (enabled by default) #COOKIES_ENABLED = False

设置是否可远程登录控制台，目前默认是可以的，去掉注释则禁用

# Disable Telnet Console (enabled by default)
#TELNETCONSOLE_ENABLED = False

1 2	# Disable Telnet Console (enabled by default) #TELNETCONSOLE_ENABLED = False

用来设置请求头，一般不用，因为请求头可以动态设置

# Override the default request headers:
#DEFAULT_REQUEST_HEADERS = {
#   'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
#   'Accept-Language': 'en',
#}

# Override the default request headers:

#DEFAULT_REQUEST_HEADERS = {

# 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',

# 'Accept-Language': 'en',

是否开启使用爬虫spider的中间件,默认不启用，解除注释后启用，后面的数字代表优先级，数字越小，优先级越高

# Enable or disable spider middlewares
# See http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html
#SPIDER_MIDDLEWARES = {
#    'new_center.middlewares.NewCenterSpiderMiddleware': 543,
#}

# Enable or disable spider middlewares

# See http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html

#SPIDER_MIDDLEWARES = {

# 'new_center.middlewares.NewCenterSpiderMiddleware': 543,

是否开启爬虫下载器的中间件,默认不启用，解除注释后启用

# Enable or disable downloader middlewares
# See http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html
#DOWNLOADER_MIDDLEWARES = {
#    'new_center.middlewares.MyCustomDownloaderMiddleware': 543,
#}

# Enable or disable downloader middlewares

# See http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html

#DOWNLOADER_MIDDLEWARES = {

# 'new_center.middlewares.MyCustomDownloaderMiddleware': 543,

是否禁用爬虫扩展，默认禁用，解除注释后将None改成数字，如500，扩展的优先级一般不重要，因为他们并不相互依赖，多个扩展的value值可以写相同。

# Enable or disable extensions
# See http://scrapy.readthedocs.org/en/latest/topics/extensions.html
#EXTENSIONS = {
#    'scrapy.extensions.telnet.TelnetConsole': None,
#}

# Enable or disable extensions

# See http://scrapy.readthedocs.org/en/latest/topics/extensions.html

#EXTENSIONS = {

# 'scrapy.extensions.telnet.TelnetConsole': None,

是否开启管道，默认关闭，开启則解除注释。是一个字典，默认为空，值任意，不过一般设置在0-1000之间，值越小优先级越高。

# Configure item pipelines
# See http://scrapy.readthedocs.org/en/latest/topics/item-pipeline.html
#ITEM_PIPELINES = {
#    'new_center.pipelines.NewCenterPipeline': 300,
#}

# Configure item pipelines

# See http://scrapy.readthedocs.org/en/latest/topics/item-pipeline.html

#ITEM_PIPELINES = {

# 'new_center.pipelines.NewCenterPipeline': 300,

设置自动限速，根据Scrapy服务器及爬取的网站的负载自动限制爬取速度，默认关闭，开启需解除注释。

# Enable and configure the AutoThrottle extension (disabled by default)
# See http://doc.scrapy.org/en/latest/topics/autothrottle.html
#AUTOTHROTTLE_ENABLED = True  # 自动限速的开关
# The initial download delay  
#AUTOTHROTTLE_START_DELAY = 5  # 初始下载延时
# The maximum download delay to be set in case of high latencies
#AUTOTHROTTLE_MAX_DELAY = 60  # 最大下载延时
# The average number of requests Scrapy should be sending in parallel to
# each remote server
#AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
# Enable showing throttling stats for every response received:
#AUTOTHROTTLE_DEBUG = False

# Enable and configure the AutoThrottle extension (disabled by default)

# See http://doc.scrapy.org/en/latest/topics/autothrottle.html

#AUTOTHROTTLE_ENABLED = True # 自动限速的开关

# The initial download delay

#AUTOTHROTTLE_START_DELAY = 5 # 初始下载延时

# The maximum download delay to be set in case of high latencies

#AUTOTHROTTLE_MAX_DELAY = 60 # 最大下载延时

# The average number of requests Scrapy should be sending in parallel to

# each remote server

#AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0

# Enable showing throttling stats for every response received:

#AUTOTHROTTLE_DEBUG = False

启用和配置HTTP缓存

# Enable and configure HTTP caching (disabled by default)
# See http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html#httpcache-middleware-settings
#HTTPCACHE_ENABLED = True
#HTTPCACHE_EXPIRATION_SECS = 0
#HTTPCACHE_DIR = 'httpcache'
#HTTPCACHE_IGNORE_HTTP_CODES = []
#HTTPCACHE_STORAGE = 'scrapy.extensions.httpcache.FilesystemCacheStorage'

# Enable and configure HTTP caching (disabled by default)

# See http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html#httpcache-middleware-settings

#HTTPCACHE_ENABLED = True

#HTTPCACHE_EXPIRATION_SECS = 0

#HTTPCACHE_DIR = 'httpcache'

#HTTPCACHE_IGNORE_HTTP_CODES = []

#HTTPCACHE_STORAGE = 'scrapy.extensions.httpcache.FilesystemCacheStorage'

另外的重要的设置：

DEPTH_LIMIT：爬取网站最大的允许深度，默认值为0，表示没有限制；如果为1，表示只允许解析一层的url。
DOWNLOAD_TIMEOUT:下载器超时时间，默认180s.
LOG_ENABLED:是否启用logging，即日志文件，默认为True。
LOG_ENCODING：设置日志文件的编码，默认使用UTF_8。
LOG_LEVEL：对日志的内容进行等级的设置，有五个级别可以选择:

CRITICAL: 严重错误；
ERROR：一般错误；
WARNING：警告信息；
INFO：一般信息；
DEBUG：调试信息

CRITICAL: 严重错误；

ERROR：一般错误；

WARNING：警告信息；

INFO：一般信息；

DEBUG：调试信息

PROXIES：设置代理；这不是scrapy内置的参数，但可以定义在setting模块，使用的时候可以导入。

「三年博客，如果觉得我的文章对您有用，请帮助本站成长」

赞(0) 打赏

标签：

上一篇：正则表达式(括号)、[中括号]、{大括号}的区别小结

下一篇：实体entity、JavaBean、Model、POJO、domain的区别

你可能感兴趣

python -m SimpleHTTPServer 实现快速上传下载

redis 批量拉取与更新

python 变量名命名规则

Python 多线程，threading模块，创建子线程的两种方式

Python 【收藏夹】

python gevnet

Python3 next send yield

Python多进程和多线程（跑满CPU）

Centos7 安装Python3和scrapy（正确安装姿势）

python 正则

共有 0 - scrapy爬虫框架setting模块解析

点击这里取消回复。

博客简介

ai & trae 2026年05月07日 3
git 从本地到远程 2026年04月19日 10
什么是跨域？如何解决跨域问题？ 2021年07月30日 73
git 三路合并范例 2026年01月27日 45
git 常用命令 2022年01月20日 117

admin（6年前 (2020-03-09)）
分别用不同厚度的筏板定义，画图后这设置筏板变截面处理。 http://f.fwxgx.co...
评：新文章！
admin（6年前 (2020-03-09)）
分别用不同厚度的筏板定义，画图后这设置筏板变截面处理。 http://f.fwxgx.co...
评：新文章！
admin（6年前 (2020-03-09)）
新增一个框架图！ http://biji.jinli.vip/wp-content/upl...
评：新文章！
一位WordPress评论者（6年前 (2020-02-13)）
嗨，这是一条评论。要开始审核、编辑及删除评论，请访问仪表盘的“评论”页面。评论者头像来自...
评：世界，您好！

scrapy爬虫框架setting模块解析

你可能感兴趣

共有 0 - scrapy爬虫框架setting模块解析

点击这里取消回复。

博客简介

最近更新

随机文章

热门标签

精彩评论