python请求超时。获得完整的响应

我正在收集网站列表上的统计数据，为了简单起见，我正在使用请求。这是我的代码:

data=[]
websites=['http://google.com', 'http://bbc.co.uk']
for w in websites:
    r= requests.get(w, verify=False)
    data.append( (r.url, len(r.content), r.elapsed.total_seconds(), str([(l.status_code, l.url) for l in r.history]), str(r.headers.items()), str(r.cookies.items())) )

现在，我想要请求。10秒后进入超时，这样循环就不会卡住。

这个问题以前也很有趣，但没有一个答案是干净的。

我听说可能不使用请求是一个好主意，但我应该如何得到请求提供的好东西(元组中的那些)。

当前回答

这可能有点过分，但是芹菜分布式任务队列对超时有很好的支持。

特别是，您可以定义一个软时间限制，它只在您的流程中引发一个异常(这样您就可以清理)和/或一个硬时间限制，它在超过时间限制时终止任务。

在封面之下，这使用了与你的“之前”帖子中引用的相同的信号方法，但以一种更可用和更易于管理的方式。如果你监控的网站列表很长，你可能会从它的主要功能中受益——各种各样的方法来管理大量任务的执行。

2014-02-27 05:47:58

其他回答

最大的问题是，如果无法建立连接，请求包会等待太长时间，并阻塞程序的其余部分。

有几种方法来解决这个问题，但当我寻找类似请求的联机程序时，我找不到任何东西。这就是为什么我为请求构建了一个名为reqto(“请求超时”)的包装器，它支持来自请求的所有标准方法的适当超时。

pip install reqto

语法与请求相同

import reqto

response = reqto.get(f'https://pypi.org/pypi/reqto/json',timeout=1)
# Will raise an exception on Timeout
print(response)

此外，还可以设置自定义超时函数

def custom_function(parameter):
    print(parameter)


response = reqto.get(f'https://pypi.org/pypi/reqto/json',timeout=5,timeout_function=custom_function,timeout_args="Timeout custom function called")
#Will call timeout_function instead of raising an exception on Timeout
print(response)

重要的注意事项是导入行

import reqto

由于monkey_patch在后台运行，需要比所有其他导入更早地导入请求，线程等。

2022-03-18 04:01:40

有一个叫做timeout-decorator的包，你可以用它让任何python函数超时。

@timeout_decorator.timeout(5)
def mytest():
    print("Start")
    for i in range(1,10):
        time.sleep(1)
        print("{} seconds have passed".format(i))

它使用这里的一些答案所建议的信号方法。或者，你可以告诉它使用多处理而不是信号(例如，如果你在多线程环境中)。

2018-09-13 19:37:46

更新:https://requests.readthedocs.io/en/master/user/advanced/超时

在新版本的请求:

如果你为超时指定一个单独的值，像这样:

r = requests.get('https://github.com', timeout=5)

超时值将应用于连接超时和读取超时。如果你想分别设置值，请指定一个元组:

r = requests.get('https://github.com', timeout=(3.05, 27))

如果远程服务器非常慢，您可以告诉Requests永远等待响应，方法是将None作为超时值，然后检索一杯咖啡。

r = requests.get('https://github.com', timeout=None)

我以前的答案(可能已经过时了)(很久以前贴出来的):

还有其他方法可以克服这个问题:

1. 使用TimeoutSauce内部类

来自:https://github.com/kennethreitz/requests/issues/1928 # issuecomment - 35811896

import requests from requests.adapters import TimeoutSauce class MyTimeout(TimeoutSauce): def __init__(self, *args, **kwargs): connect = kwargs.get('connect', 5) read = kwargs.get('read', connect) super(MyTimeout, self).__init__(connect=connect, read=read) requests.adapters.TimeoutSauce = MyTimeout This code should cause us to set the read timeout as equal to the connect timeout, which is the timeout value you pass on your Session.get() call. (Note that I haven't actually tested this code, so it may need some quick debugging, I just wrote it straight into the GitHub window.)

2. 使用kevinburke请求的分支:https://github.com/kevinburke/requests/tree/connect-timeout

从它的文档:https://github.com/kevinburke/requests/blob/connect-timeout/docs/user/advanced.rst

如果你为超时指定一个单独的值，像这样: R = requests.get('https://github.com'， timeout=5) 超时值将应用于连接和读取超时。如果要设置值，请指定一个元组另外: R = requests.get('https://github.com'， timeout=(3.05, 27))

Kevinburke已请求将其合并到主要请求项目中，但尚未被接受。

2014-03-13 11:42:11

我想到了一个更直接的解决方案，虽然很难看，但能解决真正的问题。它是这样的:

resp = requests.get(some_url, stream=True)
resp.raw._fp.fp._sock.settimeout(read_timeout)
# This will load the entire response even though stream is set
content = resp.content

你可以在这里阅读完整的解释

2015-12-03 18:29:58

使用eventlet怎么样?如果你想在10秒后超时请求，即使数据正在接收，下面的代码段将为你工作:

import requests
import eventlet
eventlet.monkey_patch()

with eventlet.Timeout(10):
    requests.get("http://ipv4.download.thinkbroadband.com/1GB.zip", verify=False)

2014-02-28 13:43:58

python请求超时。获得完整的响应

推荐文章

最新文章

标签