silkworm.httpcache¶
On-disk HTTP response cache for fast, polite development re-runs.
Wrap the engine’s HTTP client with an HttpCache so repeated runs of a
spider (while writing selectors, for example) read pages from disk instead of
downloading them again:
run_spider(MySpider, http_cache=HttpCache(".silkworm/httpcache"))
Entries are keyed by request_fingerprint(), so the same method,
canonical URL, and body share one entry regardless of headers or metadata.
Cached responses carry an x-silkworm-cache: hit header.
- class silkworm.httpcache.HttpCache[source]¶
Bases:
objectConfiguration and storage for cached HTTP responses.
- Parameters:
directory – Cache directory; created when missing.
expiration – Maximum entry age (seconds or
timedelta);Nonekeeps entries forever.ignore_statuses – Response statuses that are never stored (e.g. 5xx).
methods – HTTP methods eligible for caching.
Set
request.meta["dont_cache"] = Trueto bypass the cache for one request.- __init__(directory, *, expiration=None, ignore_statuses=(500, 502, 503, 504, 522, 524, 408, 429), methods=('GET', 'HEAD'))[source]¶
- wrap(client)[source]¶
Return a client that serves
client’s responses from this cache.- Parameters:
client (FetchClient)
- Return type:
- class silkworm.httpcache.CachingHttpClient[source]¶
Bases:
objectHTTP client wrapper that reads and writes an
HttpCache.Created by
HttpCache.wrap(); exposes the same interface as the wrapped client so the engine can use it transparently.- __init__(client, cache)[source]¶
- Parameters:
client (FetchClient)
cache (HttpCache)
- Return type:
None