silkworm.middlewares¶
Request/response middlewares.
Implementations live in silkworm._middlewares; this module re-exports
the public API.
- class silkworm.middlewares.AutoThrottleMiddleware[source]¶
Bases:
objectSpace requests per host and adapt the spacing to the host’s latency.
Register the same instance as both a request and a response middleware:
throttle = AutoThrottleMiddleware(start_delay=1.0, max_delay=30.0) run_spider( MySpider, request_middlewares=[throttle], response_middlewares=[throttle], )
Requests to one host are released at least
delayseconds apart. After each response the delay moves towardlatency / target_concurrency(an average of the previous and new estimate), so slow hosts are crawled more gently and fast hosts more quickly. Throttling statuses (429/503 by default) double the delay and honour aRetry-Afterheader; other error responses never lower it. The delay always stays within[min_delay, max_delay].- Parameters:
start_delay – Initial per-host delay in seconds.
min_delay – Lower bound for the delay.
max_delay – Upper bound for the delay.
target_concurrency – Average number of requests to keep in flight per host; higher values crawl faster.
backoff_statuses – Statuses that signal the host is overloaded.
- __init__(*, start_delay=1.0, min_delay=0.0, max_delay=60.0, target_concurrency=1.0, backoff_statuses=(429, 503))[source]¶
- class silkworm.middlewares.CloudflareCrawlMiddleware[source]¶
Bases:
objectRoute opt-in requests through Cloudflare Browser Rendering’s crawl API.
Set request.meta[“cloudflare_crawl”] = True to crawl a URL with the middleware defaults, or assign a dict to provide per-request crawl options. The spider callback receives a synthetic JSON Response containing the final Cloudflare API payload.
- Parameters:
account_id – Cloudflare account identifier.
api_token – API token authorized for Browser Rendering.
crawl_options – Default crawl API fields merged into each submission.
api_base_url – Cloudflare API root, overridable for testing.
poll_interval – Seconds between job-status polls.
timeout – Overall crawl job timeout in seconds.
api_timeout – Timeout for each Cloudflare API request.
- __init__(account_id, api_token, *, crawl_options=None, api_base_url='https://api.cloudflare.com/client/v4', poll_interval=1.0, timeout=300.0, api_timeout=30.0)[source]¶
- class silkworm.middlewares.CookiesMiddleware[source]¶
Bases:
objectStateful cookie middleware using Python’s standards-compliant CookieJar.
By default all requests share one cookie jar. Set request.meta[“cookiejar”] to a string or integer to isolate sessions, set request.meta[“cookies”] to add per-request cookies, and set request.meta[“dont_merge_cookies”] = True to bypass cookie handling for a request/response pair.
- Parameters:
cookies – Initial cookies placed in the default jar.
enabled – Whether request and response cookie processing is active.
allow_domains – Optional domain allowlist passed to the cookie policy.
block_domains – Optional domain blocklist passed to the cookie policy.
rfc2965 – Enable RFC 2965 cookie handling in addition to Netscape cookies.
hide_cookie_header – Hide generated cookie headers from redirect handling performed by the underlying HTTP client.
- __init__(*, cookies=None, enabled=True, allow_domains=None, block_domains=None, rfc2965=False, hide_cookie_header=True)[source]¶
- async process_request(request, spider)[source]¶
Merge explicit and stored cookies into the outgoing request.
The request is returned unchanged when cookie handling is disabled or
dont_merge_cookiesmetadata is true.
- async process_response(response, spider)[source]¶
Store response
Set-Cookievalues in the selected cookie jar.
- clear_session_cookies(cookiejar=None)[source]¶
Discard session cookies from all jars, or a single named jar.
- set_cookie(name, value, *, domain, path='/', secure=False, expires=None, cookiejar=None)[source]¶
Add or replace one cookie in the selected jar.
- save(path, *, cookiejar=None, ignore_discard=True, ignore_expires=True)[source]¶
Save cookies from one jar to a Netscape/Mozilla cookie file.
- class silkworm.middlewares.DelayMiddleware[source]¶
Bases:
objectMiddleware to add configurable delays between requests.
Supports three delay strategies:
Fixed delay: Always wait the same amount of time
Random delay: Wait a random time between min and max
Custom delay: Use a callable that returns delay duration
- Parameters:
delay – Fixed delay in seconds, or None if using delay_func
min_delay – Minimum delay for random strategy (requires max_delay)
max_delay – Maximum delay for random strategy (requires min_delay)
delay_func – Custom callable that returns delay in seconds. Called with
(request, spider)and should return a float.
Example:
# Fixed delay of 1 second DelayMiddleware(delay=1.0) # Random delay between 0.5 and 2 seconds DelayMiddleware(min_delay=0.5, max_delay=2.0) # Custom delay function def my_delay(request, spider): return 1.0 if "fast" in request.url else 2.0 DelayMiddleware(delay_func=my_delay)
- class silkworm.middlewares.ExceptionMiddleware[source]¶
Bases:
ProtocolProtocol for middleware that may recover from request failures.
- async process_exception(request, exception, spider)[source]¶
Return a retry request, or
Noneto leave the error unhandled.
- __init__(*args, **kwargs)¶
- class silkworm.middlewares.ProxyMiddleware[source]¶
Bases:
objectAssign proxies to requests and rotate away from failed proxies.
- Parameters:
proxies – Proxy URLs loaded directly.
proxy_file – UTF-8 file containing one proxy URL per non-empty line.
random_selection – Choose randomly instead of round-robin.
Exactly one proxy source is required. A pre-existing string in
request.meta["proxy"]takes precedence. Register the same instance as exception middleware to retry failures through unused proxies.
- class silkworm.middlewares.RequestMiddleware[source]¶
Bases:
ProtocolProtocol for middleware applied before an HTTP request is sent.
- async process_request(request, spider)[source]¶
Return the request to send, optionally modified or replaced.
- __init__(*args, **kwargs)¶
- class silkworm.middlewares.RequestResponseStreamMiddleware[source]¶
Bases:
objectStream request/response telemetry to a remote HTTP endpoint.
Each outbound request receives a unique exchange_id. The middleware emits a request event before fetch and a response event after fetch, both carrying the same identifier so downstream analysis can join them.
Use the same middleware instance in both request_middlewares and response_middlewares so a single sender queue can stream the full exchange lifecycle:
stream = RequestResponseStreamMiddleware( "https://collector.example.com/events", auth_token="secret-token", batch_size=50, ) run_spider( MySpider, request_middlewares=[stream], response_middlewares=[stream], )
- Parameters:
url – Remote event collector endpoint.
method – HTTP method used for delivery.
headers – Collector request headers.
timeout – Delivery timeout.
max_body_bytes – Maximum request or response body bytes serialized.
queue_size – Bounded in-memory event queue capacity.
auth_token – Optional authorization credential.
auth_scheme – Authorization scheme prepended to
auth_token.batch_size – Events per delivery request.
batch_envelope_key – JSON key containing batched events.
- __init__(url, *, method='POST', headers=None, timeout=10.0, max_body_bytes=64_000, queue_size=1_000, auth_token=None, auth_scheme='Bearer', batch_size=1, batch_envelope_key='events')[source]¶
- async open(spider)[source]¶
Create the sender client, bounded queue, and background task.
- Parameters:
spider (Spider)
- Return type:
None
- async close(spider)[source]¶
Flush queued events, stop the sender, and close its HTTP client.
- Parameters:
spider (Spider)
- Return type:
None
- async process_request(request, spider)[source]¶
Assign an exchange ID and enqueue a serialized request event.
- class silkworm.middlewares.ResponseMiddleware[source]¶
Bases:
ProtocolProtocol for middleware applied after an HTTP response is received.
- async process_response(response, spider)[source]¶
Return a response to dispatch or a request to enqueue instead.
- __init__(*args, **kwargs)¶
- class silkworm.middlewares.RetryMiddleware[source]¶
Bases:
objectRetry failed requests with optional exponential backoff.
Retries both responses with selected HTTP statuses (as a response middleware) and transient transport failures such as timeouts and connection resets (as an exception middleware).
- Parameters:
max_times – Maximum retries after the initial request.
retry_http_codes – Status codes that produce a replacement request.
backoff_base – Base seconds for
base * 2 ** (attempt - 1).sleep_http_codes – Retry statuses that also wait before enqueueing. These codes are automatically added to the retry set.
retry_exceptions – Exception types retried with backoff. Defaults to
DEFAULT_RETRY_EXCEPTIONS; pass()to retry statuses only.
Attempts are stored in
request.meta["retry_times"], shared by status and exception retries, and retry requests bypass deduplication.- __init__(max_times=3, retry_http_codes=None, backoff_base=0.5, sleep_http_codes=None, retry_exceptions=None)[source]¶
- class silkworm.middlewares.RobotsTxtDelayMiddleware[source]¶
Bases:
objectRequest middleware that loads robots.txt and applies its delay directives.
The middleware currently uses Crawl-delay first and falls back to Request-rate when present. Delays are scoped to the origin that served the robots.txt file, and concurrent requests are serialized so the configured spacing is respected under engine concurrency.
- Parameters:
website_url – Absolute HTTP(S) site URL whose origin is throttled.
user_agent – robots.txt group used to resolve directives.
fallback_delay – Delay used when no directive exists or loading fails.
timeout – robots.txt fetch timeout.
ignore_fetch_errors – Apply the fallback instead of propagating errors.
fetcher – Optional async robots.txt loader for custom transports or tests.
- __init__(website_url, *, user_agent='*', fallback_delay=None, timeout=10.0, ignore_fetch_errors=True, fetcher=None)[source]¶
- class silkworm.middlewares.RobotsTxtMiddleware[source]¶
Bases:
objectObey robots.txt rules for every site the crawl visits.
Each origin’s
/robots.txtis fetched once, on its first request. Disallowed requests are dropped withIgnoreRequest(counted asignored_requestswith reasonrobots_txt); withobey_crawl_delaythe origin’sCrawl-delay/Request-ratealso spaces its requests. Requests forrobots.txtitself and requests withmeta["dont_obey_robotstxt"]are never blocked.Following RFC 9309, a 4xx robots.txt response means “no restrictions”. Server errors and unreachable robots.txt files follow
on_unavailable:"allow"(the default, logged as a warning) or"disallow"to skip the origin entirely, as the RFC recommends.- Parameters:
user_agent – Product token matched against robots.txt groups.
obey_crawl_delay – Apply crawl delays per origin.
on_unavailable – Policy when robots.txt cannot be fetched.
timeout – robots.txt fetch timeout.
fetcher – Optional async loader returning robots.txt text; raise an
HttpErrorsubclass with astatusattribute for HTTP errors.
- class silkworm.middlewares.SkipNonHTMLMiddleware[source]¶
Bases:
objectResponse middleware that drops callbacks for non-HTML payloads.
It checks the Content-Type header first, then falls back to a quick body sniff for “<html”. Non-HTML responses keep flowing through the engine but execute a no-op callback so spider parse methods are skipped. Set request.meta[“allow_non_html”] = True to bypass filtering for a request (useful for XML sitemaps, robots.txt fetches, etc.).
- Parameters:
allowed_types – Lowercase tokens accepted in the Content-Type header.
sniff_bytes – Leading body bytes inspected for an HTML tag when headers are inconclusive.
- class silkworm.middlewares.UserAgentMiddleware[source]¶
Bases:
objectSet a missing
User-Agentheader from a pool or fixed default.- Parameters:
user_agents – Values sampled independently for each request.
default – Value used when the pool is empty. Defaults to
"silkworm/0.1".
Existing request headers are never overwritten.