silkworm.middlewares

Request/response middlewares.

Implementations live in silkworm._middlewares; this module re-exports the public API.

class silkworm.middlewares.AutoThrottleMiddleware[source]

Bases: object

Space requests per host and adapt the spacing to the host’s latency.

Register the same instance as both a request and a response middleware:

throttle = AutoThrottleMiddleware(start_delay=1.0, max_delay=30.0)
run_spider(
    MySpider,
    request_middlewares=[throttle],
    response_middlewares=[throttle],
)

Requests to one host are released at least delay seconds apart. After each response the delay moves toward latency / target_concurrency (an average of the previous and new estimate), so slow hosts are crawled more gently and fast hosts more quickly. Throttling statuses (429/503 by default) double the delay and honour a Retry-After header; other error responses never lower it. The delay always stays within [min_delay, max_delay].

Parameters:
  • start_delay – Initial per-host delay in seconds.

  • min_delay – Lower bound for the delay.

  • max_delay – Upper bound for the delay.

  • target_concurrency – Average number of requests to keep in flight per host; higher values crawl faster.

  • backoff_statuses – Statuses that signal the host is overloaded.

__init__(*, start_delay=1.0, min_delay=0.0, max_delay=60.0, target_concurrency=1.0, backoff_statuses=(429, 503))[source]
Parameters:
  • start_delay (float)

  • min_delay (float)

  • max_delay (float)

  • target_concurrency (float)

  • backoff_statuses (Iterable[int])

Return type:

None

delay_for(url)[source]

Return the current delay for url’s host (the start delay if unseen).

Parameters:

url (str)

Return type:

float

async process_request(request, spider)[source]

Wait for the host’s next slot, then release the request.

Parameters:
Return type:

Request

async process_response(response, spider)[source]

Adjust the host’s delay from the response latency and status.

Parameters:
Return type:

Response | Request

class silkworm.middlewares.CloudflareCrawlMiddleware[source]

Bases: object

Route opt-in requests through Cloudflare Browser Rendering’s crawl API.

Set request.meta[“cloudflare_crawl”] = True to crawl a URL with the middleware defaults, or assign a dict to provide per-request crawl options. The spider callback receives a synthetic JSON Response containing the final Cloudflare API payload.

Parameters:
  • account_id – Cloudflare account identifier.

  • api_token – API token authorized for Browser Rendering.

  • crawl_options – Default crawl API fields merged into each submission.

  • api_base_url – Cloudflare API root, overridable for testing.

  • poll_interval – Seconds between job-status polls.

  • timeout – Overall crawl job timeout in seconds.

  • api_timeout – Timeout for each Cloudflare API request.

__init__(account_id, api_token, *, crawl_options=None, api_base_url='https://api.cloudflare.com/client/v4', poll_interval=1.0, timeout=300.0, api_timeout=30.0)[source]
Parameters:
Return type:

None

async process_request(request, spider)[source]

Replace opted-in network work with a synthetic crawl API response.

Raises:
  • TypeError – If cloudflare_crawl metadata is not a supported value.

  • HttpError – If submission, polling, or result retrieval fails.

Parameters:
Return type:

Request

async close(spider)[source]

Close the internal client used for Cloudflare API calls.

Parameters:

spider (Spider)

Return type:

None

class silkworm.middlewares.CookiesMiddleware[source]

Bases: object

Stateful cookie middleware using Python’s standards-compliant CookieJar.

By default all requests share one cookie jar. Set request.meta[“cookiejar”] to a string or integer to isolate sessions, set request.meta[“cookies”] to add per-request cookies, and set request.meta[“dont_merge_cookies”] = True to bypass cookie handling for a request/response pair.

Parameters:
  • cookies – Initial cookies placed in the default jar.

  • enabled – Whether request and response cookie processing is active.

  • allow_domains – Optional domain allowlist passed to the cookie policy.

  • block_domains – Optional domain blocklist passed to the cookie policy.

  • rfc2965 – Enable RFC 2965 cookie handling in addition to Netscape cookies.

  • hide_cookie_header – Hide generated cookie headers from redirect handling performed by the underlying HTTP client.

__init__(*, cookies=None, enabled=True, allow_domains=None, block_domains=None, rfc2965=False, hide_cookie_header=True)[source]
Parameters:
  • cookies (Mapping[str, object] | None)

  • enabled (bool)

  • allow_domains (Iterable[str] | None)

  • block_domains (Iterable[str] | None)

  • rfc2965 (bool)

  • hide_cookie_header (bool)

Return type:

None

async process_request(request, spider)[source]

Merge explicit and stored cookies into the outgoing request.

The request is returned unchanged when cookie handling is disabled or dont_merge_cookies metadata is true.

Parameters:
Return type:

Request

async process_response(response, spider)[source]

Store response Set-Cookie values in the selected cookie jar.

Parameters:
Return type:

Response | Request

clear(cookiejar=None)[source]

Clear all cookies, or only the named cookie jar.

Parameters:

cookiejar (str | int | None)

Return type:

None

clear_session_cookies(cookiejar=None)[source]

Discard session cookies from all jars, or a single named jar.

Parameters:

cookiejar (str | int | None)

Return type:

None

Add or replace one cookie in the selected jar.

Parameters:
Return type:

None

save(path, *, cookiejar=None, ignore_discard=True, ignore_expires=True)[source]

Save cookies from one jar to a Netscape/Mozilla cookie file.

Parameters:
Return type:

None

load(path, *, cookiejar=None, ignore_discard=True, ignore_expires=True, clear_existing=False)[source]

Load cookies from a Netscape/Mozilla cookie file into one jar.

Parameters:
Return type:

None

class silkworm.middlewares.DelayMiddleware[source]

Bases: object

Middleware to add configurable delays between requests.

Supports three delay strategies:

  1. Fixed delay: Always wait the same amount of time

  2. Random delay: Wait a random time between min and max

  3. Custom delay: Use a callable that returns delay duration

Parameters:
  • delay – Fixed delay in seconds, or None if using delay_func

  • min_delay – Minimum delay for random strategy (requires max_delay)

  • max_delay – Maximum delay for random strategy (requires min_delay)

  • delay_func – Custom callable that returns delay in seconds. Called with (request, spider) and should return a float.

Example:

# Fixed delay of 1 second
DelayMiddleware(delay=1.0)

# Random delay between 0.5 and 2 seconds
DelayMiddleware(min_delay=0.5, max_delay=2.0)

# Custom delay function
def my_delay(request, spider):
    return 1.0 if "fast" in request.url else 2.0

DelayMiddleware(delay_func=my_delay)
__init__(delay=None, min_delay=None, max_delay=None, delay_func=None)[source]
Parameters:
Return type:

None

async process_request(request, spider)[source]

Calculate and apply delay before processing the request.

Parameters:
Return type:

Request

class silkworm.middlewares.ExceptionMiddleware[source]

Bases: Protocol

Protocol for middleware that may recover from request failures.

async process_exception(request, exception, spider)[source]

Return a retry request, or None to leave the error unhandled.

Parameters:
Return type:

Request | None

__init__(*args, **kwargs)
class silkworm.middlewares.ProxyMiddleware[source]

Bases: object

Assign proxies to requests and rotate away from failed proxies.

Parameters:
  • proxies – Proxy URLs loaded directly.

  • proxy_file – UTF-8 file containing one proxy URL per non-empty line.

  • random_selection – Choose randomly instead of round-robin.

Exactly one proxy source is required. A pre-existing string in request.meta["proxy"] takes precedence. Register the same instance as exception middleware to retry failures through unused proxies.

__init__(proxies=None, proxy_file=None, random_selection=False)[source]
Parameters:
  • proxies (Iterable[str] | None)

  • proxy_file (str | Path | None)

  • random_selection (bool)

Return type:

None

async process_request(request, spider)[source]

Preserve an explicit proxy or assign the next configured proxy.

Parameters:
Return type:

Request

async process_exception(request, exception, spider)[source]

Retry with an unused proxy, or return None when none remain.

Parameters:
Return type:

Request | None

class silkworm.middlewares.RequestMiddleware[source]

Bases: Protocol

Protocol for middleware applied before an HTTP request is sent.

async process_request(request, spider)[source]

Return the request to send, optionally modified or replaced.

Parameters:
Return type:

Request

__init__(*args, **kwargs)
class silkworm.middlewares.RequestResponseStreamMiddleware[source]

Bases: object

Stream request/response telemetry to a remote HTTP endpoint.

Each outbound request receives a unique exchange_id. The middleware emits a request event before fetch and a response event after fetch, both carrying the same identifier so downstream analysis can join them.

Use the same middleware instance in both request_middlewares and response_middlewares so a single sender queue can stream the full exchange lifecycle:

stream = RequestResponseStreamMiddleware(
    "https://collector.example.com/events",
    auth_token="secret-token",
    batch_size=50,
)

run_spider(
    MySpider,
    request_middlewares=[stream],
    response_middlewares=[stream],
)
Parameters:
  • url – Remote event collector endpoint.

  • method – HTTP method used for delivery.

  • headers – Collector request headers.

  • timeout – Delivery timeout.

  • max_body_bytes – Maximum request or response body bytes serialized.

  • queue_size – Bounded in-memory event queue capacity.

  • auth_token – Optional authorization credential.

  • auth_scheme – Authorization scheme prepended to auth_token.

  • batch_size – Events per delivery request.

  • batch_envelope_key – JSON key containing batched events.

__init__(url, *, method='POST', headers=None, timeout=10.0, max_body_bytes=64_000, queue_size=1_000, auth_token=None, auth_scheme='Bearer', batch_size=1, batch_envelope_key='events')[source]
Parameters:
Return type:

None

async open(spider)[source]

Create the sender client, bounded queue, and background task.

Parameters:

spider (Spider)

Return type:

None

async close(spider)[source]

Flush queued events, stop the sender, and close its HTTP client.

Parameters:

spider (Spider)

Return type:

None

async process_request(request, spider)[source]

Assign an exchange ID and enqueue a serialized request event.

Parameters:
Return type:

Request

async process_response(response, spider)[source]

Enqueue a response event linked to its request exchange ID.

Parameters:
Return type:

Response | Request

async process_exception(request, exception, spider)[source]

Enqueue a request-error event without handling the exception.

Parameters:
Return type:

Request | None

class silkworm.middlewares.ResponseMiddleware[source]

Bases: Protocol

Protocol for middleware applied after an HTTP response is received.

async process_response(response, spider)[source]

Return a response to dispatch or a request to enqueue instead.

Parameters:
Return type:

Response | Request

__init__(*args, **kwargs)
class silkworm.middlewares.RetryMiddleware[source]

Bases: object

Retry failed requests with optional exponential backoff.

Retries both responses with selected HTTP statuses (as a response middleware) and transient transport failures such as timeouts and connection resets (as an exception middleware).

Parameters:
  • max_times – Maximum retries after the initial request.

  • retry_http_codes – Status codes that produce a replacement request.

  • backoff_base – Base seconds for base * 2 ** (attempt - 1).

  • sleep_http_codes – Retry statuses that also wait before enqueueing. These codes are automatically added to the retry set.

  • retry_exceptions – Exception types retried with backoff. Defaults to DEFAULT_RETRY_EXCEPTIONS; pass () to retry statuses only.

Attempts are stored in request.meta["retry_times"], shared by status and exception retries, and retry requests bypass deduplication.

__init__(max_times=3, retry_http_codes=None, backoff_base=0.5, sleep_http_codes=None, retry_exceptions=None)[source]
Parameters:
  • max_times (int)

  • retry_http_codes (Iterable[int] | None)

  • backoff_base (float)

  • sleep_http_codes (Iterable[int] | None)

  • retry_exceptions (Iterable[type[BaseException]] | None)

Return type:

None

async process_exception(request, exception, spider)[source]

Return a retry request for transient transport failures.

The retry waits the exponential backoff delay first. Returns None for other exceptions or once max_times retries were made.

Parameters:
Return type:

Request | None

async process_response(response, spider)[source]

Return a retry request for eligible statuses until the limit is met.

Parameters:
Return type:

Response | Request

type silkworm.middlewares.RobotsOrigin = tuple[str, str, int | None]
class silkworm.middlewares.RobotsTxtDelayMiddleware[source]

Bases: object

Request middleware that loads robots.txt and applies its delay directives.

The middleware currently uses Crawl-delay first and falls back to Request-rate when present. Delays are scoped to the origin that served the robots.txt file, and concurrent requests are serialized so the configured spacing is respected under engine concurrency.

Parameters:
  • website_url – Absolute HTTP(S) site URL whose origin is throttled.

  • user_agent – robots.txt group used to resolve directives.

  • fallback_delay – Delay used when no directive exists or loading fails.

  • timeout – robots.txt fetch timeout.

  • ignore_fetch_errors – Apply the fallback instead of propagating errors.

  • fetcher – Optional async robots.txt loader for custom transports or tests.

__init__(website_url, *, user_agent='*', fallback_delay=None, timeout=10.0, ignore_fetch_errors=True, fetcher=None)[source]
Parameters:
  • website_url (str)

  • user_agent (str)

  • fallback_delay (float | None)

  • timeout (float | timedelta | None)

  • ignore_fetch_errors (bool)

  • fetcher (RobotsTxtFetcher | None)

Return type:

None

async open(spider)[source]

Load and parse robots.txt before crawl requests begin.

Parameters:

spider (Spider)

Return type:

None

async process_request(request, spider)[source]

Apply origin-scoped spacing from the loaded robots.txt directives.

Parameters:
Return type:

Request

type silkworm.middlewares.RobotsTxtFetcher = Callable[[str], Awaitable[str]]
class silkworm.middlewares.RobotsTxtMiddleware[source]

Bases: object

Obey robots.txt rules for every site the crawl visits.

Each origin’s /robots.txt is fetched once, on its first request. Disallowed requests are dropped with IgnoreRequest (counted as ignored_requests with reason robots_txt); with obey_crawl_delay the origin’s Crawl-delay/Request-rate also spaces its requests. Requests for robots.txt itself and requests with meta["dont_obey_robotstxt"] are never blocked.

Following RFC 9309, a 4xx robots.txt response means “no restrictions”. Server errors and unreachable robots.txt files follow on_unavailable: "allow" (the default, logged as a warning) or "disallow" to skip the origin entirely, as the RFC recommends.

Parameters:
  • user_agent – Product token matched against robots.txt groups.

  • obey_crawl_delay – Apply crawl delays per origin.

  • on_unavailable – Policy when robots.txt cannot be fetched.

  • timeout – robots.txt fetch timeout.

  • fetcher – Optional async loader returning robots.txt text; raise an HttpError subclass with a status attribute for HTTP errors.

__init__(*, user_agent='*', obey_crawl_delay=True, on_unavailable='allow', timeout=10.0, fetcher=None)[source]
Parameters:
  • user_agent (str)

  • obey_crawl_delay (bool)

  • on_unavailable (Literal['allow', 'disallow'])

  • timeout (float | timedelta | None)

  • fetcher (RobotsTxtFetcher | None)

Return type:

None

async process_request(request, spider)[source]

Drop requests robots.txt disallows and apply crawl delays.

Parameters:
Return type:

Request

class silkworm.middlewares.SkipNonHTMLMiddleware[source]

Bases: object

Response middleware that drops callbacks for non-HTML payloads.

It checks the Content-Type header first, then falls back to a quick body sniff for “<html”. Non-HTML responses keep flowing through the engine but execute a no-op callback so spider parse methods are skipped. Set request.meta[“allow_non_html”] = True to bypass filtering for a request (useful for XML sitemaps, robots.txt fetches, etc.).

Parameters:
  • allowed_types – Lowercase tokens accepted in the Content-Type header.

  • sniff_bytes – Leading body bytes inspected for an HTML tag when headers are inconclusive.

__init__(allowed_types=None, sniff_bytes=2048)[source]
Parameters:
  • allowed_types (Iterable[str] | None)

  • sniff_bytes (int)

Return type:

None

async process_response(response, spider)[source]

Replace the callback with a no-op when the payload is not HTML.

Parameters:
Return type:

Response | Request

class silkworm.middlewares.UserAgentMiddleware[source]

Bases: object

Set a missing User-Agent header from a pool or fixed default.

Parameters:
  • user_agents – Values sampled independently for each request.

  • default – Value used when the pool is empty. Defaults to "silkworm/0.1".

Existing request headers are never overwritten.

__init__(user_agents=None, *, default=None)[source]
Parameters:
  • user_agents (Sequence[str] | None)

  • default (str | None)

Return type:

None

async process_request(request, spider)[source]

Set a user agent if the request does not already define one.

Parameters:
Return type:

Request