Middlewares¶
Middlewares let you intercept requests and responses. The engine applies them in order. See src/silkworm/middlewares.py.
Interfaces¶
Middlewares implement protocol-style async methods:
class RequestMiddleware:
async def process_request(self, request, spider) -> Request: ...
class ResponseMiddleware:
async def process_response(self, response, spider) -> Response | Request: ...
class ExceptionMiddleware:
async def process_exception(self, request, exception, spider) -> Request | None: ...
Order of execution:
Request middlewares (before HTTP fetch)
Response middlewares (after HTTP fetch)
Callback (
parseor custom callback)
If request processing fails, the engine calls process_exception on middleware
instances from both middleware lists, deduplicating shared instances. Returning a
Request schedules a retry; returning None leaves the exception for the next
exception middleware or the request’s errback.
run_spider(
MySpider,
request_middlewares=[UserAgentMiddleware(), DelayMiddleware(delay=0.5)],
response_middlewares=[RetryMiddleware(max_times=3), SkipNonHTMLMiddleware()],
)
Built-in Middlewares¶
UserAgentMiddleware¶
Picks a random user agent from a list or uses the default
silkworm/0.1.
UserAgentMiddleware(user_agents=["UA1", "UA2"], default="silkworm/0.1")
ProxyMiddleware¶
Rotates proxies (round-robin or random).
Reads from a list or file.
Writes
request.meta["proxy"]for the HTTP client.Retries fetch exceptions with another proxy when one is available, preserving failed proxies in request metadata.
ProxyMiddleware(proxies=["http://proxy1:8080", "http://proxy2:8080"])
ProxyMiddleware(proxy_file="proxies.txt", random_selection=True)
DelayMiddleware¶
Fixed, random range, or custom delay function.
Uses
asyncio.sleep(non-blocking).
DelayMiddleware(delay=1.0)
DelayMiddleware(min_delay=0.3, max_delay=1.0)
def custom_delay(request, spider) -> float:
return 0.5
DelayMiddleware(delay_func=custom_delay)
RobotsTxtMiddleware¶
Obeys robots.txt for every origin the crawl visits: each
/robots.txtis fetched once, on the origin’s first request.Disallowed requests raise
IgnoreRequest(counted asignored_requestswith reasonrobots_txt); they are not errors.Applies
Crawl-delay(decimal values included) orRequest-rateper origin unlessobey_crawl_delay=False.A 4xx robots.txt means “no restrictions”. When robots.txt cannot be fetched,
on_unavailable="allow"(default) crawls anyway with a warning and"disallow"skips the origin.meta["dont_obey_robotstxt"] = Trueexempts a request.
from silkworm.middlewares import RobotsTxtMiddleware
run_spider(MySpider, request_middlewares=[RobotsTxtMiddleware(user_agent="silkbot")])
AutoThrottleMiddleware¶
Spaces requests per host and adapts the delay to the host’s latency (
latency / target_concurrency, averaged with the previous delay).Throttling statuses (429/503 by default) double the delay and honour
Retry-After; error responses never make crawling faster.Options:
start_delay,min_delay,max_delay,target_concurrency,backoff_statuses.Register the same instance as a request and a response middleware.
from silkworm.middlewares import AutoThrottleMiddleware
throttle = AutoThrottleMiddleware(start_delay=1.0, max_delay=30.0)
run_spider(MySpider, request_middlewares=[throttle], response_middlewares=[throttle])
RobotsTxtDelayMiddleware¶
Downloads
robots.txtfrom the provided website origin and applies delay settings from that file (delays only; useRobotsTxtMiddlewareto also obeyDisallowrules).Uses
Crawl-delayfirst. If it is absent, usesRequest-rateasseconds / requests.Applies only to requests for the same scheme/host/port as the robots.txt origin.
Serializes same-origin requests with an internal async lock so engine concurrency cannot bypass the robots delay.
Fetches robots.txt during middleware
open; if used without the engine lifecycle, it loads lazily on the first request.
from silkworm import RobotsTxtDelayMiddleware, run_spider
run_spider(
MySpider,
request_middlewares=[
RobotsTxtDelayMiddleware(
"https://example.com",
user_agent="silkworm",
fallback_delay=1.0,
timeout=10.0,
)
],
)
Constructor options:
website_url: absolutehttporhttpssite URL; Silkworm fetches/robots.txtfor that origin.user_agent: user-agent token used when readingCrawl-delayandRequest-rate; defaults to"*".fallback_delay: optional delay to apply when robots.txt has no delay directive or cannot be fetched.timeout: robots.txt fetch timeout in seconds ortimedelta; defaults to10.0.ignore_fetch_errors: whenTrue(default), fetch failures fall back tofallback_delay; whenFalse, the fetch error is raised.fetcher: optionalRobotsTxtFetcher(async (robots_url: str) -> str) that loads robots.txt text, for custom transports or tests.
silkworm.middlewares also exports the RobotsTxtFetcher and RobotsOrigin ((scheme, host, port)) type aliases.
RetryMiddleware¶
Retries on HTTP codes in
retry_http_codes(defaults: 500, 502, 503, 504, 522, 524, 408, 429).Also retries transient transport failures through its
process_exceptionhook:HttpTimeoutError,HttpConnectionError(refused, reset, or dropped connections), and built-inTimeoutError/ConnectionErrorby default. Passretry_exceptions=()to retry statuses only. Callback errors and other failures are never retried.max_timescaps retries after the initial request (default 3).Codes in
sleep_http_codeswaitbackoff_base * 2 ** (attempt - 1)seconds (non-blocking) before re-enqueueing. By default every retry code sleeps; codes listed only insleep_http_codesare added to the retry set automatically.Uses
request.meta["retry_times"](shared by status and exception retries), setsdont_filter=Trueon retries, and counts each retry in theretriesstatistic.
RetryMiddleware(max_times=3, backoff_base=0.5, sleep_http_codes=[429, 503])
SkipNonHTMLMiddleware¶
Skips callbacks for non-HTML responses unless
allow_non_htmlis set in request meta.Checks content-type and optional body sniff.
SkipNonHTMLMiddleware(allowed_types=["html"], sniff_bytes=2048)
RequestResponseStreamMiddleware¶
Streams paired
requestandresponsetelemetry events to a collector endpoint.Use the same instance in
request_middlewaresandresponse_middlewares.Adds internal exchange IDs to request metadata so downstream systems can join events.
Supports authorization headers, bounded sender queue, body truncation, batching, and
open/closelifecycle flushing.Options:
url,method(default"POST"),headers,timeout(default 10s),max_body_bytes(default 64,000),queue_size(default 1,000),auth_token,auth_scheme(default"Bearer"),batch_size(default 1), andbatch_envelope_key(JSON key wrapping batched events, default"events").
from silkworm import RequestResponseStreamMiddleware
stream = RequestResponseStreamMiddleware(
"https://collector.example.com/events",
auth_token="secret-token",
batch_size=50,
max_body_bytes=8_192,
)
run_spider(
MySpider,
request_middlewares=[stream],
response_middlewares=[stream],
)
CloudflareCrawlMiddleware¶
Routes opt-in requests through Cloudflare Browser Rendering’s crawl API.
Enable per request with
request.meta["cloudflare_crawl"] = Trueor pass a dict of per-request crawl options.The callback receives a synthetic JSON
Responsecontaining the final Cloudflare API payload.Requires Cloudflare account credentials; there is no package extra for this middleware.
Options:
account_id,api_token,crawl_options(defaults merged into every crawl submission),api_base_url,poll_interval(seconds between job polls, default 1.0),timeout(overall job timeout, default 300s), andapi_timeout(per API call, default 30s).Example: examples/cloudflare_crawl_spider.py
from silkworm import Request
from silkworm.middlewares import CloudflareCrawlMiddleware
await self.follow(
Request(
url="https://example.com/",
callback=self.parse,
meta={"cloudflare_crawl": {"limit": 25, "render": True}},
)
)
run_spider(
MySpider,
request_middlewares=[
CloudflareCrawlMiddleware(
account_id="...",
api_token="...",
timeout=300,
)
],
)
Custom Middleware Example¶
from silkworm.request import Request
class AddHeaderMiddleware:
async def process_request(self, request: Request, spider):
headers = {**request.headers}
headers.setdefault("x-trace", "1")
return request.replace(headers=headers)
from silkworm.request import Request
class RetryOnceMiddleware:
async def process_request(self, request: Request, spider):
return request
async def process_exception(
self,
request: Request,
exception: Exception,
spider,
) -> Request | None:
if request.meta.get("retried"):
return None
return request.replace(meta={**request.meta, "retried": True}, dont_filter=True)