silkworm¶
Public entry point for building and running asynchronous web spiders.
The package root exposes the core request, response, spider, engine, runner,
browser-backed client, middleware, logging, and HTML-to-Markdown APIs most
applications need. Specialized middleware and pipeline implementations are
available from silkworm.middlewares and silkworm.pipelines.
Spiders¶
- class silkworm.Spider[source]¶
Bases:
objectBase class defining a crawl’s initial requests and response callback.
Subclasses usually set
nameandstart_urls, then overrideparse(). Callbacks areasyncfunctions returningNone; they report results by awaitingemit()for items andfollow()for further requests. Both apply backpressure, so eachawaitcompletes once the item passed every pipeline or the request entered the queue.- Parameters:
name – Per-instance name overriding the class attribute.
start_urls – Per-instance starting URLs.
allowed_domains – Per-instance domains the crawl may visit.
custom_settings – JSON-compatible settings copied for this instance. Keys naming engine options (for example
concurrencyormax_depth) configure runs started through the runners and CLI.logger – Existing structured logger or context mapping used to create one lazily.
- stats_payload¶
User-defined statistics included in engine summaries.
Example
>>> from silkworm import HTMLResponse, Response, Spider >>> class TitlesSpider(Spider): ... name = "titles" ... start_urls = ("https://example.com",) ... ... async def parse(self, response: Response) -> None: ... if isinstance(response, HTMLResponse): ... title = await response.select_first("title") ... if title is not None: ... await self.emit({"title": title.text})
- __init__(*, name=None, start_urls=None, allowed_domains=None, custom_settings=None, logger=None)[source]¶
- allowed_domains: tuple[str, ...] = ()¶
Domains (and their subdomains) the crawl may visit; requests to other hosts are dropped unless they set
dont_filter. Empty allows any host.
- async start_requests()[source]¶
Schedule one request per
start_urlsentry.Override this hook to customize methods, headers, metadata, or callbacks for initial requests; schedule each one with
follow().- Return type:
None
- async parse(response)[source]¶
Process a starting response.
Subclasses must implement this coroutine and report results with
emit()andfollow(). The engine always passes anHTMLResponseto this callback.- Raises:
NotImplementedError – When the base implementation is called.
- Parameters:
response (Response)
- Return type:
None
- async emit(item)[source]¶
Send a scraped item through the item pipelines.
With the default per-item processing, returns after every pipeline has processed the item. When engine batching is enabled, returns after the item is queued; pending batches drain and surface errors before the callback completes.
- Raises:
SpiderError – If awaited outside an engine-run callback, or after the callback that owns the current scope finished.
- Parameters:
item (JSONLike)
- Return type:
None
- async follow(target, callback=None, **kwargs)[source]¶
Schedule a request for crawling.
targetis either a readyRequestor a URL. Inside a response callback a URL is resolved against the response and inherits its callback, exactly likesilkworm.Response.follow(). Requests pass engine deduplication and wait for queue capacity.- Parameters:
- Raises:
SpiderError – If awaited outside an engine-run callback, or after the callback that owns the current scope finished.
TypeError – If request fields are combined with a
Requesttarget.
- Return type:
None
- async follow_all(targets, callback=None, **kwargs)[source]¶
Schedule every non-
Nonetarget in input order viafollow().
Requests and Responses¶
- class silkworm.Request[source]¶
Bases:
objectDescribe one HTTP request and how its result should be handled.
- Parameters:
url – Absolute request URL. Relative links can be resolved with
silkworm.Response.url_join()or scheduled directly withsilkworm.Response.follow().method – HTTP method name.
headers – Per-request headers merged over client defaults.
params – Query parameters merged with any query already in
url.data – Form, byte, or text request body.
json – JSON-compatible request body. Do not combine with
data.meta – Framework and user metadata propagated with the request.
timeout – Per-request timeout in seconds or as a
timedelta.callback – Async callable that consumes the response and reports results with
silkworm.Spider.emit()andsilkworm.Spider.follow(). When omitted on a followed request, the parent request’s callback is reused.errback – Async callable invoked when request processing fails. It may also emit items and follow requests.
dont_filter – Bypass engine request deduplication when true.
priority – Scheduling priority; larger values are processed first.
Note
Built-in metadata keys include
proxy,retry_times,allow_non_html, andredirect_times. Applications may store additional JSON-compatible values alongside them.- replace(**kwargs)[source]¶
Return a shallow copy with the named dataclass fields replaced.
Unspecified mutable fields are shared with the original request.
- __init__(url, method='GET', headers=<factory>, params=<factory>, data=None, json=None, meta=<factory>, timeout=None, callback=None, errback=None, dont_filter=False, priority=0)¶
- class silkworm.Response[source]¶
Bases:
objectRepresent a completed HTTP response and its originating request.
- Parameters:
url – Final URL after redirects.
status – Integer HTTP status code.
headers – Normalized response headers with lowercase names.
body – Raw response payload.
request – Request that produced this response.
The decoded
textand detectedencodingare computed lazily and cached. Callclose()when retaining response objects to release their payload memory early.- property encoding: str¶
Return the normalized encoding used by
text.Detection checks byte-order marks, HTTP headers, HTML/XML declarations, and finally
charset-normbefore falling back to UTF-8 with replacement characters.
- async follow(href, callback=None, **kwargs)[source]¶
Schedule a request for a link relative to this response.
Must be awaited inside a spider callback, errback, or
start_requests(). The request passes through engine deduplication and waits for queue capacity (backpressure).- Parameters:
- Raises:
SpiderError – If awaited outside an engine-run callback.
- Return type:
None
- async follow_all(hrefs, callback=None, **kwargs)[source]¶
Schedule requests for every non-
Nonelink inhrefs.The callback and extra request fields are applied to every request, which are scheduled in input order.
- class silkworm.HTMLResponse[source]¶
Bases:
ResponseHTTP response with lazy asynchronous HTML selector helpers.
- Parameters:
doc_max_size_bytes – Maximum source size accepted by
scraper-rswhen the document is first parsed.
The parsed document is created once and shared by selector calls. Selector and parsing failures are normalized to
SelectorError.- async select(selector)[source]¶
Return every element matching a CSS selector.
- Raises:
SelectorError – If the document cannot be parsed or the selector is invalid.
- Parameters:
selector (str)
- Return type:
list[AsyncElement]
- async select_first(selector)[source]¶
Return the first CSS match, or
Nonewhen no element matches.- Raises:
SelectorError – If parsing or selector evaluation fails.
- Parameters:
selector (str)
- Return type:
AsyncElement | None
- async find(selector)[source]¶
Return the first CSS match as an alias for
select_first().- Parameters:
selector (str)
- Return type:
AsyncElement | None
- async css_first(selector)[source]¶
Return the first CSS match as an alias for
select_first().- Parameters:
selector (str)
- Return type:
AsyncElement | None
- async xpath(xpath)[source]¶
Return every element matching an XPath expression.
- Raises:
SelectorError – If parsing or XPath evaluation fails.
- Parameters:
xpath (str)
- Return type:
list[AsyncElement]
- async xpath_first(xpath)[source]¶
Return the first XPath match, or
Nonewhen no element matches.- Raises:
SelectorError – If parsing or XPath evaluation fails.
- Parameters:
xpath (str)
- Return type:
AsyncElement | None
- async prettify()[source]¶
Return the parsed document formatted as readable HTML.
- Raises:
SelectorError – If parsing or serialization fails.
- Return type:
- async to_markdown(*, mode='full', options=None)[source]¶
Convert the decoded HTML body to Markdown off the event-loop thread.
- Parameters:
mode (MarkdownMode) –
fast-h2mconversion strategy.options (MarkdownOptions | None) – Additional converter options.
- Returns:
Converted Markdown text.
- Return type:
- async to_markdown_result(*, mode='full', options=None)[source]¶
Return the structured
fast-h2mconversion result.- Parameters:
mode (MarkdownMode) –
fast-h2mconversion strategy.options (MarkdownOptions | None) – Additional converter options.
- Return type:
Engine¶
- class silkworm.Engine[source]¶
Bases:
objectCoordinate request scheduling, HTTP I/O, callbacks, and item pipelines.
- Parameters:
spider – Spider instance to execute.
concurrency – Maximum simultaneous HTTP requests for the default client.
max_pending_requests – Queue capacity used for backpressure. Defaults to ten times the effective HTTP client concurrency.
start_requests()waits while the queue is full; callbacks wait only when they are the sole producer, otherwise they enqueue past the bound so workers keep crawling and can never deadlock.emulation – Browser profile used by the default
wreqclient; passNoneto disable impersonation.request_timeout – Default per-request timeout for the default client: 60 seconds (
DEFAULT_REQUEST_TIMEOUT) unless given;Nonedisables it.Request.timeoutoverrides it per request. Ignored whenhttp_clientis supplied.html_max_size_bytes – Maximum document size parsed by HTML responses.
max_response_size_bytes – Largest response body the default client downloads (
Nonefor no limit); larger bodies fail withResponseTooLargeError.request_middlewares – Request processors applied in list order.
response_middlewares – Response processors applied in list order.
item_pipelines – Item processors applied in list order, passing each returned value to the next pipeline.
item_batch_size – Number of emitted items processed together. The default
1preserves immediate per-item processing.item_batch_wait – Maximum seconds to wait for a partial batch when
item_batch_sizeis greater than one.log_stats_interval – Seconds between statistics messages, or
Noneto disable periodic summaries.keep_alive – Request connection reuse from the default client when supported by
wreq.http_client – Preconfigured client replacing the default client.
engine_logger – Event logger customization.
dedup_key – Function mapping a request to its deduplication key. Defaults to
default_dedup_key().concurrency_per_domain – Maximum simultaneous fetches per host, or
Nonefor no per-host limit.max_depth – Drop requests more than this many links away from a start request (start requests have depth
0).max_requests – Stop after sending this many requests.
max_items – Stop after this many items passed every pipeline; later items are dropped.
max_errors – Stop after this many unrecovered failures.
max_duration – Stop after this much wall-clock time.
max_error_rate – Fail the crawl when
errors / requests_sentexceeds this fraction.min_items – Fail the crawl when fewer items were scraped.
max_item_drop_rate – Fail the crawl when the share of items dropped by pipelines exceeds this fraction.
job_dir – Directory persisting the seen-set and unfinished requests so an interrupted crawl resumes where it stopped.
http_cache – Serve and store responses through this on-disk cache.
metrics_port – Serve Prometheus metrics at
/metricson this port while crawling (0picks a free port).metrics_host – Interface for the metrics server.
Requests with
dont_filterbypass deduplication and off-site filtering. Higher request priorities are dequeued before lower ones, while insertion order breaks ties. The stop limits end the crawl gracefully: pending requests are discarded (or kept injob_dir), in-flight requests finish, andrun()reports the limit as the close reason. Failure-policy violations makerun()raiseCrawlFailedError; they are not evaluated for crawls stopped withstop().- __init__(spider, *, concurrency=16, max_pending_requests=None, emulation=DEFAULT_EMULATION, request_timeout=DEFAULT_REQUEST_TIMEOUT, html_max_size_bytes=5_000_000, max_response_size_bytes=DEFAULT_MAX_RESPONSE_SIZE_BYTES, request_middlewares=None, response_middlewares=None, item_pipelines=None, item_batch_size=1, item_batch_wait=0.05, log_stats_interval=None, keep_alive=False, http_client=None, engine_logger=None, dedup_key=None, concurrency_per_domain=None, max_depth=None, max_requests=None, max_items=None, max_errors=None, max_duration=None, max_error_rate=None, min_items=None, max_item_drop_rate=None, job_dir=None, http_cache=None, metrics_port=None, metrics_host='127.0.0.1')[source]¶
- Parameters:
spider (Spider)
concurrency (int)
max_pending_requests (int | None)
emulation (Emulation | Profile | None)
request_timeout (float | timedelta | None)
html_max_size_bytes (int)
max_response_size_bytes (int | None)
request_middlewares (Iterable[RequestMiddleware] | None)
response_middlewares (Iterable[ResponseMiddleware] | None)
item_pipelines (Iterable[ItemPipeline] | None)
item_batch_size (int)
item_batch_wait (float)
log_stats_interval (float | None)
keep_alive (bool)
http_client (FetchClient | None)
engine_logger (EngineLogger | None)
dedup_key (DedupKey | None)
concurrency_per_domain (int | None)
max_depth (int | None)
max_requests (int | None)
max_items (int | None)
max_errors (int | None)
max_duration (float | timedelta | None)
max_error_rate (float | None)
min_items (int | None)
max_item_drop_rate (float | None)
job_dir (str | os.PathLike[str] | None)
http_cache (HttpCache | None)
metrics_port (int | None)
metrics_host (str)
- Return type:
None
- property close_reason: str | None¶
Return why the crawl is stopping, or
Nonewhile it runs normally.
- stop(reason='shutdown')[source]¶
Stop the crawl gracefully.
New requests are no longer scheduled and queued ones are discarded (they stay saved when a job directory is configured, so the crawl can resume). Requests already being processed finish, pipelines close normally, and
run()returns withreasonas the close reason. Calling it again has no effect.- Parameters:
reason (str)
- Return type:
None
- async open_spider()[source]¶
Open middleware, spider, and pipelines, then enqueue initial requests.
Middleware opens before the spider; pipelines open afterward in their configured order. When resuming a job, saved requests are queued before
start_requests()runs (already-seen start requests are skipped).- Return type:
None
- async close_spider()[source]¶
Close pipelines, the spider, and middleware lifecycle hooks.
Components close in reverse startup order. Middleware instances close once even when registered for both request and response work.
- Return type:
None
- metrics_text()[source]¶
Return current statistics in the Prometheus text exposition format.
- Return type:
- async run()[source]¶
Run the crawl until the queue drains or a stop condition, then clean up.
Worker tasks, periodic statistics, and lifecycle hooks are managed as a task group. The HTTP client, spider components, job state, and metrics server are always closed, and a final statistics record is emitted.
- Returns:
The crawl’s
CrawlResult.- Raises:
CrawlFailedError – If the crawl violated its failure policy (
max_error_rate,min_items,max_item_drop_rate).Exception – Errors from lifecycle hooks, pipelines’
open/close, or cleanup propagate after structured error logging.
- Return type:
- class silkworm.EngineOptions[source]¶
Bases:
TypedDictKeyword options for
Engine, also accepted by every runner.run_spider(MySpider, concurrency=32, request_timeout=10)forwards these toEngine; omitted keys use theEnginedefaults. Supplyhttp_clientto inject a compatible client; its concurrency then controls worker count and default queue capacity.
- class silkworm.EngineLogger[source]¶
Bases:
objectCustomizable engine event logger.
Subclass this when you need to redact or reshape selected engine log events. Set an event level to
Noneto suppress that event.- fetched_response(logger, request, response, spider)[source]¶
Log a completed response with status and optional request URL.
- retrying_request(logger, request, spider, *, source)[source]¶
Log that a middleware or response requested another attempt.
- running_item_pipeline(logger, pipeline, spider)[source]¶
Log item dispatch using the pipeline’s effective log level.
- Parameters:
logger (Logger)
pipeline (ItemPipeline)
spider (Spider)
- Return type:
None
- __init__(fetched_response_level='INFO', fetching_request_level='DEBUG', item_pipeline_level='DEBUG', retry_request_level='DEBUG', include_request_url=True)¶
- silkworm.default_dedup_key(req)[source]¶
Return the engine’s default deduplication key for
req.This is
request_fingerprint(): the HTTP method, the canonical URL (normalized case, default port, sorted query, no fragment) withparamsmerged in, and the request body. Headers and metadata do not affect it.
- silkworm.request_fingerprint(request)[source]¶
Return a stable hex digest identifying what
requestfetches.The fingerprint covers the HTTP method, the canonical URL with
paramsmerged in, and the request body (form data or JSON). Headers, metadata, callbacks, and URL fragments are ignored, so two requests with the same fingerprint are treated as duplicates by the engine’s default deduplicator and share HTTP cache entries.
- silkworm.canonicalize_url(url, *, keep_fragments=False)[source]¶
Return a normalized form of
urlfor deduplication and caching.The scheme and host are lowercased, default ports and (unless
keep_fragments) fragments are removed, percent-encoding is normalized, an empty path becomes/, and query parameters are sorted while keeping blank values and repeated keys. Two URLs that address the same resource in these respects map to the same string.Example
>>> canonicalize_url("HTTP://Example.com:80/a?b=2&a=1#top") 'http://example.com/a?a=1&b=2'
Runners¶
- async silkworm.crawl(spider, *, handle_signals=False, **options)[source]¶
Run
spiderto completion on the current event loop.Options are resolved through
silkworm.settings.resolve_options(), soSILKWORM_*environment variables and the spider’scustom_settingsapply beneath the options passed here.- Parameters:
spider (Spider | type[Spider]) – Spider instance, or a no-argument spider class.
handle_signals (bool) – Stop gracefully on the first SIGINT/SIGTERM and cancel on the second. Off by default because the calling application owns the event loop and may handle signals itself.
**options (Unpack[EngineOptions]) –
EngineOptionsforwarded to the engine.
- Returns:
The crawl’s
CrawlResult.- Raises:
CrawlFailedError – If the crawl violated its failure policy.
KeyboardInterrupt – If a second shutdown signal forced cancellation.
- Return type:
Use this coroutine when the application already owns an event loop; use a
run_spider*function from synchronous code.
- silkworm.run_spider(spider, *, loop_factory=None, handle_signals=True, **options)[source]¶
Run
spiderwithasyncio, blocking until the crawl finishes.The first SIGINT (Ctrl+C) or SIGTERM stops the crawl gracefully: pending requests are discarded (or saved with
job_dir), in-flight requests finish, and pipelines close normally. A second signal cancels immediately.- Parameters:
spider (Spider | type[Spider]) – Spider instance, or a spider class to instantiate without arguments.
loop_factory (LoopFactory | None) – Optional event loop factory, e.g. from uvloop.
handle_signals (bool) – Install the graceful shutdown handlers described above.
**options (Unpack[EngineOptions]) – Engine options; see
EngineOptions.
- Returns:
The crawl’s
CrawlResult.- Raises:
CrawlFailedError – If the crawl violated its failure policy.
- Return type:
- silkworm.run_spider_rsloop(spider, **options)[source]¶
Run
spideron an rsloop event loop (pip install silkworm-rs[rsloop]).- Raises:
ImportError – If rsloop is not installed.
- Parameters:
options (Unpack[EngineOptions])
- Return type:
- silkworm.run_spider_uvloop(spider, **options)[source]¶
Run
spideron a uvloop event loop (pip install silkworm-rs[uvloop]).- Raises:
ImportError – If uvloop is not installed.
- Parameters:
options (Unpack[EngineOptions])
- Return type:
- silkworm.run_spider_winloop(spider, **options)[source]¶
Run
spideron a winloop event loop, optimized for Windows (pip install silkworm-rs[winloop]).- Raises:
ImportError – If winloop is not installed.
- Parameters:
options (Unpack[EngineOptions])
- Return type:
- silkworm.run_spider_trio(spider, **options)[source]¶
Run
spiderwith trio as the async backend (pip install silkworm-rs[trio]).The engine uses asyncio primitives, so it runs inside trio via trio-asyncio. This runner is currently available on Python 3.13 only because trio-asyncio 0.16 is incompatible with Python 3.14 and newer.
- Raises:
ImportError – If trio or trio-asyncio is not installed.
- Parameters:
options (Unpack[EngineOptions])
- Return type:
Crawl Results¶
- class silkworm.CrawlResult[source]¶
Bases:
objectOutcome of a finished crawl, returned by
Engine.run()and runners.- close_reason¶
Why the crawl ended:
"finished"when the queue drained,"shutdown"after a stop request or signal, a limit name such as"max_items", or the reason passed toCloseSpider.- Type:
- labeled_stats¶
Per-label breakdowns (see
BASE_LABELED_COUNTERS).
- __init__(spider, close_reason, elapsed_seconds, stats, labeled_stats, custom_stats=<factory>, failures=())¶
Convenience Helpers¶
- async silkworm.fetch_html(url, *, emulation=DEFAULT_EMULATION, timeout=DEFAULT_REQUEST_TIMEOUT)[source]¶
Fetch and asynchronously parse one HTML document with
wreq.- Parameters:
url (str) – Absolute URL to fetch.
emulation (Emulation | Profile | None) – Browser profile to impersonate, or
Noneto disable it.timeout (float | timedelta | None) – Request timeout in seconds or as a
timedelta; 60 seconds (DEFAULT_REQUEST_TIMEOUT) unless given, andNonedisables it.
- Returns:
A
(text, AsyncDocument)tuple with awaitable selector helpers.- Return type:
Note
This convenience API does not apply spider middleware, retries, deduplication, or pipelines.
- async silkworm.fetch_html_cdp(url, *, ws_endpoint='ws://127.0.0.1:9222', timeout=None)[source]¶
Fetch HTML from a URL using CDP (Chrome DevTools Protocol).
This function connects to a CDP-compatible browser (like Lightpanda, Chrome, or Chromium) and fetches the rendered HTML after JavaScript execution.
- Parameters:
- Returns:
A tuple of (text, AsyncDocument) with awaitable selector helpers.
- Raises:
ImportError – If websockets package is not installed
HttpError – If the request fails
- Return type:
Example
>>> import asyncio >>> from silkworm import fetch_html_cdp >>> >>> async def main(): ... text, doc = await fetch_html_cdp("https://example.com") ... title = await doc.select_first("title") ... print(title.text if title else "No title") >>> >>> asyncio.run(main())
- async silkworm.fetch_html_servo(url, *, timeout=None, settle_ms=0, user_agent=None, javascript=None, allow_private_addresses=False)[source]¶
Fetch and parse rendered HTML with
servofetchand Servo.- Parameters:
url (str) – Absolute URL to render.
timeout (float | timedelta | None) – Render timeout in seconds or as a
timedelta.settle_ms (int) – Delay after loading before capturing the document.
user_agent (str | None) – Optional browser user agent override.
javascript (str | None) – Optional JavaScript evaluated by the rendered-page client.
allow_private_addresses (bool) – Permit navigation to private network addresses.
- Returns:
A
(text, AsyncDocument)tuple with awaitable selector helpers.- Raises:
ImportError – If a compatible
servofetchbuild is unavailable.HttpError – If rendering fails.
- Return type:
Client Adapters¶
- class silkworm.http.HttpClient[source]¶
Bases:
objectSend Silkworm requests through a browser-impersonating
wreqclient.- Parameters:
concurrency – Maximum requests in flight.
emulation – Browser profile to impersonate, or
Noneto disable it.default_headers – Headers merged below per-request headers.
timeout – Default request timeout in seconds or as a
timedelta(DEFAULT_REQUEST_TIMEOUT, 60 seconds, unless given);Nonedisables it.Request.timeoutoverrides it per request. The budget covers sending the request and downloading the whole body, restarts for each redirect hop, and excludes time spent waiting for a concurrency slot.html_max_size_bytes – Maximum document size parsed by HTML selectors.
follow_redirects – Follow redirect responses internally.
max_redirects – Maximum redirect hops.
keep_alive – Request connection reuse when the installed
wreqversion supports it.max_response_size_bytes – Largest body downloaded, or
Nonefor no limit. A largerContent-Lengthfails before the body is read, and streamed bodies stop as soon as they exceed the limit. Override it per request withrequest.meta["max_response_size"].**client_kwargs – Additional options forwarded to
wreq.Client.
Requests are converted to
HTMLResponsewhen headers or a small body sniff indicate HTML; all other payloads becomeResponse.- __init__(*, concurrency=16, emulation=DEFAULT_EMULATION, default_headers=None, timeout=DEFAULT_REQUEST_TIMEOUT, html_max_size_bytes=5_000_000, follow_redirects=True, max_redirects=10, keep_alive=False, max_response_size_bytes=DEFAULT_MAX_RESPONSE_SIZE_BYTES, **client_kwargs)[source]¶
- Parameters:
- Return type:
None
- async fetch(req)[source]¶
Send one request, follow redirects, and return a normalized response.
The per-request timeout and proxy metadata override client defaults. Synthetic response metadata is honored for middleware integrations.
- Raises:
HttpTimeoutError – If the request times out.
HttpConnectionError – If the connection fails or is reset.
ResponseTooLargeError – If the body exceeds the size limit.
HttpError – If redirects loop or exceed the configured limit, or the request fails for another reason.
- Parameters:
req (Request)
- Return type:
- class silkworm.http.FetchClient[source]¶
Bases:
ProtocolInterface the engine needs from an HTTP client.
HttpClient,CDPClient,ServoFetchClient,OnionLinkClient, andCachingHttpClientimplement it; pass any conforming object ashttp_client.- __init__(*args, **kwargs)¶
- class silkworm.HttpCache[source]¶
Bases:
objectConfiguration and storage for cached HTTP responses.
- Parameters:
directory – Cache directory; created when missing.
expiration – Maximum entry age (seconds or
timedelta);Nonekeeps entries forever.ignore_statuses – Response statuses that are never stored (e.g. 5xx).
methods – HTTP methods eligible for caching.
Set
request.meta["dont_cache"] = Trueto bypass the cache for one request.- __init__(directory, *, expiration=None, ignore_statuses=(500, 502, 503, 504, 522, 524, 408, 429), methods=('GET', 'HEAD'))[source]¶
- wrap(client)[source]¶
Return a client that serves
client’s responses from this cache.- Parameters:
client (FetchClient)
- Return type:
- silkworm.http.DEFAULT_REQUEST_TIMEOUT = 60.0¶
Convert a string or number to a floating-point number, if possible.
Default per-request timeout in seconds (
60.0) forHttpClientand the engine’srequest_timeout.
- silkworm.http.MOCK_RESPONSE_META_KEY¶
str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.
Request
metakey holding a{"status", "headers", "body", "url"}mapping thatHttpClientreturns instead of making a network call.
- class silkworm.CDPClient[source]¶
Bases:
objectFetch rendered pages through a CDP-compatible browser.
- Parameters:
ws_endpoint – Browser WebSocket endpoint. A bare
ws://host:port(orhttp://host:port) also works with Chrome/Chromium, which only accept their per-session/devtools/browser/<id>URL: when the direct connection fails, the URL advertised at/json/versionis used, keeping the host and port given here.concurrency – Maximum simultaneous page fetches.
timeout – Default command and navigation timeout in seconds.
html_max_size_bytes – Maximum rendered document size accepted by the HTML parser and WebSocket transport.
- Raises:
ImportError – If the
cdpextra is not installed.
Example
>>> client = CDPClient(ws_endpoint="ws://127.0.0.1:9222") >>> await client.connect() >>> try: ... response = await client.fetch(request) ... finally: ... await client.close()
- __init__(*, ws_endpoint='ws://127.0.0.1:9222', concurrency=16, timeout=None, html_max_size_bytes=5_000_000)[source]¶
- async connect()[source]¶
Connect to the browser and create an isolated page target.
Calling this method more than once is harmless.
- Raises:
HttpError – If connection or target initialization fails.
- Return type:
None
- class silkworm.ServoFetchClient[source]¶
Bases:
objectRender pages with
servofetchfor use as an engine HTTP client.- Parameters:
concurrency – Maximum simultaneous renders.
timeout – Default render timeout.
settle_ms – Default delay after page load before capture.
user_agent – Default browser user agent.
allow_private_addresses – Permit navigation to private network addresses.
html_max_size_bytes – Maximum rendered document size parsed by selectors.
onion_bootstrap – Optional Tor bootstrap endpoint forwarded to Servo.
onion_consensus_file – Optional cached Tor consensus file.
onion_verbose – Enable verbose Tor integration output.
onion_response_limit – Maximum Tor response size in bytes.
Request metadata can override JavaScript, settle delay, user agent, and screenshot behavior through the exported
SERVO_*_META_KEYconstants.- Raises:
ImportError – If a compatible
servofetchbuild is unavailable.
- __init__(*, concurrency=16, timeout=None, settle_ms=0, user_agent=None, allow_private_addresses=False, html_max_size_bytes=5_000_000, onion_bootstrap=None, onion_consensus_file=None, onion_verbose=False, onion_response_limit=4 * 1024 * 1024)[source]¶
- async fetch(req)[source]¶
Render
req.urland return its HTML response.Per-request timeout and supported Servo metadata override client defaults. Screenshot requests still return the page HTML and expose screenshot metadata through synthetic response headers.
- Raises:
- Parameters:
req (Request)
- Return type:
- silkworm.servo.SERVO_JAVASCRIPT_META_KEY = 'servo_javascript'¶
str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.
- silkworm.servo.SERVO_SETTLE_MS_META_KEY = 'servo_settle_ms'¶
str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.
- silkworm.servo.SERVO_USER_AGENT_META_KEY = 'servo_user_agent'¶
str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.
- silkworm.servo.SERVO_SCREENSHOT_META_KEY = 'servo_screenshot'¶
str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.
- silkworm.servo.SERVO_FULL_PAGE_META_KEY = 'servo_full_page'¶
str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.
- class silkworm.OnionLinkClient[source]¶
Bases:
HttpClientFetch Tor v3 onion services through
onionlink.- Parameters:
concurrency – Maximum simultaneous requests.
default_headers – Headers merged into every request.
timeout – Default request timeout.
html_max_size_bytes – Maximum HTML size accepted by selectors.
follow_redirects – Follow HTTP redirect responses when true.
max_redirects – Maximum redirects before raising an error.
bootstrap – Tor directory authority bootstrap endpoint.
consensus_file – Optional consensus cache path.
verbose – Enable verbose onionlink output.
response_limit – Default maximum response body size in bytes. Override it per request with
onionlink_response_limitmetadata.
- Raises:
ImportError – If a compatible
onionlinkextra is unavailable.ValueError – If
max_redirectsis negative.
- __init__(*, concurrency=16, default_headers=None, timeout=None, html_max_size_bytes=5_000_000, follow_redirects=True, max_redirects=10, bootstrap='128.31.0.39:9131', consensus_file='', verbose=False, response_limit=4 * 1024 * 1024)[source]¶