silkworm

Public entry point for building and running asynchronous web spiders.

The package root exposes the core request, response, spider, engine, runner, browser-backed client, middleware, logging, and HTML-to-Markdown APIs most applications need. Specialized middleware and pipeline implementations are available from silkworm.middlewares and silkworm.pipelines.

Spiders

class silkworm.Spider[source]

Bases: object

Base class defining a crawl’s initial requests and response callback.

Subclasses usually set name and start_urls, then override parse(). Callbacks are async functions returning None; they report results by awaiting emit() for items and follow() for further requests. Both apply backpressure, so each await completes once the item passed every pipeline or the request entered the queue.

Parameters:
  • name – Per-instance name overriding the class attribute.

  • start_urls – Per-instance starting URLs.

  • allowed_domains – Per-instance domains the crawl may visit.

  • custom_settings – JSON-compatible settings copied for this instance. Keys naming engine options (for example concurrency or max_depth) configure runs started through the runners and CLI.

  • logger – Existing structured logger or context mapping used to create one lazily.

stats_payload

User-defined statistics included in engine summaries.

Example

>>> from silkworm import HTMLResponse, Response, Spider
>>> class TitlesSpider(Spider):
...     name = "titles"
...     start_urls = ("https://example.com",)
...
...     async def parse(self, response: Response) -> None:
...         if isinstance(response, HTMLResponse):
...             title = await response.select_first("title")
...             if title is not None:
...                 await self.emit({"title": title.text})
__init__(*, name=None, start_urls=None, allowed_domains=None, custom_settings=None, logger=None)[source]
Parameters:
  • name (str | None)

  • start_urls (Iterable[str] | None)

  • allowed_domains (Iterable[str] | None)

  • custom_settings (MetaData | None)

  • logger (Logger | dict[str, object] | None)

Return type:

None

allowed_domains: tuple[str, ...] = ()

Domains (and their subdomains) the crawl may visit; requests to other hosts are dropped unless they set dont_filter. Empty allows any host.

property log: Logger

Return the spider logger, creating one bound to its name if needed.

async start_requests()[source]

Schedule one request per start_urls entry.

Override this hook to customize methods, headers, metadata, or callbacks for initial requests; schedule each one with follow().

Return type:

None

async parse(response)[source]

Process a starting response.

Subclasses must implement this coroutine and report results with emit() and follow(). The engine always passes an HTMLResponse to this callback.

Raises:

NotImplementedError – When the base implementation is called.

Parameters:

response (Response)

Return type:

None

async emit(item)[source]

Send a scraped item through the item pipelines.

With the default per-item processing, returns after every pipeline has processed the item. When engine batching is enabled, returns after the item is queued; pending batches drain and surface errors before the callback completes.

Raises:
  • SpiderError – If awaited outside an engine-run callback, or after the callback that owns the current scope finished.

  • TypeError – If item is a Request.

Parameters:

item (JSONLike)

Return type:

None

async follow(target, callback=None, **kwargs)[source]

Schedule a request for crawling.

target is either a ready Request or a URL. Inside a response callback a URL is resolved against the response and inherits its callback, exactly like silkworm.Response.follow(). Requests pass engine deduplication and wait for queue capacity.

Parameters:
  • target (Request | str) – Request to schedule, or an absolute or relative URL.

  • callback (Callback | None) – Response callback for a URL target.

  • **kwargs (object) – Additional Request fields for a URL target.

Raises:
  • SpiderError – If awaited outside an engine-run callback, or after the callback that owns the current scope finished.

  • TypeError – If request fields are combined with a Request target.

Return type:

None

async follow_all(targets, callback=None, **kwargs)[source]

Schedule every non-None target in input order via follow().

Parameters:
  • targets (Iterable[Request | str | None])

  • callback (Callback | None)

  • kwargs (object)

Return type:

None

async open()[source]

Run once after middleware setup and before initial requests enqueue.

Return type:

None

async close()[source]

Run once after pipelines close during normal engine shutdown.

Return type:

None

Requests and Responses

class silkworm.Request[source]

Bases: object

Describe one HTTP request and how its result should be handled.

Parameters:
  • url – Absolute request URL. Relative links can be resolved with silkworm.Response.url_join() or scheduled directly with silkworm.Response.follow().

  • method – HTTP method name.

  • headers – Per-request headers merged over client defaults.

  • params – Query parameters merged with any query already in url.

  • data – Form, byte, or text request body.

  • json – JSON-compatible request body. Do not combine with data.

  • meta – Framework and user metadata propagated with the request.

  • timeout – Per-request timeout in seconds or as a timedelta.

  • callback – Async callable that consumes the response and reports results with silkworm.Spider.emit() and silkworm.Spider.follow(). When omitted on a followed request, the parent request’s callback is reused.

  • errback – Async callable invoked when request processing fails. It may also emit items and follow requests.

  • dont_filter – Bypass engine request deduplication when true.

  • priority – Scheduling priority; larger values are processed first.

Note

Built-in metadata keys include proxy, retry_times, allow_non_html, and redirect_times. Applications may store additional JSON-compatible values alongside them.

replace(**kwargs)[source]

Return a shallow copy with the named dataclass fields replaced.

Unspecified mutable fields are shared with the original request.

Raises:

TypeError – If a keyword is not a request field.

Parameters:

kwargs (object)

Return type:

Self

__init__(url, method='GET', headers=<factory>, params=<factory>, data=None, json=None, meta=<factory>, timeout=None, callback=None, errback=None, dont_filter=False, priority=0)
Parameters:
  • url (str)

  • method (str)

  • headers (Headers)

  • params (QueryParams)

  • data (BodyData)

  • json (JSONValue | None)

  • meta (MetaData)

  • timeout (float | timedelta | None)

  • callback (Callback | None)

  • errback (Errback | None)

  • dont_filter (bool)

  • priority (int)

Return type:

None

class silkworm.Response[source]

Bases: object

Represent a completed HTTP response and its originating request.

Parameters:
  • url – Final URL after redirects.

  • status – Integer HTTP status code.

  • headers – Normalized response headers with lowercase names.

  • body – Raw response payload.

  • request – Request that produced this response.

The decoded text and detected encoding are computed lazily and cached. Call close() when retaining response objects to release their payload memory early.

property text: str

Return the response body decoded with the detected character set.

property encoding: str

Return the normalized encoding used by text.

Detection checks byte-order marks, HTTP headers, HTML/XML declarations, and finally charset-norm before falling back to UTF-8 with replacement characters.

url_join(href)[source]

Resolve href relative to the final response URL.

Parameters:

href (str)

Return type:

str

async follow(href, callback=None, **kwargs)[source]

Schedule a request for a link relative to this response.

Must be awaited inside a spider callback, errback, or start_requests(). The request passes through engine deduplication and waits for queue capacity (backpressure).

Parameters:
  • href (str) – Absolute or relative target URL.

  • callback (Callback | None) – Response callback. The originating request’s callback is inherited when this is omitted.

  • **kwargs (object) – Additional Request fields.

Raises:

SpiderError – If awaited outside an engine-run callback.

Return type:

None

async follow_all(hrefs, callback=None, **kwargs)[source]

Schedule requests for every non-None link in hrefs.

The callback and extra request fields are applied to every request, which are scheduled in input order.

Parameters:
  • hrefs (Iterable[str | None])

  • callback (Callback | None)

  • kwargs (object)

Return type:

None

close()[source]

Release payload references so a retained response does not pin memory.

Closing is idempotent. It clears the body, headers, and cached decoded text; callers should finish reading the response first.

Return type:

None

__init__(url, status, headers, body, request)
Parameters:
Return type:

None

class silkworm.HTMLResponse[source]

Bases: Response

HTTP response with lazy asynchronous HTML selector helpers.

Parameters:

doc_max_size_bytes – Maximum source size accepted by scraper-rs when the document is first parsed.

The parsed document is created once and shared by selector calls. Selector and parsing failures are normalized to SelectorError.

async select(selector)[source]

Return every element matching a CSS selector.

Raises:

SelectorError – If the document cannot be parsed or the selector is invalid.

Parameters:

selector (str)

Return type:

list[AsyncElement]

async select_first(selector)[source]

Return the first CSS match, or None when no element matches.

Raises:

SelectorError – If parsing or selector evaluation fails.

Parameters:

selector (str)

Return type:

AsyncElement | None

async find(selector)[source]

Return the first CSS match as an alias for select_first().

Parameters:

selector (str)

Return type:

AsyncElement | None

async css(selector)[source]

Return all CSS matches as an alias for select().

Parameters:

selector (str)

Return type:

list[AsyncElement]

async css_first(selector)[source]

Return the first CSS match as an alias for select_first().

Parameters:

selector (str)

Return type:

AsyncElement | None

async xpath(xpath)[source]

Return every element matching an XPath expression.

Raises:

SelectorError – If parsing or XPath evaluation fails.

Parameters:

xpath (str)

Return type:

list[AsyncElement]

async xpath_first(xpath)[source]

Return the first XPath match, or None when no element matches.

Raises:

SelectorError – If parsing or XPath evaluation fails.

Parameters:

xpath (str)

Return type:

AsyncElement | None

async prettify()[source]

Return the parsed document formatted as readable HTML.

Raises:

SelectorError – If parsing or serialization fails.

Return type:

str

async to_markdown(*, mode='full', options=None)[source]

Convert the decoded HTML body to Markdown off the event-loop thread.

Parameters:
Returns:

Converted Markdown text.

Return type:

str

async to_markdown_result(*, mode='full', options=None)[source]

Return the structured fast-h2m conversion result.

Parameters:
Return type:

MarkdownResult

close()[source]

Close the parsed document and release response payload references.

Return type:

None

__init__(url, status, headers, body, request, doc_max_size_bytes=5000000)
Parameters:
Return type:

None

Engine

class silkworm.Engine[source]

Bases: object

Coordinate request scheduling, HTTP I/O, callbacks, and item pipelines.

Parameters:
  • spider – Spider instance to execute.

  • concurrency – Maximum simultaneous HTTP requests for the default client.

  • max_pending_requests – Queue capacity used for backpressure. Defaults to ten times the effective HTTP client concurrency. start_requests() waits while the queue is full; callbacks wait only when they are the sole producer, otherwise they enqueue past the bound so workers keep crawling and can never deadlock.

  • emulation – Browser profile used by the default wreq client; pass None to disable impersonation.

  • request_timeout – Default per-request timeout for the default client: 60 seconds (DEFAULT_REQUEST_TIMEOUT) unless given; None disables it. Request.timeout overrides it per request. Ignored when http_client is supplied.

  • html_max_size_bytes – Maximum document size parsed by HTML responses.

  • max_response_size_bytes – Largest response body the default client downloads (None for no limit); larger bodies fail with ResponseTooLargeError.

  • request_middlewares – Request processors applied in list order.

  • response_middlewares – Response processors applied in list order.

  • item_pipelines – Item processors applied in list order, passing each returned value to the next pipeline.

  • item_batch_size – Number of emitted items processed together. The default 1 preserves immediate per-item processing.

  • item_batch_wait – Maximum seconds to wait for a partial batch when item_batch_size is greater than one.

  • log_stats_interval – Seconds between statistics messages, or None to disable periodic summaries.

  • keep_alive – Request connection reuse from the default client when supported by wreq.

  • http_client – Preconfigured client replacing the default client.

  • engine_logger – Event logger customization.

  • dedup_key – Function mapping a request to its deduplication key. Defaults to default_dedup_key().

  • concurrency_per_domain – Maximum simultaneous fetches per host, or None for no per-host limit.

  • max_depth – Drop requests more than this many links away from a start request (start requests have depth 0).

  • max_requests – Stop after sending this many requests.

  • max_items – Stop after this many items passed every pipeline; later items are dropped.

  • max_errors – Stop after this many unrecovered failures.

  • max_duration – Stop after this much wall-clock time.

  • max_error_rate – Fail the crawl when errors / requests_sent exceeds this fraction.

  • min_items – Fail the crawl when fewer items were scraped.

  • max_item_drop_rate – Fail the crawl when the share of items dropped by pipelines exceeds this fraction.

  • job_dir – Directory persisting the seen-set and unfinished requests so an interrupted crawl resumes where it stopped.

  • http_cache – Serve and store responses through this on-disk cache.

  • metrics_port – Serve Prometheus metrics at /metrics on this port while crawling (0 picks a free port).

  • metrics_host – Interface for the metrics server.

Requests with dont_filter bypass deduplication and off-site filtering. Higher request priorities are dequeued before lower ones, while insertion order breaks ties. The stop limits end the crawl gracefully: pending requests are discarded (or kept in job_dir), in-flight requests finish, and run() reports the limit as the close reason. Failure-policy violations make run() raise CrawlFailedError; they are not evaluated for crawls stopped with stop().

__init__(spider, *, concurrency=16, max_pending_requests=None, emulation=DEFAULT_EMULATION, request_timeout=DEFAULT_REQUEST_TIMEOUT, html_max_size_bytes=5_000_000, max_response_size_bytes=DEFAULT_MAX_RESPONSE_SIZE_BYTES, request_middlewares=None, response_middlewares=None, item_pipelines=None, item_batch_size=1, item_batch_wait=0.05, log_stats_interval=None, keep_alive=False, http_client=None, engine_logger=None, dedup_key=None, concurrency_per_domain=None, max_depth=None, max_requests=None, max_items=None, max_errors=None, max_duration=None, max_error_rate=None, min_items=None, max_item_drop_rate=None, job_dir=None, http_cache=None, metrics_port=None, metrics_host='127.0.0.1')[source]
Parameters:
  • spider (Spider)

  • concurrency (int)

  • max_pending_requests (int | None)

  • emulation (Emulation | Profile | None)

  • request_timeout (float | timedelta | None)

  • html_max_size_bytes (int)

  • max_response_size_bytes (int | None)

  • request_middlewares (Iterable[RequestMiddleware] | None)

  • response_middlewares (Iterable[ResponseMiddleware] | None)

  • item_pipelines (Iterable[ItemPipeline] | None)

  • item_batch_size (int)

  • item_batch_wait (float)

  • log_stats_interval (float | None)

  • keep_alive (bool)

  • http_client (FetchClient | None)

  • engine_logger (EngineLogger | None)

  • dedup_key (DedupKey | None)

  • concurrency_per_domain (int | None)

  • max_depth (int | None)

  • max_requests (int | None)

  • max_items (int | None)

  • max_errors (int | None)

  • max_duration (float | timedelta | None)

  • max_error_rate (float | None)

  • min_items (int | None)

  • max_item_drop_rate (float | None)

  • job_dir (str | os.PathLike[str] | None)

  • http_cache (HttpCache | None)

  • metrics_port (int | None)

  • metrics_host (str)

Return type:

None

property close_reason: str | None

Return why the crawl is stopping, or None while it runs normally.

property in_flight: int

Return the number of requests currently being processed.

stop(reason='shutdown')[source]

Stop the crawl gracefully.

New requests are no longer scheduled and queued ones are discarded (they stay saved when a job directory is configured, so the crawl can resume). Requests already being processed finish, pipelines close normally, and run() returns with reason as the close reason. Calling it again has no effect.

Parameters:

reason (str)

Return type:

None

async open_spider()[source]

Open middleware, spider, and pipelines, then enqueue initial requests.

Middleware opens before the spider; pipelines open afterward in their configured order. When resuming a job, saved requests are queued before start_requests() runs (already-seen start requests are skipped).

Return type:

None

async close_spider()[source]

Close pipelines, the spider, and middleware lifecycle hooks.

Components close in reverse startup order. Middleware instances close once even when registered for both request and response work.

Return type:

None

metrics_text()[source]

Return current statistics in the Prometheus text exposition format.

Return type:

str

async run()[source]

Run the crawl until the queue drains or a stop condition, then clean up.

Worker tasks, periodic statistics, and lifecycle hooks are managed as a task group. The HTTP client, spider components, job state, and metrics server are always closed, and a final statistics record is emitted.

Returns:

The crawl’s CrawlResult.

Raises:
  • CrawlFailedError – If the crawl violated its failure policy (max_error_rate, min_items, max_item_drop_rate).

  • Exception – Errors from lifecycle hooks, pipelines’ open/close, or cleanup propagate after structured error logging.

Return type:

CrawlResult

class silkworm.EngineOptions[source]

Bases: TypedDict

Keyword options for Engine, also accepted by every runner.

run_spider(MySpider, concurrency=32, request_timeout=10) forwards these to Engine; omitted keys use the Engine defaults. Supply http_client to inject a compatible client; its concurrency then controls worker count and default queue capacity.

class silkworm.EngineLogger[source]

Bases: object

Customizable engine event logger.

Subclass this when you need to redact or reshape selected engine log events. Set an event level to None to suppress that event.

fetching_request(logger, request, spider)[source]

Log that request is about to be sent.

Parameters:
Return type:

None

fetched_response(logger, request, response, spider)[source]

Log a completed response with status and optional request URL.

Parameters:
Return type:

None

retrying_request(logger, request, spider, *, source)[source]

Log that a middleware or response requested another attempt.

Parameters:
Return type:

None

running_item_pipeline(logger, pipeline, spider)[source]

Log item dispatch using the pipeline’s effective log level.

Parameters:
Return type:

None

__init__(fetched_response_level='INFO', fetching_request_level='DEBUG', item_pipeline_level='DEBUG', retry_request_level='DEBUG', include_request_url=True)
Parameters:
Return type:

None

silkworm.default_dedup_key(req)[source]

Return the engine’s default deduplication key for req.

This is request_fingerprint(): the HTTP method, the canonical URL (normalized case, default port, sorted query, no fragment) with params merged in, and the request body. Headers and metadata do not affect it.

Parameters:

req (Request)

Return type:

str

silkworm.request_fingerprint(request)[source]

Return a stable hex digest identifying what request fetches.

The fingerprint covers the HTTP method, the canonical URL with params merged in, and the request body (form data or JSON). Headers, metadata, callbacks, and URL fragments are ignored, so two requests with the same fingerprint are treated as duplicates by the engine’s default deduplicator and share HTTP cache entries.

Parameters:

request (Request)

Return type:

str

silkworm.canonicalize_url(url, *, keep_fragments=False)[source]

Return a normalized form of url for deduplication and caching.

The scheme and host are lowercased, default ports and (unless keep_fragments) fragments are removed, percent-encoding is normalized, an empty path becomes /, and query parameters are sorted while keeping blank values and repeated keys. Two URLs that address the same resource in these respects map to the same string.

Example

>>> canonicalize_url("HTTP://Example.com:80/a?b=2&a=1#top")
'http://example.com/a?a=1&b=2'
Parameters:
Return type:

str

type silkworm.DedupKey = Callable[[Request], str]

Runners

async silkworm.crawl(spider, *, handle_signals=False, **options)[source]

Run spider to completion on the current event loop.

Options are resolved through silkworm.settings.resolve_options(), so SILKWORM_* environment variables and the spider’s custom_settings apply beneath the options passed here.

Parameters:
  • spider (Spider | type[Spider]) – Spider instance, or a no-argument spider class.

  • handle_signals (bool) – Stop gracefully on the first SIGINT/SIGTERM and cancel on the second. Off by default because the calling application owns the event loop and may handle signals itself.

  • **options (Unpack[EngineOptions]) – EngineOptions forwarded to the engine.

Returns:

The crawl’s CrawlResult.

Raises:
Return type:

CrawlResult

Use this coroutine when the application already owns an event loop; use a run_spider* function from synchronous code.

silkworm.run_spider(spider, *, loop_factory=None, handle_signals=True, **options)[source]

Run spider with asyncio, blocking until the crawl finishes.

The first SIGINT (Ctrl+C) or SIGTERM stops the crawl gracefully: pending requests are discarded (or saved with job_dir), in-flight requests finish, and pipelines close normally. A second signal cancels immediately.

Parameters:
  • spider (Spider | type[Spider]) – Spider instance, or a spider class to instantiate without arguments.

  • loop_factory (LoopFactory | None) – Optional event loop factory, e.g. from uvloop.

  • handle_signals (bool) – Install the graceful shutdown handlers described above.

  • **options (Unpack[EngineOptions]) – Engine options; see EngineOptions.

Returns:

The crawl’s CrawlResult.

Raises:

CrawlFailedError – If the crawl violated its failure policy.

Return type:

CrawlResult

silkworm.run_spider_rsloop(spider, **options)[source]

Run spider on an rsloop event loop (pip install silkworm-rs[rsloop]).

Raises:

ImportError – If rsloop is not installed.

Parameters:
Return type:

CrawlResult

silkworm.run_spider_uvloop(spider, **options)[source]

Run spider on a uvloop event loop (pip install silkworm-rs[uvloop]).

Raises:

ImportError – If uvloop is not installed.

Parameters:
Return type:

CrawlResult

silkworm.run_spider_winloop(spider, **options)[source]

Run spider on a winloop event loop, optimized for Windows (pip install silkworm-rs[winloop]).

Raises:

ImportError – If winloop is not installed.

Parameters:
Return type:

CrawlResult

silkworm.run_spider_trio(spider, **options)[source]

Run spider with trio as the async backend (pip install silkworm-rs[trio]).

The engine uses asyncio primitives, so it runs inside trio via trio-asyncio. This runner is currently available on Python 3.13 only because trio-asyncio 0.16 is incompatible with Python 3.14 and newer.

Raises:

ImportError – If trio or trio-asyncio is not installed.

Parameters:
Return type:

CrawlResult

Crawl Results

class silkworm.CrawlResult[source]

Bases: object

Outcome of a finished crawl, returned by Engine.run() and runners.

spider

Spider name.

Type:

str

close_reason

Why the crawl ended: "finished" when the queue drained, "shutdown" after a stop request or signal, a limit name such as "max_items", or the reason passed to CloseSpider.

Type:

str

elapsed_seconds

Wall-clock crawl duration.

Type:

float

stats

Final counters (see BASE_COUNTERS).

Type:

Mapping[str, int]

labeled_stats

Per-label breakdowns (see BASE_LABELED_COUNTERS).

Type:

Mapping[str, Mapping[str, int]]

custom_stats

A copy of the spider’s stats_payload.

Type:

Mapping[str, JSONValue]

failures

Failure-policy violations; empty when the crawl succeeded.

Type:

tuple[str, …]

property ok: bool

Return whether the crawl met its failure policy.

property requests_sent: int

Return the number of requests handed to the HTTP client.

property responses_received: int

Return the number of responses received.

property items_scraped: int

Return the number of items that passed every pipeline.

property items_dropped: int

Return the number of items discarded by pipelines or limits.

property errors: int

Return the number of unrecovered request or callback failures.

property error_rate: float

Return errors / requests_sent (0.0 when nothing was sent).

property item_drop_rate: float

Return items_dropped / (items_scraped + items_dropped).

__init__(spider, close_reason, elapsed_seconds, stats, labeled_stats, custom_stats=<factory>, failures=())
Parameters:
  • spider (str)

  • close_reason (str)

  • elapsed_seconds (float)

  • stats (Mapping[str, int])

  • labeled_stats (Mapping[str, Mapping[str, int]])

  • custom_stats (Mapping[str, JSONValue])

  • failures (tuple[str, ...])

Return type:

None

Convenience Helpers

async silkworm.fetch_html(url, *, emulation=DEFAULT_EMULATION, timeout=DEFAULT_REQUEST_TIMEOUT)[source]

Fetch and asynchronously parse one HTML document with wreq.

Parameters:
  • url (str) – Absolute URL to fetch.

  • emulation (Emulation | Profile | None) – Browser profile to impersonate, or None to disable it.

  • timeout (float | timedelta | None) – Request timeout in seconds or as a timedelta; 60 seconds (DEFAULT_REQUEST_TIMEOUT) unless given, and None disables it.

Returns:

A (text, AsyncDocument) tuple with awaitable selector helpers.

Return type:

tuple[str, AsyncDocument]

Note

This convenience API does not apply spider middleware, retries, deduplication, or pipelines.

async silkworm.fetch_html_cdp(url, *, ws_endpoint='ws://127.0.0.1:9222', timeout=None)[source]

Fetch HTML from a URL using CDP (Chrome DevTools Protocol).

This function connects to a CDP-compatible browser (like Lightpanda, Chrome, or Chromium) and fetches the rendered HTML after JavaScript execution.

Parameters:
  • url (str) – The URL to fetch

  • ws_endpoint (str) – WebSocket endpoint for CDP connection (default: ws://127.0.0.1:9222)

  • timeout (float | None) – Optional timeout in seconds

Returns:

A tuple of (text, AsyncDocument) with awaitable selector helpers.

Raises:
Return type:

tuple[str, AsyncDocument]

Example

>>> import asyncio
>>> from silkworm import fetch_html_cdp
>>>
>>> async def main():
...     text, doc = await fetch_html_cdp("https://example.com")
...     title = await doc.select_first("title")
...     print(title.text if title else "No title")
>>>
>>> asyncio.run(main())
async silkworm.fetch_html_servo(url, *, timeout=None, settle_ms=0, user_agent=None, javascript=None, allow_private_addresses=False)[source]

Fetch and parse rendered HTML with servofetch and Servo.

Parameters:
  • url (str) – Absolute URL to render.

  • timeout (float | timedelta | None) – Render timeout in seconds or as a timedelta.

  • settle_ms (int) – Delay after loading before capturing the document.

  • user_agent (str | None) – Optional browser user agent override.

  • javascript (str | None) – Optional JavaScript evaluated by the rendered-page client.

  • allow_private_addresses (bool) – Permit navigation to private network addresses.

Returns:

A (text, AsyncDocument) tuple with awaitable selector helpers.

Raises:
  • ImportError – If a compatible servofetch build is unavailable.

  • HttpError – If rendering fails.

Return type:

tuple[str, AsyncDocument]

Client Adapters

class silkworm.http.HttpClient[source]

Bases: object

Send Silkworm requests through a browser-impersonating wreq client.

Parameters:
  • concurrency – Maximum requests in flight.

  • emulation – Browser profile to impersonate, or None to disable it.

  • default_headers – Headers merged below per-request headers.

  • timeout – Default request timeout in seconds or as a timedelta (DEFAULT_REQUEST_TIMEOUT, 60 seconds, unless given); None disables it. Request.timeout overrides it per request. The budget covers sending the request and downloading the whole body, restarts for each redirect hop, and excludes time spent waiting for a concurrency slot.

  • html_max_size_bytes – Maximum document size parsed by HTML selectors.

  • follow_redirects – Follow redirect responses internally.

  • max_redirects – Maximum redirect hops.

  • keep_alive – Request connection reuse when the installed wreq version supports it.

  • max_response_size_bytes – Largest body downloaded, or None for no limit. A larger Content-Length fails before the body is read, and streamed bodies stop as soon as they exceed the limit. Override it per request with request.meta["max_response_size"].

  • **client_kwargs – Additional options forwarded to wreq.Client.

Requests are converted to HTMLResponse when headers or a small body sniff indicate HTML; all other payloads become Response.

__init__(*, concurrency=16, emulation=DEFAULT_EMULATION, default_headers=None, timeout=DEFAULT_REQUEST_TIMEOUT, html_max_size_bytes=5_000_000, follow_redirects=True, max_redirects=10, keep_alive=False, max_response_size_bytes=DEFAULT_MAX_RESPONSE_SIZE_BYTES, **client_kwargs)[source]
Parameters:
  • concurrency (int)

  • emulation (Emulation | Profile | None)

  • default_headers (Headers | None)

  • timeout (float | timedelta | None)

  • html_max_size_bytes (int)

  • follow_redirects (bool)

  • max_redirects (int)

  • keep_alive (bool)

  • max_response_size_bytes (int | None)

  • client_kwargs (object)

Return type:

None

property concurrency: int

Return the maximum number of requests allowed in flight.

property html_max_size_bytes: int

Return the HTML document parsing limit in bytes.

property max_response_size_bytes: int | None

Return the default response body size limit in bytes.

async fetch(req)[source]

Send one request, follow redirects, and return a normalized response.

The per-request timeout and proxy metadata override client defaults. Synthetic response metadata is honored for middleware integrations.

Raises:
Parameters:

req (Request)

Return type:

Response

async close()[source]

Close the underlying transport.

Return type:

None

class silkworm.http.FetchClient[source]

Bases: Protocol

Interface the engine needs from an HTTP client.

HttpClient, CDPClient, ServoFetchClient, OnionLinkClient, and CachingHttpClient implement it; pass any conforming object as http_client.

property concurrency: int

Return the maximum number of requests in flight.

property html_max_size_bytes: int

Return the HTML document parsing limit in bytes.

async fetch(req)[source]

Send req and return its response.

Parameters:

req (Request)

Return type:

Response

async close()[source]

Release transport resources.

Return type:

None

__init__(*args, **kwargs)
class silkworm.HttpCache[source]

Bases: object

Configuration and storage for cached HTTP responses.

Parameters:
  • directory – Cache directory; created when missing.

  • expiration – Maximum entry age (seconds or timedelta); None keeps entries forever.

  • ignore_statuses – Response statuses that are never stored (e.g. 5xx).

  • methods – HTTP methods eligible for caching.

Set request.meta["dont_cache"] = True to bypass the cache for one request.

__init__(directory, *, expiration=None, ignore_statuses=(500, 502, 503, 504, 522, 524, 408, 429), methods=('GET', 'HEAD'))[source]
Parameters:
Return type:

None

wrap(client)[source]

Return a client that serves client’s responses from this cache.

Parameters:

client (FetchClient)

Return type:

CachingHttpClient

is_cacheable(request)[source]

Return whether request may be read from or written to the cache.

Parameters:

request (Request)

Return type:

bool

load(request)[source]

Return (metadata, body) for a fresh entry, else None.

Parameters:

request (Request)

Return type:

tuple[dict[str, object], bytes] | None

store(request, response)[source]

Write response for request; return whether it was stored.

Parameters:
Return type:

bool

silkworm.http.DEFAULT_REQUEST_TIMEOUT = 60.0

Convert a string or number to a floating-point number, if possible.

Default per-request timeout in seconds (60.0) for HttpClient and the engine’s request_timeout.

silkworm.http.MOCK_RESPONSE_META_KEY

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.

Request meta key holding a {"status", "headers", "body", "url"} mapping that HttpClient returns instead of making a network call.

class silkworm.CDPClient[source]

Bases: object

Fetch rendered pages through a CDP-compatible browser.

Parameters:
  • ws_endpoint – Browser WebSocket endpoint. A bare ws://host:port (or http://host:port) also works with Chrome/Chromium, which only accept their per-session /devtools/browser/<id> URL: when the direct connection fails, the URL advertised at /json/version is used, keeping the host and port given here.

  • concurrency – Maximum simultaneous page fetches.

  • timeout – Default command and navigation timeout in seconds.

  • html_max_size_bytes – Maximum rendered document size accepted by the HTML parser and WebSocket transport.

Raises:

ImportError – If the cdp extra is not installed.

Example

>>> client = CDPClient(ws_endpoint="ws://127.0.0.1:9222")
>>> await client.connect()
>>> try:
...     response = await client.fetch(request)
... finally:
...     await client.close()
__init__(*, ws_endpoint='ws://127.0.0.1:9222', concurrency=16, timeout=None, html_max_size_bytes=5_000_000)[source]
Parameters:
  • ws_endpoint (str)

  • concurrency (int)

  • timeout (float | None)

  • html_max_size_bytes (int)

Return type:

None

property concurrency: int

Return the maximum number of simultaneous page fetches.

property html_max_size_bytes: int

Return the rendered HTML size limit in bytes.

async connect()[source]

Connect to the browser and create an isolated page target.

Calling this method more than once is harmless.

Raises:

HttpError – If connection or target initialization fails.

Return type:

None

async fetch(req)[source]

Navigate to a URL and return its rendered HTML.

The response status is reported as 200 because CDP does not reliably expose the navigation status. The final document URL is detected after redirects when supported by the browser.

Raises:

HttpError – If the client is disconnected, navigation times out, or rendered HTML cannot be retrieved.

Parameters:

req (Request)

Return type:

Response

async close()[source]

Cancel background work and close the page target and WebSocket.

Return type:

None

class silkworm.ServoFetchClient[source]

Bases: object

Render pages with servofetch for use as an engine HTTP client.

Parameters:
  • concurrency – Maximum simultaneous renders.

  • timeout – Default render timeout.

  • settle_ms – Default delay after page load before capture.

  • user_agent – Default browser user agent.

  • allow_private_addresses – Permit navigation to private network addresses.

  • html_max_size_bytes – Maximum rendered document size parsed by selectors.

  • onion_bootstrap – Optional Tor bootstrap endpoint forwarded to Servo.

  • onion_consensus_file – Optional cached Tor consensus file.

  • onion_verbose – Enable verbose Tor integration output.

  • onion_response_limit – Maximum Tor response size in bytes.

Request metadata can override JavaScript, settle delay, user agent, and screenshot behavior through the exported SERVO_*_META_KEY constants.

Raises:

ImportError – If a compatible servofetch build is unavailable.

__init__(*, concurrency=16, timeout=None, settle_ms=0, user_agent=None, allow_private_addresses=False, html_max_size_bytes=5_000_000, onion_bootstrap=None, onion_consensus_file=None, onion_verbose=False, onion_response_limit=4 * 1024 * 1024)[source]
Parameters:
  • concurrency (int)

  • timeout (float | timedelta | None)

  • settle_ms (int)

  • user_agent (str | None)

  • allow_private_addresses (bool)

  • html_max_size_bytes (int)

  • onion_bootstrap (str | None)

  • onion_consensus_file (str | None)

  • onion_verbose (bool)

  • onion_response_limit (int)

Return type:

None

property concurrency: int

Return the maximum number of simultaneous renders.

property html_max_size_bytes: int

Return the rendered HTML parsing limit in bytes.

async fetch(req)[source]

Render req.url and return its HTML response.

Per-request timeout and supported Servo metadata override client defaults. Screenshot requests still return the page HTML and expose screenshot metadata through synthetic response headers.

Raises:
  • TypeError – If a recognized metadata value has the wrong type.

  • HttpError – If rendering fails or no HTML is returned.

Parameters:

req (Request)

Return type:

HTMLResponse

async close()[source]

Close the underlying Servo browser using its available close hook.

Return type:

None

silkworm.servo.SERVO_JAVASCRIPT_META_KEY = 'servo_javascript'

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.

silkworm.servo.SERVO_SETTLE_MS_META_KEY = 'servo_settle_ms'

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.

silkworm.servo.SERVO_USER_AGENT_META_KEY = 'servo_user_agent'

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.

silkworm.servo.SERVO_SCREENSHOT_META_KEY = 'servo_screenshot'

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.

silkworm.servo.SERVO_FULL_PAGE_META_KEY = 'servo_full_page'

str(object=’’) -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.__str__() (if defined) or repr(object). encoding defaults to ‘utf-8’. errors defaults to ‘strict’.

class silkworm.OnionLinkClient[source]

Bases: HttpClient

Fetch Tor v3 onion services through onionlink.

Parameters:
  • concurrency – Maximum simultaneous requests.

  • default_headers – Headers merged into every request.

  • timeout – Default request timeout.

  • html_max_size_bytes – Maximum HTML size accepted by selectors.

  • follow_redirects – Follow HTTP redirect responses when true.

  • max_redirects – Maximum redirects before raising an error.

  • bootstrap – Tor directory authority bootstrap endpoint.

  • consensus_file – Optional consensus cache path.

  • verbose – Enable verbose onionlink output.

  • response_limit – Default maximum response body size in bytes. Override it per request with onionlink_response_limit metadata.

Raises:
  • ImportError – If a compatible onionlink extra is unavailable.

  • ValueError – If max_redirects is negative.

__init__(*, concurrency=16, default_headers=None, timeout=None, html_max_size_bytes=5_000_000, follow_redirects=True, max_redirects=10, bootstrap='128.31.0.39:9131', consensus_file='', verbose=False, response_limit=4 * 1024 * 1024)[source]
Parameters:
  • concurrency (int)

  • default_headers (Headers | None)

  • timeout (float | timedelta | None)

  • html_max_size_bytes (int)

  • follow_redirects (bool)

  • max_redirects (int)

  • bootstrap (str)

  • consensus_file (str)

  • verbose (bool)

  • response_limit (int)

Return type:

None

async fetch(req)[source]

Fetch one HTTP(S) .onion URL.

Query parameters, redirects, timeouts, bodies, and response type detection follow the standard client contract.

Raises:
  • TypeError – If the response-limit metadata is not an integer.

  • HttpError – If the URL is not an onion service or fetching fails.

Parameters:

req (Request)

Return type:

Response

async close()[source]

Complete client cleanup; onionlink sessions need no close operation.

Return type:

None