Core Concepts¶
This section covers Silkworm’s Spider/Request/Response model, callback semantics, and how data flows through the engine.
Spider¶
Spider is the base class you subclass for each crawl. See src/silkworm/spiders.py.
Key attributes and hooks:
name: Spider identifier used in logs and stats.start_urls: Seed URLs forstart_requests().allowed_domains: Domains (and subdomains) the crawl may visit; other hosts are dropped. See Staying on-site.custom_settings: Per-spider settings storage (copied on init).start_requests(): Coroutine that schedules initial requests withawait self.follow(...).parse(response): Main callback (auto-wrapped toHTMLResponse).await emit(item): Send a scraped item through the item pipelines.await follow(request_or_url, callback=None, **kwargs)/await follow_all(...): Schedule requests for crawling.open()/close(): Lifecycle hooks called by the engine.
from silkworm import Response, Spider
class MySpider(Spider):
name = "my_spider"
start_urls = ("https://example.com",)
async def parse(self, response: Response) -> None:
await self.emit({"url": response.url, "status": response.status})
Request¶
Request is a slotted dataclass used to describe HTTP work. See src/silkworm/request.py.
Important fields:
url,method,headers,params,data,jsontimeout: Per-request timeout (seconds ortimedelta).callback: Async callback to run with the response.errback: Async callback to run when the request fails and no middleware retries it.meta: Free-form dict for middlewares and custom logic.dont_filter: Bypass request deduplication.priority: Higher values are dequeued first; requests with the same priority keep FIFO order.
Request.replace(**kwargs) is the safest way to create updated requests, while
headers and meta are mutable dicts that middlewares may update in-place.
from silkworm import Request
request = Request(
url="https://example.com/search",
method="GET",
params={"q": "silkworm"},
headers={"accept": "text/html"},
timeout=5,
)
Request Error Handling¶
Use Request.errback for per-request recovery from fetch, middleware, or callback
exceptions that were not handled by exception middlewares. The errback receives the
failed Request and the raised exception. Like a normal callback it is an async
function returning None that may await self.emit(...) items and
await self.follow(...) follow-up requests.
from silkworm import Request
async def start_requests(self) -> None:
await self.follow(
"https://example.com/maybe-down",
callback=self.parse,
errback=self.handle_error,
)
async def handle_error(self, request: Request, exception: Exception) -> None:
await self.emit(
{
"url": request.url,
"error_type": exception.__class__.__name__,
}
)
Built-in meta Keys¶
These are used by built-in components (you can add your own as well):
proxy: Used byProxyMiddlewareand HttpClient for proxy routing.retry_times: Used byRetryMiddlewareto track attempts.allow_non_html: Used bySkipNonHTMLMiddlewareto bypass filtering.cookiejar: Used byCookiesMiddlewareto isolate named cookie sessions.cookies: Used byCookiesMiddlewareto add cookies for one request.dont_merge_cookies: Used byCookiesMiddlewareto bypass cookie storage and header merging for one exchange.redirect_times: Set by HttpClient when following redirects.cloudflare_crawl: Used byCloudflareCrawlMiddleware; set toTrueor a dict of crawl options.servo_javascript,servo_settle_ms,servo_user_agent,servo_screenshot,servo_full_page: Used byServoFetchClient.onionlink_response_limit: Used byOnionLinkClientto override the per-response byte cap.
Response and HTMLResponse¶
Response contains the response payload; HTMLResponse adds selector helpers. See src/silkworm/response.py.
Core APIs:
text: Decoded body text with charset detection.encoding: Detected or default encoding.url_join(href): Resolve a relative URL against the response URL.await follow(href, callback=None, **kwargs): URL join + callback reuse, then schedule the request.await follow_all(hrefs, callback=None, **kwargs): Schedule every non-Nonelink in order.close(): Release payload references to save memory.await to_markdown(mode="full" | "minimal" | "mdream", options=None): Convert anHTMLResponseto Markdown viafast-h2m(runs off the event-loop thread).await to_markdown_result(...): Returnfast-h2m’s structured conversion result.
from silkworm import HTMLResponse, Response
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
title = await response.select_first("title")
if title:
await self.emit({"title": title.text.strip()})
Selector helpers on HTMLResponse (async):
select(selector)select_first(selector)css(selector)css_first(selector)xpath(xpath)xpath_first(xpath)find(selector): Alias forselect_first.prettify(): Return the parsed document formatted as readable HTML.
Elements returned from these helpers also expose async selectors, so nested lookups should be awaited:
for card in await response.select(".card"):
title = await card.select_first("h2")
The selector engine uses scraper-rs and respects doc_max_size_bytes (see HttpClient). Errors are raised as SelectorError in src/silkworm/exceptions.py.
HTML to Markdown¶
Markdown conversion is available on HTMLResponse (above) and as standalone helpers from silkworm. See src/silkworm/markdown.py.
html_to_markdown(html, mode="full", options=None): Convert a complete document to a Markdown string.convert_html_to_markdown(html, mode="full", options=None): Returnfast-h2m’s structured result (MarkdownResult, a dict).MarkdownStream(mode="minimal", options=None): Incremental converter; callprocess_chunk(html)for each fragment andfinish()at the end. Each call returns the Markdown available so far.stream_html_to_markdown(chunks, ...)/stream_html_to_markdown_async(chunks, ...): Convert an iterable / async iterable of HTML chunks into one string.
MarkdownMode is "full" (rich converter, the default for whole documents), "minimal" (lean Fast DOM path, the default for streaming), or "mdream" (mdream-backed lean path). options (MarkdownOptions) is a mapping of extra fast-h2m options that override the mode defaults. Conversion failures raise MarkdownConversionError; an unknown mode raises ValueError.
from silkworm import MarkdownStream, html_to_markdown
markdown = html_to_markdown("<h1>Title</h1><p>Hello</p>")
stream = MarkdownStream()
parts = [stream.process_chunk(chunk) for chunk in ("<h1>Ti", "tle</h1><p>Hi</p>")]
parts.append(stream.finish())
markdown = "".join(parts)
Reporting Results: emit and follow¶
Callbacks are async functions that return None. Instead of yielding or
returning results, they push them to the engine while they run. See
src/silkworm/spiders.py.
Call |
Effect |
|---|---|
|
Runs |
|
Schedules a ready |
|
Builds and schedules a |
|
Calls |
|
Schedules links relative to that response. |
With the default item_batch_size=1, each await emit finishes after the item
has passed every pipeline. When batching is enabled, it finishes after bounded
enqueueing; the callback scope drains its pending batches before returning, so
pipeline errors still fail the callback. follow
waits while the queue is full if it is the only callback waiting; otherwise it
enqueues past the bound so workers keep crawling and cannot deadlock (see
Queue Capacity).
Both modes apply backpressure. Per-item pipeline errors surface at the emit
call site; batch failures surface while the callback scope drains.
from silkworm import Request, Response, Spider
class MixedSpider(Spider):
name = "mixed"
start_urls = ("https://example.com",)
async def parse(self, response: Response) -> None:
await self.emit({"url": response.url})
await self.follow("/page2") # relative to response.url, reuses parse
await self.follow(
Request(url="https://example.com/api", callback=self.parse_api)
)
async def parse_api(self, response: Response) -> None:
await self.emit({"ok": response.status == 200})
Concurrency inside a callback¶
emit and follow are bound to the running callback through a context
variable, so tasks spawned inside a callback can use them too. Wait for those
tasks before the callback returns; asyncio.TaskGroup does this for you:
import asyncio
async def parse(self, response: HTMLResponse) -> None:
async with asyncio.TaskGroup() as tg:
for card in await response.select(".card"):
tg.create_task(self.parse_card(card))
async def parse_card(self, card) -> None:
title = await card.select_first("h2")
if title is not None:
await self.emit({"title": title.text})
Rules the engine enforces¶
emit/followraiseSpiderErrorwhen awaited outsidestart_requests(), a request callback, or an errback, and when a task calls them after its callback has returned. Calls already in progress when the callback returns are awaited before the request is marked done.A callback that is an async generator (uses
yield), is notasync, or returns a value other thanNonefails with aSpiderErrorexplaining the fix.emitrejectsRequestobjects (usefollow), andfollowaccepts request fields only together with a URL target.
Migrating from 0.10 (yield-based callbacks)¶
0.10 |
0.11 |
|---|---|
|
|
|
|
|
|
|
|
|
one |
|
|
|
|
Note: The engine auto-wraps only the spider’s
parsecallback toHTMLResponse. Other callbacks receive theResponseproduced by the HTTP client, which may already be anHTMLResponsefor HTML content.
Errors¶
All framework exceptions derive from SilkwormError and are importable from silkworm. See src/silkworm/exceptions.py.
Exception |
Raised when |
|---|---|
|
A fetch fails: redirect loops, a closed client, or CDP/Servo/OnionLink failures. Subclasses: |
|
A spider callback (or |
|
CSS/XPath selector evaluation or HTML parsing fails. |
|
HTML-to-Markdown conversion fails. |
|
Declarative item extraction fails; see Declarative Extraction. |
|
The crawl violated its failure policy; |
|
A bulk destination reported one or more rejected items; count details are available on the exception. |
Three exceptions are control-flow signals rather than errors: IgnoreRequest
(drop a request from a middleware), DropItem (discard an item from a pipeline),
and CloseSpider (stop the crawl gracefully). See
Control-flow exceptions.
Recover from failed requests with Request.errback or an exception middleware (see Middlewares).
Deduplication¶
The engine keeps a set of seen request keys. The default key is
request_fingerprint(request): the HTTP method, the canonical URL (lowercase
scheme and host, no default port or fragment, sorted query parameters) with
Request.params merged in, and the request body. Headers and meta are ignored.
See Deduplication for details, and pass
dedup_key to Engine, crawl, or run_spider for a custom rule.
await self.follow(same_url, dont_filter=True)
Or customize the key globally for a run, for example to treat URLs as equal regardless of their query string:
from silkworm import Request, canonicalize_url, run_spider
def dedup_without_query(req: Request) -> str:
return canonicalize_url(req.url).split("?", 1)[0]
run_spider(MySpider, dedup_key=dedup_without_query)
Data Types¶
Public type aliases and protocols live in silkworm.types. Import them from there to annotate your own code, for example:
from silkworm.types import Callback, JSONValue, Logger, MetaData
It covers:
JSON item shapes:
JSONScalar,JSONValue, and the read-onlyJSONLikeaccepted bySpider.emit.Request data:
Headers,QueryParams,QueryValue,MetaData,BodyData.Callbacks:
CallbackandErrback(async callables returningNone).Engine and runners:
EngineOptions,DedupKey,LoopFactory,CrawlResult, and theFetchClientprotocol for custom HTTP clients.Logging: the
Loggerprotocol andLogLevel.Middleware and pipeline protocols:
RequestMiddleware,ResponseMiddleware,ExceptionMiddleware,ItemPipeline,BatchItemPipeline, plusItemCallback(forCallbackPipeline),ItemSchema/ItemValidator/ModelSchema(forValidationPipeline), andZenohKeyResolver(forZenohPipeline).
See the silkworm.types reference for each definition.