Runners¶
Runners are convenience helpers that build an Engine and start the crawl. See src/silkworm/runner.py.
Spiders and Engine Options¶
Every runner takes a spider and the same keyword engine options; the runners differ only in the event loop they use.
Spider: pass a spider class to use its no-argument constructor, or an instance when the spider takes arguments. Constructor arguments then go to the spider directly, so type checkers verify them:
run_spider(QuotesSpider, concurrency=16) run_spider(SitemapSpider(sitemap_url=url, max_pages=5), concurrency=16)
Engine options are the keyword parameters of
Engine, described by theEngineOptionstyped dict (from silkworm import EngineOptions), for exampleconcurrency,request_timeout,request_middlewares,item_pipelines,max_items, orjob_dir(see Engine and HTTP Client for the full list). Omitted options fall back toSILKWORM_<OPTION>environment variables, then the spider’scustom_settings, then theEnginedefaults (see Settings). To share options between runs, build them once:from silkworm import EngineOptions, run_spider options: EngineOptions = {"concurrency": 32, "request_timeout": 10, "keep_alive": True} run_spider(QuotesSpider, **options)
emulation selects the browser profile wreq impersonates (Emulation.Firefox139 by default); pass emulation=None to disable it.
Upgrading: runners no longer forward unknown keyword arguments to the spider constructor. Replace
run_spider(MySpider, pages=3, concurrency=8)withrun_spider(MySpider(pages=3), concurrency=8).
Results and Failures¶
Every runner returns the crawl’s CrawlResult (close reason, counters, and
failure-policy violations) and raises CrawlFailedError when the crawl violates
its failure policy:
result = run_spider(MySpider, max_error_rate=0.1, min_items=1)
print(result.close_reason, result.items_scraped, result.errors)
See Production Crawling.
Async Entry Point: crawl¶
crawl is an async helper that runs the spider on the current event loop and
returns its CrawlResult.
from silkworm import crawl
result = await crawl(MySpider, concurrency=16, request_timeout=10)
crawl leaves signal handling to the application that owns the event loop; pass
handle_signals=True to stop gracefully on SIGINT/SIGTERM as the sync runners do.
Sync Entry Point: run_spider¶
run_spider wraps crawl with asyncio.run. Pass loop_factory= (a LoopFactory, i.e. a zero-argument callable returning an event loop) to run on a custom asyncio event loop.
import asyncio
from silkworm import run_spider
run_spider(MySpider, concurrency=16, request_timeout=10)
run_spider(MySpider, loop_factory=asyncio.new_event_loop)
The sync runners stop gracefully on the first SIGINT (Ctrl+C) or SIGTERM,
finishing in-flight requests and closing pipelines; a second signal cancels
immediately. Pass handle_signals=False to run_spider to keep Python’s
default handling. See Graceful shutdown.
rsloop¶
run_spider_rsloop installs rsloop and then runs the spider.
from silkworm import run_spider_rsloop
run_spider_rsloop(MySpider, concurrency=32)
Requires:
pip install silkworm-rs[rsloop]
uvloop (Unix)¶
run_spider_uvloop installs uvloop and then runs the spider.
from silkworm import run_spider_uvloop
run_spider_uvloop(MySpider, concurrency=32)
Requires:
pip install silkworm-rs[uvloop]
winloop (Windows)¶
run_spider_winloop installs winloop on Windows.
from silkworm import run_spider_winloop
run_spider_winloop(MySpider, concurrency=32)
Requires:
pip install silkworm-rs[winloop]
Trio¶
run_spider_trio uses trio + trio-asyncio for those who prefer trio semantics.
from silkworm import run_spider_trio
run_spider_trio(MySpider, concurrency=16)
Requires: Python 3.13 and
pip install silkworm-rs[trio]. trio-asyncio 0.16 is not compatible with Python 3.14 or newer.
Engine Direct Usage¶
If you want to manage the lifecycle directly:
from silkworm.engine import Engine
from silkworm import Response, Spider
class CustomSpider(Spider):
start_urls = ("https://example.com",)
async def parse(self, response: Response):
return None
spider = CustomSpider(name="custom")
engine = Engine(spider, concurrency=4)
# await engine.run()
Engine.run() opens middlewares, the spider, and pipelines (open_spider()), processes the queue until it is empty, then closes everything in reverse order (close_spider()) and logs final stats. Exceptions from requests, callbacks, middlewares, or pipelines propagate to the caller after being logged.
Engine details: src/silkworm/engine.py