Getting Started¶
Requirements¶
Python: 3.13, 3.14 or 3.15. On Python 3.15 the
msgpack,vortexandonionlinkextras install without their backing library until it ships 3.15 wheels, so those pipelines/clients are unavailable there for now.run_spider_triois available on Python 3.13 only until trio-asyncio supports Python 3.14+ (thetrioextra skips trio-asyncio there). Theiggyandcassandraextras are limited to Python 3.13.Project metadata: pyproject.toml
Installation¶
From PyPI (pip)¶
pip install silkworm-rs
From PyPI (uv)¶
uv pip install silkworm-rs
# or if using uv project management
uv add silkworm-rs
From Source¶
uv venv --python python3.13
source .venv/bin/activate
uv pip install -e .
Optional Extras¶
Silkworm ships many integrations behind extras. Install only what you need.
pip install "silkworm-rs[rsloop,polars]"
Extra |
Purpose |
Related Code |
|---|---|---|
|
Faster event loop via |
|
|
Faster event loop on Unix |
|
|
Faster event loop on Windows |
|
|
Trio backend |
|
|
MsgPack export |
|
|
Parquet export |
|
|
Excel export |
|
|
YAML export |
|
|
Avro export |
|
|
Elasticsearch export |
|
|
MongoDB export |
|
|
S3 export (OpenDAL) |
|
|
Vortex export |
|
|
MySQL export |
|
|
PostgreSQL export |
|
|
Google Sheets export |
|
|
Snowflake export |
|
|
FTP export |
|
|
SFTP export |
|
|
Cassandra export (not on Windows) |
|
|
CouchDB export |
|
|
DynamoDB export |
|
|
DuckDB export |
|
|
Taskiq queue pipeline |
|
|
Zenoh pub/sub pipeline |
|
|
Apache Iggy message-streaming pipeline (Python 3.13) |
|
|
Memory profiling |
|
|
CDP browser client and rendered HTML fetch |
|
|
Tor v3 onion-service client |
OnionLinkClient is also available as an optional integration for .onion sites:
pip install "silkworm-rs[onionlink]"
ServoFetchClient is available from silkworm, but servofetch is distributed separately from the package extras. Install a compatible wheel from the servofetch releases before using fetch_html_servo or ServoFetchClient.
Your First Spider¶
This is a minimal spider that extracts quotes and writes JSON Lines output. For a full version, see examples/quotes_spider.py.
from silkworm import HTMLResponse, Response, Spider, run_spider
from silkworm.middlewares import RetryMiddleware, UserAgentMiddleware
from silkworm.pipelines import JsonLinesPipeline
class QuotesSpider(Spider):
name = "quotes"
start_urls = ("https://quotes.toscrape.com/",)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
html = response
for el in await html.select(".quote"):
text_el = await el.select_first(".text")
author_el = await el.select_first(".author")
if text_el is None or author_el is None:
continue
tags = await el.select(".tag")
await self.emit(
{
"text": text_el.text,
"author": author_el.text,
"tags": [t.text for t in tags],
}
)
if next_link := await html.select_first("li.next > a"):
if href := next_link.attr("href"):
await html.follow(href, callback=self.parse)
run_spider(
QuotesSpider,
request_middlewares=[UserAgentMiddleware()],
response_middlewares=[RetryMiddleware(max_times=3)],
item_pipelines=[JsonLinesPipeline("data/quotes.jl")],
concurrency=16,
request_timeout=10,
log_stats_interval=30,
)
Running Examples¶
Examples are in examples/. A few popular ones:
python examples/quotes_spider.py
python examples/quotes_spider_xpath.py
python examples/url_titles_spider.py --urls-file data/url_titles.jl --output data/titles.jl
See Examples for a full list and what each one demonstrates.
One-off HTML Fetch¶
For quick, standalone fetches, use fetch_html in src/silkworm/api.py. For rendered pages, install servofetch and use fetch_html_servo, or pass ServoFetchClient as http_client:
from silkworm import ServoFetchClient, run_spider
run_spider(MySpider, http_client=ServoFetchClient(settle_ms=500))
Servo request overrides use Request.meta: servo_javascript, servo_settle_ms, servo_user_agent, servo_screenshot, and servo_full_page.
import asyncio
from silkworm import fetch_html
async def main():
text, doc = await fetch_html("https://example.com")
title = await doc.select_first("title")
print(title.text if title else "no title")
asyncio.run(main())
For JS-rendered pages via a CDP-compatible browser (Lightpanda/Chrome), use
fetch_html_cdp:
import asyncio
from silkworm import fetch_html_cdp
async def main():
text, doc = await fetch_html_cdp(
"https://example.com",
ws_endpoint="ws://127.0.0.1:9222",
)
title = await doc.select_first("title")
print(title.text if title else "no title")
asyncio.run(main())
Running From the Command Line¶
The silkworm command runs a spider file without a runner script and writes
items in the format given by the output extension:
silkworm crawl quotes_spider.py -o data/quotes.jl -s max_items=100
silkworm parse https://quotes.toscrape.com/ --spider quotes_spider.py
See Command Line, and Production Crawling before scheduling spiders unattended.
Development Workflow¶
Developer commands are defined in the justfile.
just help
just fmt
just lint
just typecheck
just test
just docs
`just docs` builds this documentation into `docs/_build/html` with the same settings the GitHub Pages deployment uses.