Command Line¶
Installing silkworm-rs provides a silkworm command for running and debugging
spiders without writing a runner script.
silkworm --version
silkworm --log-level DEBUG crawl ... # --log-level applies to every command
Spider references¶
Commands that take a spider accept a Python file or an importable module,
optionally followed by :ClassName:
silkworm crawl examples/quotes_spider.py
silkworm crawl myproject.spiders.news
silkworm crawl myproject/spiders.py:NewsSpider
Without a class name the module must define exactly one spider class (spiders it
imports are ignored). Files are loaded with their directory on sys.path, so
sibling modules can be imported.
silkworm crawl¶
silkworm crawl SPIDER [-o FILE]... [-s NAME=VALUE]... [-a NAME=VALUE]...
[--job-dir DIR] [--http-cache DIR] [--loop LOOP]
-o FILEwrites items to a file; the format follows the extension:.jl,.jsonl,.ndjson,.csv,.xml,.db/.sqlite/.sqlite3,.msgpack,.parquet,.yaml/.yml, or.xlsx. Use-o -for JSON lines on stdout. Repeat-oto write several formats; parent directories are created.-s NAME=VALUEsets an engine option (any scalar option, see Settings), e.g.-s concurrency=8 -s max_items=100. These overrideSILKWORM_*environment variables andcustom_settings.-a NAME=VALUEpasses a keyword argument to the spider’s constructor (as a string).--job-dir DIRpersists state so an interrupted crawl resumes (details).--http-cache DIRcaches responses on disk (details).--looppicksasyncio(default),uvloop,rsloop, orwinloop.
The first Ctrl+C or SIGTERM stops gracefully; a second one cancels immediately. A summary line is printed to stderr at the end.
silkworm crawl examples/quotes_spider.py -o data/quotes.jl -o data/quotes.csv \
-s max_error_rate=0.1 -s min_items=50 --job-dir state/quotes
silkworm parse¶
Fetches one URL, runs it through a spider callback, and prints what the callback
produced as JSON lines ({"type": "item", ...} and {"type": "request", ...}).
Handy while writing selectors:
silkworm parse https://quotes.toscrape.com/ --spider examples/quotes_spider.py
silkworm parse https://example.com/product/1 --spider shop.py -c parse_product \
--meta '{"category": "books"}'
The spider’s open()/close() hooks run around the call, and parse receives an
HTMLResponse just as in a crawl.
silkworm fetch¶
Downloads a URL with the default client and writes the body to stdout (or a file
with -o); --headers prints the status and headers to stderr.
silkworm fetch https://example.com/ --headers -o page.html
Exit codes¶
Code |
Meaning |
|---|---|
|
Success. |
|
The crawl violated its failure policy ( |
|
Usage error: bad arguments or settings, or a spider that cannot be loaded. |
|
Interrupted by a signal. |