silkworm.cli

Command-line interface: silkworm crawl, parse, fetch, version.

Examples:

silkworm crawl examples/quotes_spider.py -o data/quotes.jl -s max_items=50
silkworm crawl myproject.spiders:QuotesSpider -o - -a category=books
silkworm parse https://quotes.toscrape.com/ --spider examples/quotes_spider.py
silkworm fetch https://example.com/ --headers

A spider reference is a Python file or an importable module, optionally followed by :ClassName when the module defines several spiders.

Exit codes: 0 success, 1 failure policy violated or a callback failed (parse), 2 usage or spider-loading error, 130 interrupted.

exception silkworm.cli.CliError[source]

Bases: Exception

A user-facing command-line error (exit code 2).

silkworm.cli.load_spider_class(reference)[source]

Return the spider class named by path_or_module[:ClassName].

Without ClassName the module must define exactly one spider class (imported spiders are ignored); otherwise the choices are listed.

Raises:

CliError – If the module cannot be loaded or no single spider matches.

Parameters:

reference (str)

Return type:

type[Spider]

silkworm.cli.output_pipeline(target)[source]

Return the pipeline writing items to target (- means stdout).

The format follows the file extension: .jl/.jsonl/.ndjson, .csv, .xml, .db/.sqlite/.sqlite3, .msgpack, .parquet, .yaml/.yml, or .xlsx.

Parameters:

target (str)

Return type:

ItemPipeline

silkworm.cli.build_parser()[source]

Return the silkworm argument parser.

Return type:

ArgumentParser

silkworm.cli.main(argv=None)[source]

Run the command line and return its exit code.

Parameters:

argv (Sequence[str] | None)

Return type:

int