silkworm.testing

Helpers for testing spider callbacks offline, without an engine or network.

Build responses from strings or saved files, run a callback exactly as the engine would (same emit/follow scope and callback contract checks), and assert on what it produced:

from silkworm.testing import response_from_file, run_callback

async def test_parse_extracts_quotes():
    response = response_from_file(
        "tests/fixtures/quotes.html", url="https://quotes.toscrape.com/"
    )
    result = await run_callback(QuotesSpider().parse, response)

    assert result.items[0] == {"text": "...", "author": "Albert Einstein"}
    assert [r.url for r in result.requests] == [
        "https://quotes.toscrape.com/page/2/"
    ]
class silkworm.testing.CallbackResult[source]

Bases: object

Items and requests a callback produced.

items

Items passed to emit(), in order.

Type:

list[JSONLike]

requests

Requests passed to follow() (URLs already resolved and callbacks inherited), in order.

Type:

list[Request]

property urls: list[str]

Return the URLs of requests.

__init__(items=<factory>, requests=<factory>)
Parameters:
Return type:

None

silkworm.testing.html_response(body, url=DEFAULT_URL, *, status=200, headers=None, callback=None, meta=None, html_max_size_bytes=DEFAULT_HTML_MAX_SIZE_BYTES)[source]

Return an HTMLResponse for body served at url.

callback and meta populate the originating request, so followed links inherit the callback exactly as during a crawl.

Parameters:
  • body (str | bytes)

  • url (str)

  • status (int)

  • headers (dict[str, str] | None)

  • callback (Callback | None)

  • meta (MetaData | None)

  • html_max_size_bytes (int)

Return type:

HTMLResponse

silkworm.testing.response_from_file(path, url=DEFAULT_URL, *, status=200, headers=None, callback=None, meta=None, html_max_size_bytes=DEFAULT_HTML_MAX_SIZE_BYTES)[source]

Return a response whose body is the bytes of path.

Files ending in .html/.htm (or whose content looks like HTML) become HTMLResponse; JSON, XML, and other files become plain Response objects, as the HTTP client would return.

Parameters:
Return type:

Response

async silkworm.testing.run_callback(callback, response)[source]

Run callback(response) and return what it emitted and followed.

Exceptions raised by the callback propagate unchanged, so tests see the real traceback. A callback that yields, is synchronous, or returns a value raises SpiderError, as in a crawl.

Parameters:
  • callback (Callback)

  • response (Response)

Return type:

CallbackResult

async silkworm.testing.run_errback(errback, request, exception)[source]

Run errback(request, exception) and return what it produced.

Parameters:
Return type:

CallbackResult

async silkworm.testing.run_start_requests(spider)[source]

Return the requests spider.start_requests() schedules.

Parameters:

spider (Spider)

Return type:

list[Request]