silkworm.testing¶
Helpers for testing spider callbacks offline, without an engine or network.
Build responses from strings or saved files, run a callback exactly as the
engine would (same emit/follow scope and callback contract checks), and
assert on what it produced:
from silkworm.testing import response_from_file, run_callback
async def test_parse_extracts_quotes():
response = response_from_file(
"tests/fixtures/quotes.html", url="https://quotes.toscrape.com/"
)
result = await run_callback(QuotesSpider().parse, response)
assert result.items[0] == {"text": "...", "author": "Albert Einstein"}
assert [r.url for r in result.requests] == [
"https://quotes.toscrape.com/page/2/"
]
- class silkworm.testing.CallbackResult[source]¶
Bases:
objectItems and requests a callback produced.
- requests¶
Requests passed to
follow()(URLs already resolved and callbacks inherited), in order.
- silkworm.testing.html_response(body, url=DEFAULT_URL, *, status=200, headers=None, callback=None, meta=None, html_max_size_bytes=DEFAULT_HTML_MAX_SIZE_BYTES)[source]¶
Return an
HTMLResponseforbodyserved aturl.callbackandmetapopulate the originating request, so followed links inherit the callback exactly as during a crawl.
- silkworm.testing.response_from_file(path, url=DEFAULT_URL, *, status=200, headers=None, callback=None, meta=None, html_max_size_bytes=DEFAULT_HTML_MAX_SIZE_BYTES)[source]¶
Return a response whose body is the bytes of
path.Files ending in
.html/.htm(or whose content looks like HTML) becomeHTMLResponse; JSON, XML, and other files become plainResponseobjects, as the HTTP client would return.
- async silkworm.testing.run_callback(callback, response)[source]¶
Run
callback(response)and return what it emitted and followed.Exceptions raised by the callback propagate unchanged, so tests see the real traceback. A callback that yields, is synchronous, or returns a value raises
SpiderError, as in a crawl.- Parameters:
callback (Callback)
response (Response)
- Return type: