silkworm.declarative

Declarative item extraction from HTMLResponse objects.

Subclass Item and assign Text or Attr descriptors to compile a reusable extraction plan with validation, transforms, and nested items.

silkworm.declarative.Attr(selector: str, name: str, *, absolute: bool = False, transform: Callable[[str], T], default: T | None | _Missing = MISSING) → Any[source]
silkworm.declarative.Attr(selector: str, name: str, *, absolute: bool = False, transform: None = None, default: str | None | _Missing = MISSING) → Any

Declare a field that extracts an HTML attribute.

Parameters:
  • selector (str) – CSS selector relative to the item root.

  • name (str) – Attribute name to read from each selected element.

  • absolute (bool) – Resolve extracted URLs against the response URL.

  • transform (Callable[[str], object] | None) – Optional synchronous value converter.

  • default (object | _Missing) – Value used when a scalar match or attribute is absent.

Return type:

Any

The field annotation determines scalar, optional, or list cardinality.

exception silkworm.declarative.DeclarativeConfigurationError[source]

Bases: DeclarativeError

Raised when an item declaration cannot be compiled.

exception silkworm.declarative.DeclarativeError[source]

Bases: SilkwormError

Base exception for declarative extraction.

exception silkworm.declarative.DeclarativeSerializationError[source]

Bases: DeclarativeError

Raised when an item cannot be represented as a JSON value.

class silkworm.declarative.ExtractionPlan[source]

Bases: object

Immutable compiled plan shared by all instances of an item class.

item_type

Item subclass described by the plan.

Type:

type[Item]

root_selector

Optional selector producing repeated item roots.

Type:

str | None

fields

Field plans in declaration order.

Type:

tuple[FieldPlan, …]

__init__(item_type, root_selector, fields)
Parameters:
Return type:

None

class silkworm.declarative.Field[source]

Bases: Generic

Base descriptor for a field in a declarative item.

Parameters:
  • selector – Non-empty CSS selector evaluated within the current item root.

  • transform – Synchronous conversion applied to extracted text.

  • default – Value used when a scalar field is absent. Without a default, required fields raise MissingFieldError.

Field is primarily a typing and extension point; applications normally declare fields with Text() or Attr().

__init__(selector, *, transform=None, default=MISSING)[source]
Parameters:
Return type:

None

property name: str

Return the attribute name assigned by the owning item class.

Raises:

RuntimeError – If the descriptor has not been bound to an item class.

property kind: str

Return a human-readable field type used in validation errors.

exception silkworm.declarative.FieldExtractionError[source]

Bases: DeclarativeError

Raised when a declared field cannot be extracted.

class silkworm.declarative.FieldPlan[source]

Bases: object

Compiled extraction rules for one declarative field.

name

Item attribute name.

Type:

str

field

Field descriptor supplying selector and conversion settings.

Type:

silkworm.declarative.plans.FieldSpec

cardinality

Whether extraction expects one, optional, or many values.

Type:

silkworm.declarative.typing.Cardinality

value_type

Runtime type accepted for each extracted value.

Type:

object

annotation

Original resolved item annotation.

Type:

object

__init__(name, field, cardinality, value_type, annotation)
Parameters:
  • name (str)

  • field (FieldSpec)

  • cardinality (Cardinality)

  • value_type (object)

  • annotation (object)

Return type:

None

exception silkworm.declarative.FieldTransformError[source]

Bases: FieldExtractionError

Raised when a field transform fails or returns an invalid value.

class silkworm.declarative.Item[source]

Bases: object

Base class for typed, declaratively extracted records.

Set __selector__ to extract one item per matching root, then annotate attributes assigned with Text() or Attr().

Example

>>> class Quote(Item):
...     __selector__ = ".quote"
...     text: str = Text(".text", strip=True)
...     author: str = Text(".author")
__init__(**values)[source]
Parameters:

values (object)

Return type:

None

classmethod extraction_plan()[source]

Return the cached extraction plan for this item class.

Raises:

DeclarativeConfigurationError – If class declarations are invalid.

Return type:

ExtractionPlan

classmethod extract(response)[source]

Yield one validated item per matching root in response.

Raises:
Parameters:

response (HTMLResponse)

Return type:

AsyncIterator[Self]

async after_extract(response)[source]

Customize an item after extraction and before it is yielded.

Subclasses may mutate fields or perform asynchronous enrichment.

Parameters:

response (HTMLResponse)

Return type:

None

to_dict()[source]

Recursively convert this item into a pipeline-compatible mapping.

Raises:

DeclarativeSerializationError – If a nested value is not JSON-compatible or a mapping key is not a string.

Return type:

dict[str, JSONValue]

exception silkworm.declarative.MissingFieldError[source]

Bases: FieldExtractionError

Raised when a required element or attribute is absent.

silkworm.declarative.Text(selector: str, *, transform: Callable[[str], T], default: T | None | _Missing = MISSING, strip: bool = False) → Any[source]
silkworm.declarative.Text(selector: str, *, transform: None = None, default: str | None | _Missing = MISSING, strip: bool = False) → Any

Declare a field that extracts element text with a CSS selector.

Parameters:
  • selector (str) – CSS selector relative to the item root.

  • transform (Callable[[str], object] | None) – Optional synchronous value converter.

  • default (object | _Missing) – Value used when a scalar match is absent.

  • strip (bool) – Strip leading and trailing whitespace before transforming.

Return type:

Any

The field annotation determines whether one value, an optional value, or a list of matches is extracted.