Skip to content

Overview

Image modality transformer implementations.

This package provides transformers that convert image inputs into text outputs using either a single converter or a routed multi-converter approach.

Submodules

Exported Classes

GenericImageModalityTransformer(converter)

Bases: ImageModalityTransformer

A generic transformer which uses a single converter to convert images to text.

Attributes:

Name Type Description
converter BaseModalityConverter

The converter used to transform images into text.

Initializes the transformer with a modality converter.

Parameters:

Name Type Description Default
converter BaseModalityConverter

An instance of a modality converter responsible for performing the image-to-text or image-to-bytes transformation.

required

from_config(converter_config=None) classmethod

Create an instance from a converter configuration dict.

This function
  1. Build the converter from the converter_config.
  2. Create the GenericImageModalityTransformer instance.

Parameters:

Name Type Description Default
converter_config dict[str, ConverterConfig] | None

Per-converter configurations keyed by converter name. Defaults to None.

None

Returns:

Name Type Description
GenericImageModalityTransformer GenericImageModalityTransformer

An instance configured from the given dict.

transform(source, query=None, **kwargs) async

Transform image input with the single configured converter pipeline.

Behavior
  1. Validates source input type as bytes or string.
  2. Calls the configured converter with optional query context.
  3. Maps converter output into TransformResult fields.
  4. Extracts normalized captions and text representation for downstream routing.

Parameters:

Name Type Description Default
source bytes | str

Image source to transform.

required
query str | None

Optional contextual query. Defaults to None.

None
**kwargs Any

Additional converter arguments forwarded as-is to self.converter.convert. Typical examples include: 1. Provider-specific generation options. 2. Caption or formatter arguments supported by the selected converter. 3. Runtime controls such as event_emitter when supported.

{}

Returns:

Name Type Description
TransformResult TransformResult

Unified transformer output.

Example
result = await transformer.transform(
    source="invoice.png",
    query="extract key entities",
    text_context="Document is an Indonesian supplier invoice.",
    formatter_kwargs={"include_markdown_table": True},
    event_emitter=my_event_emitter,
)
print(result.result)