Overview
Image modality transformer implementations.
This package provides transformers that convert image inputs into text outputs using either a single converter or a routed multi-converter approach.
Submodules
generic_image_modality_transformer-- Single-converter transformer.standard_image_modality_transformer-- Router-based multi-converter transformer.
Exported Classes
GenericImageModalityTransformer-- Single-converter image transformer.
GenericImageModalityTransformer(converter)
Bases: ImageModalityTransformer
A generic transformer which uses a single converter to convert images to text.
Attributes:
| Name | Type | Description |
|---|---|---|
converter |
BaseModalityConverter
|
The converter used to transform images into text. |
Initializes the transformer with a modality converter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
converter
|
BaseModalityConverter
|
An instance of a modality converter responsible for performing the image-to-text or image-to-bytes transformation. |
required |
from_config(converter_config=None)
classmethod
Create an instance from a converter configuration dict.
This function
- Build the converter from the converter_config.
- Create the GenericImageModalityTransformer instance.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
converter_config
|
dict[str, ConverterConfig] | None
|
Per-converter configurations keyed by converter name. Defaults to None. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
GenericImageModalityTransformer |
GenericImageModalityTransformer
|
An instance configured from the given dict. |
transform(source, query=None, **kwargs)
async
Transform image input with the single configured converter pipeline.
Behavior
- Validates source input type as bytes or string.
- Calls the configured converter with optional query context.
- Maps converter output into
TransformResultfields. - Extracts normalized captions and text representation for downstream routing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
bytes | str
|
Image source to transform. |
required |
query
|
str | None
|
Optional contextual query. Defaults to None. |
None
|
**kwargs
|
Any
|
Additional converter arguments forwarded as-is to
|
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
TransformResult |
TransformResult
|
Unified transformer output. |
Example
result = await transformer.transform(
source="invoice.png",
query="extract key entities",
text_context="Document is an Indonesian supplier invoice.",
formatter_kwargs={"include_markdown_table": True},
event_emitter=my_event_emitter,
)
print(result.result)