Skip to content

Overview

Image-to-text conversion module providing OCR and captioning capabilities.

Submodules

GLAIRVisionImageToOcr(username, password, api_key, endpoint=DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT, include_detected_objects=False, timeout=DEFAULT_TIMEOUT_SECONDS)

Bases: BaseImageToOcr

Managed OCR implementation backed by the GLAIR Vision OCR API.

Sends one authenticated multipart request per conversion to a GLAIR Vision OCR endpoint and maps the validated response into OcrResult: ordered text lines and, when include_detected_objects is True, detected_objects holding one object per text line (class text or handwriting, with confidence and one-based page number) followed by one object per table (class table, with the table rendered as HTML). Tables are always requested from the API.

Usage

from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.glair_vision_image_to_ocr import (
    GLAIRVisionImageToOcr,
)

converter = GLAIRVisionImageToOcr(
    username="username",
    password="password",
    api_key="api-key",
)
result = await converter.convert("invoice.png")
print(result.result)
for detected_object in result.metadata.detected_objects:
    if detected_object.class_name == DetectedObjectClass.TABLE:
        print(detected_object.metadata["page_number"], detected_object.content)
    else:
        print(detected_object.metadata["page_number"], detected_object.content, detected_object.confidence)

Initializes the GLAIRVisionImageToOcr instance with GLAIR Vision OCR credentials.

Parameters:

Name Type Description Default
username str

Basic-auth username for the GLAIR Vision OCR API.

required
password str

Basic-auth password for the GLAIR Vision OCR API.

required
api_key str

Value sent as the x-api-key request header.

required
endpoint str

Full GLAIR Vision OCR endpoint URL. Defaults to DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT.

DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT
include_detected_objects bool

Whether to populate OcrResult.detected_objects with the recognized text lines and tables and set the image dimensions. When False, line and table geometry is neither parsed nor validated. Defaults to False.

False
timeout int | float

Request timeout in seconds, applied to each attempt. Defaults to 300.

DEFAULT_TIMEOUT_SECONDS

Raises:

Type Description
ValueError

If username, password, api_key, or endpoint is blank, or if timeout is not positive.

LMBasedImageToCaption(lm_request_processor, formatter=DEFAULT_FORMATTER)

Bases: BaseImageToCaption, UsesLM

Image captioning implementation using Language Models.

This class implements the ImageToCaption interface using LMs for generating natural language captions.

Usage

Use from_preset to easily instantiate with predefined configuration:

from gllm_multimodal.modality_converter.image_to_text.image_to_caption.lm_based_image_to_caption import LMBasedImageToCaption

converter = LMBasedImageToCaption.from_preset()

Initialize the LM based image captioning component.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

Language model request processor instance that supports multimodal inputs.

required
formatter BaseCaptionOutputFormatter

The formatter to use for formatting the output.

DEFAULT_FORMATTER

convert(source, **kwargs) async

Convert image into caption text using the LM captioning pipeline.

Behavior
  1. Validates and loads media bytes from path, URL, base64, or bytes source.
  2. Optionally extracts image metadata and builds the Caption context payload.
  3. Builds prompt params from supported prompt keys and caption fields.
  4. Sends image plus multimodal context attachments to the LM request processor.
  5. Formats the generated caption using the configured output formatter.
  6. Returns formatted caption text and metadata as TextResult.

Parameters:

Name Type Description Default
source str | bytes

Image source to caption. Supported forms: 1. Raw image bytes. 2. Local file path. 3. URL string. 4. Base64 encoded image string.

required
**kwargs Any

See Other Parameters below.

{}

Other Parameters:

Name Type Description
number_of_captions int

Number of captions to generate.

text_one_liner str

Short one-line context/title.

text_context str

Additional descriptive context.

domain_knowledge str

Domain hints for caption quality.

multimodal_context list[Attachment | str]

Additional context attachments.

use_metadata bool

Whether image metadata is extracted and injected.

formatter_kwargs dict[str, Any]

Extra kwargs for formatter output shaping.

event_emitter Any

Event emitter passed to LM request processing.

Returns:

Name Type Description
TextResult TextResult

Formatted caption result with structured metadata.

Example
from gllm_inference.schema import Attachment

converter = LMBasedImageToCaption.from_preset()
result = await converter.convert(
    source="diagram.png",
    number_of_captions=3,
    text_one_liner="System architecture diagram",
    domain_knowledge="cloud networking",
    multimodal_context=[
        "focus on request flow",
        Attachment.from_file("legend.png"),
    ],
    formatter_kwargs={"include_bullets": True, "max_sentences": 6},
    use_metadata=True,
)
print(result.result)

from_preset(preset_name=ImageCaptionPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs) classmethod

Initialize the LM based image captioning component using preset model configurations.

Parameters:

Name Type Description Default
preset_name ImageCaptionPreset | str | None

Preset name forwarded to image-caption preset registry. Defaults to ImageCaptionPreset.DEFAULT.

DEFAULT
lm_invoker_kwargs dict | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
**kwargs Any

Additional kwargs for current-class __init__ parameters. from_lm_components forwards these kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_image_to_caption. 2. Final converter constructor parameters: LMBasedImageToCaption.

Returns:

Name Type Description
LMBasedImageToCaption LMBasedImageToCaption

Initialized image captioning component using preset model.

Example
from gllm_multimodal.modality_converter.image_to_text.image_to_caption import (
    LMBasedImageToCaption,
    ImageCaptionPreset,
)

# Use default preset (Gemini flash-lite)
captioner = LMBasedImageToCaption.from_preset()

# Use structured output preset with a custom model
captioner = LMBasedImageToCaption.from_preset(
    preset_name=ImageCaptionPreset.STRUCTURED,
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await captioner.convert(source="diagram.png")
print(result.result)

LMBasedImageToMermaid(lm_request_processor)

Bases: BaseImageToMermaid, UsesLM

LM-based implementation for converting an image into Mermaid diagram syntax.

This class leverages a language model (LM) pipeline to generate structured Mermaid syntax from image inputs and optional metadata. It uses prompt builders, LM invokers, and output parsers defined via a preset system to streamline model usage.

Inherits

BaseImageToMermaid: Base class defining the image-to-mermaid interface. UsesLM: Mixin providing shared logic for components using language models.

Attributes:

Name Type Description
lm_request_processor LMRequestProcessor

Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the image-to-mermaid pipeline.

Initializes the LMBasedImageToMermaid instance with a language model request processor.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

The processor handling prompt creation, LM invocation, and output parsing.

required

convert(source, **kwargs) async

Convert image input into Mermaid syntax using the LM conversion flow.

Behavior
  1. Validates and loads image bytes from supported source formats.
  2. Builds Mermaid metadata payload from method keyword arguments.
  3. Filters prompt params to keys recognized by the prompt template.
  4. Sends the image as attachment to the LM request processor.
  5. Returns Mermaid syntax text and metadata wrapped in TextResult.

Parameters:

Name Type Description Default
source str | bytes

Image source to convert. Supported forms: 1. Raw image bytes. 2. Local file path. 3. URL string. 4. Base64 encoded image string.

required
**kwargs Any

Mermaid generation options forwarded into metadata and LM calls, including: 1. title (str, optional): Diagram title context. 2. description (str, optional): Additional scene or flow context. 3. event_emitter (Any, optional): Event emitter passed to LM request processing. 4. Any Mermaid schema fields recognized by Mermaid model and prompt template.

{}

Returns:

Name Type Description
TextResult TextResult

Mermaid syntax output with associated metadata.

Example
converter = LMBasedImageToMermaid.from_preset()
result = await converter.convert(
    source="workflow.png",
    title="Ticket Escalation Flow",
    description="Extract decision points and handoff states",
    diagram_type="flowchart",
    direction="LR",
    event_emitter=my_event_emitter,
)
print(result.result)

from_preset(preset_name='default', lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs) classmethod

Constructs an LMBasedImageToMermaid instance using a named preset configuration.

Parameters:

Name Type Description Default
preset_name str | None

Name of the predefined preset configuration to use. Defaults to "default".

'default'
lm_invoker_kwargs dict | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
**kwargs Any

Additional kwargs for current-class __init__ parameters. from_lm_components forwards these kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_image_to_mermaid. 2. Final converter constructor parameters: LMBasedImageToMermaid.

Returns:

Name Type Description
LMBasedImageToMermaid LMBasedImageToMermaid

An instance initialized with the preset's components.

Example
from gllm_multimodal.modality_converter.image_to_text.image_to_mermaid import (
    LMBasedImageToMermaid,
)

# Use default preset (Gemini flash-lite)
converter = LMBasedImageToMermaid.from_preset()

# Override model via lm_invoker_kwargs
converter = LMBasedImageToMermaid.from_preset(
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await converter.convert(source="architecture.png")
print(result.result)

LMBasedImageToOcr(lm_request_processor, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE)

Bases: BaseImageToOcr, UsesLM

LM-based implementation for transcribing text from images and documents via OCR.

This class leverages a language model (LM) pipeline to transcribe visible text from image and PDF sources into free text, reusing the source loading, attachment construction, and result wrapping defined by BaseImageToOcr.

Inherits

BaseImageToOcr: Base class defining the image/document-to-OCR interface. UsesLM: Mixin providing shared logic for components using language models.

Attributes:

Name Type Description
lm_request_processor LMRequestProcessor

Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the OCR pipeline.

include_bounding_boxes bool

Whether the LM output is parsed as the bounding box JSON object and returned as OcrResult.detected_objects.

bounding_box_range tuple[int, int]

The (min_coordinate, max_coordinate) integer range the LM uses for bounding box coordinates, used to scale the returned boxes to pixel coordinates.

Usage

Use from_preset to easily instantiate with predefined configuration:

from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.lm_based_image_to_ocr import (
    LMBasedImageToOcr,
)

converter = LMBasedImageToOcr.from_preset()

Initializes the LMBasedImageToOcr instance with a language model request processor.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

The processor handling prompt creation, LM invocation, and output parsing.

required
include_bounding_boxes bool

Whether the LM output is parsed as the bounding box JSON object. Defaults to False.

False
bounding_box_range tuple[int, int]

The (min_coordinate, max_coordinate) integer range the LM was asked to use for bounding box coordinates. It must match the range in the prompt, which from_preset keeps in sync. Defaults to (0, 1000).

DEFAULT_BOX_2D_RANGE

Raises:

Type Description
ValueError

If bounding_box_range is not a pair of integers with the minimum lower than the maximum.

convert(source, **kwargs) async

Run LM-based OCR transcription on a loaded image or document source.

This is a documented pass-through to BaseImageToOcr.convert: it does not reload media, reconstruct attachments, or reimplement source validation and fallback handling. Those behaviors, along with TextResult wrapping, are owned by the base class.

Parameters:

Name Type Description Default
source str | bytes

The document or image source to transcribe. Supported forms: 1. Raw image or PDF bytes. 2. Local file path. 3. URL, S3 URI, or Google Drive URL. 4. Base64 encoded string.

required
**kwargs Any

See Other Parameters below.

{}

Other Parameters:

Name Type Description
multimodal_context list[Attachment | bytes | str]

Reference context sent after the document to help the LM read ambiguous characters, names, or terms. Context is never transcribed. Bytes are converted into Attachments, strings that resolve to an image or PDF source are loaded as Attachments, and other strings are sent as text. Defaults to [].

event_emitter Any

Event emitter forwarded to LM request processing.

hyperparameters dict[str, Any]

Hyperparameters forwarded to the LM invocation.

Returns:

Name Type Description
TextResult TextResult

OCR result with tag ConverterResultTag.OCR and OcrResult metadata. On source loading failure, returns an empty TextResult. On LM failure, returns an OCR-tagged TextResult whose metadata is OcrResult(text="").

Raises:

Type Description
TypeError

If source is not a str or bytes, or if multimodal_context is not a list of Attachment, bytes, or str items.

ValueError

If source is an empty string.

Example
converter = LMBasedImageToOcr.from_preset()
result = await converter.convert(source="manual.pdf")
print(result.result)

result = await converter.convert(
    source="handwritten_form.png",
    multimodal_context=["Product names: Glair, Catapa.", "reference_page.png"],
)

from_preset(preset_name=ImageOcrPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE, **kwargs) classmethod

Initializes the LM-based OCR component using preset model configurations.

Parameters:

Name Type Description Default
preset_name ImageOcrPreset | str | None

Preset name forwarded to the image-OCR preset registry. Defaults to ImageOcrPreset.DEFAULT.

DEFAULT
lm_invoker_kwargs dict[str, Any] | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict[str, Any] | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
include_bounding_boxes bool

Whether the LM is additionally asked for the bounding box of every line of text. Defaults to False.

False
bounding_box_range tuple[int, int]

The (min_coordinate, max_coordinate) integer range the LM is asked to use for bounding box coordinates. It is written into the prompt and used to scale the returned boxes to pixels, so both always match. Defaults to (0, 1000).

DEFAULT_BOX_2D_RANGE
**kwargs Any

Additional kwargs for current-class __init__ parameters. from_lm_components forwards these kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_image_to_ocr. 2. Final converter constructor parameters: LMBasedImageToOcr.

Returns:

Name Type Description
LMBasedImageToOcr LMBasedImageToOcr

Initialized OCR component using the preset model.

Raises:

Type Description
ValueError

If bounding_box_range is not a pair of integers with the minimum lower than the maximum.

Example
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr import (
    LMBasedImageToOcr,
    ImageOcrPreset,
)

# Use the default preset (Gemini flash-lite).
ocr = LMBasedImageToOcr.from_preset()

# Override the model via lm_invoker_kwargs.
ocr = LMBasedImageToOcr.from_preset(
    preset_name=ImageOcrPreset.DEFAULT,
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await ocr.convert(source="invoice.png")
print(result.result)

# Also return the pixel bounding box of every line of text.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True)
result = await ocr.convert(source="invoice.png")
for detected_object in result.metadata.detected_objects:
    print(detected_object.content, detected_object.to_bounding_box())

# Ask for bounding box coordinates on a different integer range.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True, bounding_box_range=(0, 999))