Skip to content

Overview

Image-to-OCR modules for transcribing text from images and documents.

This package provides the abstract base class and a concrete LM-based implementation for OCR transcription of images and PDF documents.

Exported Classes

BaseImageToOcr()

Bases: BaseModalityConverter

Abstract base class for image and document OCR operations.

Extends BaseModalityConverter to support PDF (and multi-page TIFF via image/*) in addition to standard images, and defines the three-layer hook chain: convert → _convert → _get_ocr_result

Subclasses must implement _get_ocr_result to provide the actual text extraction logic (LM-based or specialized engine).

Initialize the base OCR converter with logging capabilities.

convert(source, **kwargs) async

Load a document or image source and run OCR.

Overrides BaseModalityConverter.convert to load via get_media_binary with DocumentConstants.ALLOWED_MIME_TYPES (images and PDF). The loaded binary is packaged as an Attachment so filename and mime_type travel with the data into _convert / _get_ocr_result.

Parameters:

Name Type Description Default
source str | bytes

The document or image source. Accepts file paths, URLs, S3 URIs, Google Drive URLs, raw bytes, and base64-encoded strings.

required
**kwargs Any

Additional keyword arguments forwarded to _convert and _get_ocr_result.

{}

Returns:

Name Type Description
TextResult TextResult

OCR result with tag ConverterResultTag.OCR. On source loading failure, returns an empty TextResult.

Raises:

Type Description
TypeError

If source is not a str or bytes.

ValueError

If source is an empty string.

GLAIRVisionImageToOcr(username, password, api_key, endpoint=DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT, include_detected_objects=False, timeout=DEFAULT_TIMEOUT_SECONDS)

Bases: BaseImageToOcr

Managed OCR implementation backed by the GLAIR Vision OCR API.

Sends one authenticated multipart request per conversion to a GLAIR Vision OCR endpoint and maps the validated response into OcrResult: ordered text lines and, when include_detected_objects is True, detected_objects holding one object per text line (class text or handwriting, with confidence and one-based page number) followed by one object per table (class table, with the table rendered as HTML). Tables are always requested from the API.

Usage

from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.glair_vision_image_to_ocr import (
    GLAIRVisionImageToOcr,
)

converter = GLAIRVisionImageToOcr(
    username="username",
    password="password",
    api_key="api-key",
)
result = await converter.convert("invoice.png")
print(result.result)
for detected_object in result.metadata.detected_objects:
    if detected_object.class_name == DetectedObjectClass.TABLE:
        print(detected_object.metadata["page_number"], detected_object.content)
    else:
        print(detected_object.metadata["page_number"], detected_object.content, detected_object.confidence)

Initializes the GLAIRVisionImageToOcr instance with GLAIR Vision OCR credentials.

Parameters:

Name Type Description Default
username str

Basic-auth username for the GLAIR Vision OCR API.

required
password str

Basic-auth password for the GLAIR Vision OCR API.

required
api_key str

Value sent as the x-api-key request header.

required
endpoint str

Full GLAIR Vision OCR endpoint URL. Defaults to DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT.

DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT
include_detected_objects bool

Whether to populate OcrResult.detected_objects with the recognized text lines and tables and set the image dimensions. When False, line and table geometry is neither parsed nor validated. Defaults to False.

False
timeout int | float

Request timeout in seconds, applied to each attempt. Defaults to 300.

DEFAULT_TIMEOUT_SECONDS

Raises:

Type Description
ValueError

If username, password, api_key, or endpoint is blank, or if timeout is not positive.

ImageOcrPreset

Bases: StrEnum

Available presets for LM-based OCR transcription.

Attributes:

Name Type Description
DEFAULT

Free-text transcription of images and PDF documents.

LMBasedImageToOcr(lm_request_processor, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE)

Bases: BaseImageToOcr, UsesLM

LM-based implementation for transcribing text from images and documents via OCR.

This class leverages a language model (LM) pipeline to transcribe visible text from image and PDF sources into free text, reusing the source loading, attachment construction, and result wrapping defined by BaseImageToOcr.

Inherits

BaseImageToOcr: Base class defining the image/document-to-OCR interface. UsesLM: Mixin providing shared logic for components using language models.

Attributes:

Name Type Description
lm_request_processor LMRequestProcessor

Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the OCR pipeline.

include_bounding_boxes bool

Whether the LM output is parsed as the bounding box JSON object and returned as OcrResult.detected_objects.

bounding_box_range tuple[int, int]

The (min_coordinate, max_coordinate) integer range the LM uses for bounding box coordinates, used to scale the returned boxes to pixel coordinates.

Usage

Use from_preset to easily instantiate with predefined configuration:

from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.lm_based_image_to_ocr import (
    LMBasedImageToOcr,
)

converter = LMBasedImageToOcr.from_preset()

Initializes the LMBasedImageToOcr instance with a language model request processor.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

The processor handling prompt creation, LM invocation, and output parsing.

required
include_bounding_boxes bool

Whether the LM output is parsed as the bounding box JSON object. Defaults to False.

False
bounding_box_range tuple[int, int]

The (min_coordinate, max_coordinate) integer range the LM was asked to use for bounding box coordinates. It must match the range in the prompt, which from_preset keeps in sync. Defaults to (0, 1000).

DEFAULT_BOX_2D_RANGE

Raises:

Type Description
ValueError

If bounding_box_range is not a pair of integers with the minimum lower than the maximum.

convert(source, **kwargs) async

Run LM-based OCR transcription on a loaded image or document source.

This is a documented pass-through to BaseImageToOcr.convert: it does not reload media, reconstruct attachments, or reimplement source validation and fallback handling. Those behaviors, along with TextResult wrapping, are owned by the base class.

Parameters:

Name Type Description Default
source str | bytes

The document or image source to transcribe. Supported forms: 1. Raw image or PDF bytes. 2. Local file path. 3. URL, S3 URI, or Google Drive URL. 4. Base64 encoded string.

required
**kwargs Any

See Other Parameters below.

{}

Other Parameters:

Name Type Description
multimodal_context list[Attachment | bytes | str]

Reference context sent after the document to help the LM read ambiguous characters, names, or terms. Context is never transcribed. Bytes are converted into Attachments, strings that resolve to an image or PDF source are loaded as Attachments, and other strings are sent as text. Defaults to [].

event_emitter Any

Event emitter forwarded to LM request processing.

hyperparameters dict[str, Any]

Hyperparameters forwarded to the LM invocation.

Returns:

Name Type Description
TextResult TextResult

OCR result with tag ConverterResultTag.OCR and OcrResult metadata. On source loading failure, returns an empty TextResult. On LM failure, returns an OCR-tagged TextResult whose metadata is OcrResult(text="").

Raises:

Type Description
TypeError

If source is not a str or bytes, or if multimodal_context is not a list of Attachment, bytes, or str items.

ValueError

If source is an empty string.

Example
converter = LMBasedImageToOcr.from_preset()
result = await converter.convert(source="manual.pdf")
print(result.result)

result = await converter.convert(
    source="handwritten_form.png",
    multimodal_context=["Product names: Glair, Catapa.", "reference_page.png"],
)

from_preset(preset_name=ImageOcrPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE, **kwargs) classmethod

Initializes the LM-based OCR component using preset model configurations.

Parameters:

Name Type Description Default
preset_name ImageOcrPreset | str | None

Preset name forwarded to the image-OCR preset registry. Defaults to ImageOcrPreset.DEFAULT.

DEFAULT
lm_invoker_kwargs dict[str, Any] | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict[str, Any] | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
include_bounding_boxes bool

Whether the LM is additionally asked for the bounding box of every line of text. Defaults to False.

False
bounding_box_range tuple[int, int]

The (min_coordinate, max_coordinate) integer range the LM is asked to use for bounding box coordinates. It is written into the prompt and used to scale the returned boxes to pixels, so both always match. Defaults to (0, 1000).

DEFAULT_BOX_2D_RANGE
**kwargs Any

Additional kwargs for current-class __init__ parameters. from_lm_components forwards these kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_image_to_ocr. 2. Final converter constructor parameters: LMBasedImageToOcr.

Returns:

Name Type Description
LMBasedImageToOcr LMBasedImageToOcr

Initialized OCR component using the preset model.

Raises:

Type Description
ValueError

If bounding_box_range is not a pair of integers with the minimum lower than the maximum.

Example
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr import (
    LMBasedImageToOcr,
    ImageOcrPreset,
)

# Use the default preset (Gemini flash-lite).
ocr = LMBasedImageToOcr.from_preset()

# Override the model via lm_invoker_kwargs.
ocr = LMBasedImageToOcr.from_preset(
    preset_name=ImageOcrPreset.DEFAULT,
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await ocr.convert(source="invoice.png")
print(result.result)

# Also return the pixel bounding box of every line of text.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True)
result = await ocr.convert(source="invoice.png")
for detected_object in result.metadata.detected_objects:
    print(detected_object.content, detected_object.to_bounding_box())

# Ask for bounding box coordinates on a different integer range.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True, bounding_box_range=(0, 999))