Overview
Image-to-OCR modules for transcribing text from images and documents.
This package provides the abstract base class and a concrete LM-based implementation for OCR transcription of images and PDF documents.
Exported Classes
BaseImageToOcr-- Abstract base class for image/document OCR.LMBasedImageToOcr-- LM-based OCR transcription.GLAIRVisionImageToOcr-- GLAIR Vision OCR-based OCR transcription.ImageOcrPreset-- Preset-based OCR configuration.
BaseImageToOcr()
Bases: BaseModalityConverter
Abstract base class for image and document OCR operations.
Extends BaseModalityConverter to support PDF (and multi-page TIFF via
image/*) in addition to standard images, and defines the three-layer
hook chain:
convert → _convert → _get_ocr_result
Subclasses must implement _get_ocr_result to provide the actual
text extraction logic (LM-based or specialized engine).
Initialize the base OCR converter with logging capabilities.
convert(source, **kwargs)
async
Load a document or image source and run OCR.
Overrides BaseModalityConverter.convert to load via get_media_binary with DocumentConstants.ALLOWED_MIME_TYPES (images and PDF). The loaded binary is packaged as an Attachment so filename and mime_type travel with the data into _convert / _get_ocr_result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
The document or image source. Accepts file paths, URLs, S3 URIs, Google Drive URLs, raw bytes, and base64-encoded strings. |
required |
**kwargs
|
Any
|
Additional keyword arguments forwarded to _convert and _get_ocr_result. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
OCR result with tag |
Raises:
| Type | Description |
|---|---|
TypeError
|
If source is not a str or bytes. |
ValueError
|
If source is an empty string. |
GLAIRVisionImageToOcr(username, password, api_key, endpoint=DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT, include_detected_objects=False, timeout=DEFAULT_TIMEOUT_SECONDS)
Bases: BaseImageToOcr
Managed OCR implementation backed by the GLAIR Vision OCR API.
Sends one authenticated multipart request per conversion to a GLAIR Vision OCR endpoint and maps the
validated response into OcrResult: ordered text lines and, when include_detected_objects is True,
detected_objects holding one object per text line (class text or handwriting, with confidence and
one-based page number) followed by one object per table (class table, with the table rendered as HTML).
Tables are always requested from the API.
Usage
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.glair_vision_image_to_ocr import (
GLAIRVisionImageToOcr,
)
converter = GLAIRVisionImageToOcr(
username="username",
password="password",
api_key="api-key",
)
result = await converter.convert("invoice.png")
print(result.result)
for detected_object in result.metadata.detected_objects:
if detected_object.class_name == DetectedObjectClass.TABLE:
print(detected_object.metadata["page_number"], detected_object.content)
else:
print(detected_object.metadata["page_number"], detected_object.content, detected_object.confidence)
Initializes the GLAIRVisionImageToOcr instance with GLAIR Vision OCR credentials.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
username
|
str
|
Basic-auth username for the GLAIR Vision OCR API. |
required |
password
|
str
|
Basic-auth password for the GLAIR Vision OCR API. |
required |
api_key
|
str
|
Value sent as the |
required |
endpoint
|
str
|
Full GLAIR Vision OCR endpoint URL. Defaults to DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT. |
DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT
|
include_detected_objects
|
bool
|
Whether to populate |
False
|
timeout
|
int | float
|
Request timeout in seconds, applied to each attempt. Defaults to 300. |
DEFAULT_TIMEOUT_SECONDS
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
ImageOcrPreset
Bases: StrEnum
Available presets for LM-based OCR transcription.
Attributes:
| Name | Type | Description |
|---|---|---|
DEFAULT |
Free-text transcription of images and PDF documents. |
LMBasedImageToOcr(lm_request_processor, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE)
Bases: BaseImageToOcr, UsesLM
LM-based implementation for transcribing text from images and documents via OCR.
This class leverages a language model (LM) pipeline to transcribe visible text from image and
PDF sources into free text, reusing the source loading, attachment construction, and result
wrapping defined by BaseImageToOcr.
Inherits
BaseImageToOcr: Base class defining the image/document-to-OCR interface. UsesLM: Mixin providing shared logic for components using language models.
Attributes:
| Name | Type | Description |
|---|---|---|
lm_request_processor |
LMRequestProcessor
|
Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the OCR pipeline. |
include_bounding_boxes |
bool
|
Whether the LM output is parsed as the bounding box JSON
object and returned as |
bounding_box_range |
tuple[int, int]
|
The |
Usage
Use from_preset to easily instantiate with predefined configuration:
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.lm_based_image_to_ocr import (
LMBasedImageToOcr,
)
converter = LMBasedImageToOcr.from_preset()
Initializes the LMBasedImageToOcr instance with a language model request processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lm_request_processor
|
LMRequestProcessor
|
The processor handling prompt creation, LM invocation, and output parsing. |
required |
include_bounding_boxes
|
bool
|
Whether the LM output is parsed as the bounding box JSON object. Defaults to False. |
False
|
bounding_box_range
|
tuple[int, int]
|
The |
DEFAULT_BOX_2D_RANGE
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
convert(source, **kwargs)
async
Run LM-based OCR transcription on a loaded image or document source.
This is a documented pass-through to BaseImageToOcr.convert: it does not reload media,
reconstruct attachments, or reimplement source validation and fallback handling. Those
behaviors, along with TextResult wrapping, are owned by the base class.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
The document or image source to transcribe. Supported forms: 1. Raw image or PDF bytes. 2. Local file path. 3. URL, S3 URI, or Google Drive URL. 4. Base64 encoded string. |
required |
**kwargs
|
Any
|
See Other Parameters below. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
multimodal_context |
list[Attachment | bytes | str]
|
Reference context sent after the document to help the LM read ambiguous characters, names, or terms. Context is never transcribed. Bytes are converted into Attachments, strings that resolve to an image or PDF source are loaded as Attachments, and other strings are sent as text. Defaults to []. |
event_emitter |
Any
|
Event emitter forwarded to LM request processing. |
hyperparameters |
dict[str, Any]
|
Hyperparameters forwarded to the LM invocation. |
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
OCR result with tag |
Raises:
| Type | Description |
|---|---|
TypeError
|
If source is not a str or bytes, or if |
ValueError
|
If source is an empty string. |
Example
converter = LMBasedImageToOcr.from_preset()
result = await converter.convert(source="manual.pdf")
print(result.result)
result = await converter.convert(
source="handwritten_form.png",
multimodal_context=["Product names: Glair, Catapa.", "reference_page.png"],
)
from_preset(preset_name=ImageOcrPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE, **kwargs)
classmethod
Initializes the LM-based OCR component using preset model configurations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
ImageOcrPreset | str | None
|
Preset name forwarded to the image-OCR
preset registry. Defaults to |
DEFAULT
|
lm_invoker_kwargs
|
dict[str, Any] | None
|
Keyword arguments passed to LM
invoker creation in the preset factory. Defaults to None. Valid keys follow
[ |
None
|
prompt_builder_kwargs
|
dict[str, Any] | None
|
Keyword arguments passed to
prompt builder creation in the preset factory. Defaults to None. Valid keys follow
|
None
|
include_bounding_boxes
|
bool
|
Whether the LM is additionally asked for the bounding box of every line of text. Defaults to False. |
False
|
bounding_box_range
|
tuple[int, int]
|
The |
DEFAULT_BOX_2D_RANGE
|
**kwargs
|
Any
|
Additional kwargs for current-class |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Any additional kwargs forwarded to |
Returns:
| Name | Type | Description |
|---|---|---|
LMBasedImageToOcr |
LMBasedImageToOcr
|
Initialized OCR component using the preset model. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr import (
LMBasedImageToOcr,
ImageOcrPreset,
)
# Use the default preset (Gemini flash-lite).
ocr = LMBasedImageToOcr.from_preset()
# Override the model via lm_invoker_kwargs.
ocr = LMBasedImageToOcr.from_preset(
preset_name=ImageOcrPreset.DEFAULT,
lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)
result = await ocr.convert(source="invoice.png")
print(result.result)
# Also return the pixel bounding box of every line of text.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True)
result = await ocr.convert(source="invoice.png")
for detected_object in result.metadata.detected_objects:
print(detected_object.content, detected_object.to_bounding_box())
# Ask for bounding box coordinates on a different integer range.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True, bounding_box_range=(0, 999))