Overview
Image-to-text conversion module providing OCR and captioning capabilities.
Submodules
image_to_text-- Abstract base class for image-to-text converters.image_to_caption-- Image captioning implementations.image_to_mermaid-- Image-to-mermaid diagram generation.image_to_ocr-- Image/document OCR transcription.
GLAIRVisionImageToOcr(username, password, api_key, endpoint=DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT, include_detected_objects=False, timeout=DEFAULT_TIMEOUT_SECONDS)
Bases: BaseImageToOcr
Managed OCR implementation backed by the GLAIR Vision OCR API.
Sends one authenticated multipart request per conversion to a GLAIR Vision OCR endpoint and maps the
validated response into OcrResult: ordered text lines and, when include_detected_objects is True,
detected_objects holding one object per text line (class text or handwriting, with confidence and
one-based page number) followed by one object per table (class table, with the table rendered as HTML).
Tables are always requested from the API.
Usage
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.glair_vision_image_to_ocr import (
GLAIRVisionImageToOcr,
)
converter = GLAIRVisionImageToOcr(
username="username",
password="password",
api_key="api-key",
)
result = await converter.convert("invoice.png")
print(result.result)
for detected_object in result.metadata.detected_objects:
if detected_object.class_name == DetectedObjectClass.TABLE:
print(detected_object.metadata["page_number"], detected_object.content)
else:
print(detected_object.metadata["page_number"], detected_object.content, detected_object.confidence)
Initializes the GLAIRVisionImageToOcr instance with GLAIR Vision OCR credentials.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
username
|
str
|
Basic-auth username for the GLAIR Vision OCR API. |
required |
password
|
str
|
Basic-auth password for the GLAIR Vision OCR API. |
required |
api_key
|
str
|
Value sent as the |
required |
endpoint
|
str
|
Full GLAIR Vision OCR endpoint URL. Defaults to DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT. |
DEFAULT_GLAIR_GENERAL_DOCUMENT_ENDPOINT
|
include_detected_objects
|
bool
|
Whether to populate |
False
|
timeout
|
int | float
|
Request timeout in seconds, applied to each attempt. Defaults to 300. |
DEFAULT_TIMEOUT_SECONDS
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
LMBasedImageToCaption(lm_request_processor, formatter=DEFAULT_FORMATTER)
Bases: BaseImageToCaption, UsesLM
Image captioning implementation using Language Models.
This class implements the ImageToCaption interface using LMs for generating natural language captions.
Usage
Use from_preset to easily instantiate with predefined configuration:
from gllm_multimodal.modality_converter.image_to_text.image_to_caption.lm_based_image_to_caption import LMBasedImageToCaption
converter = LMBasedImageToCaption.from_preset()
Initialize the LM based image captioning component.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lm_request_processor
|
LMRequestProcessor
|
Language model request processor instance that supports multimodal inputs. |
required |
formatter
|
BaseCaptionOutputFormatter
|
The formatter to use for formatting the output. |
DEFAULT_FORMATTER
|
convert(source, **kwargs)
async
Convert image into caption text using the LM captioning pipeline.
Behavior
- Validates and loads media bytes from path, URL, base64, or bytes source.
- Optionally extracts image metadata and builds the
Captioncontext payload. - Builds prompt params from supported prompt keys and caption fields.
- Sends image plus multimodal context attachments to the LM request processor.
- Formats the generated caption using the configured output formatter.
- Returns formatted caption text and metadata as
TextResult.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
Image source to caption. Supported forms: 1. Raw image bytes. 2. Local file path. 3. URL string. 4. Base64 encoded image string. |
required |
**kwargs
|
Any
|
See Other Parameters below. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
number_of_captions |
int
|
Number of captions to generate. |
text_one_liner |
str
|
Short one-line context/title. |
text_context |
str
|
Additional descriptive context. |
domain_knowledge |
str
|
Domain hints for caption quality. |
multimodal_context |
list[Attachment | str]
|
Additional context attachments. |
use_metadata |
bool
|
Whether image metadata is extracted and injected. |
formatter_kwargs |
dict[str, Any]
|
Extra kwargs for formatter output shaping. |
event_emitter |
Any
|
Event emitter passed to LM request processing. |
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
Formatted caption result with structured metadata. |
Example
from gllm_inference.schema import Attachment
converter = LMBasedImageToCaption.from_preset()
result = await converter.convert(
source="diagram.png",
number_of_captions=3,
text_one_liner="System architecture diagram",
domain_knowledge="cloud networking",
multimodal_context=[
"focus on request flow",
Attachment.from_file("legend.png"),
],
formatter_kwargs={"include_bullets": True, "max_sentences": 6},
use_metadata=True,
)
print(result.result)
from_preset(preset_name=ImageCaptionPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs)
classmethod
Initialize the LM based image captioning component using preset model configurations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
ImageCaptionPreset | str | None
|
Preset name forwarded to
image-caption preset registry. Defaults to |
DEFAULT
|
lm_invoker_kwargs
|
dict | None
|
Keyword arguments passed to LM invoker
creation in the preset factory. Defaults to None. Valid keys follow
[ |
None
|
prompt_builder_kwargs
|
dict | None
|
Keyword arguments passed to prompt
builder creation in the preset factory. Defaults to None. Valid keys follow
|
None
|
**kwargs
|
Any
|
Additional kwargs for current-class |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Any additional kwargs forwarded to |
Returns:
| Name | Type | Description |
|---|---|---|
LMBasedImageToCaption |
LMBasedImageToCaption
|
Initialized image captioning component using preset model. |
Example
from gllm_multimodal.modality_converter.image_to_text.image_to_caption import (
LMBasedImageToCaption,
ImageCaptionPreset,
)
# Use default preset (Gemini flash-lite)
captioner = LMBasedImageToCaption.from_preset()
# Use structured output preset with a custom model
captioner = LMBasedImageToCaption.from_preset(
preset_name=ImageCaptionPreset.STRUCTURED,
lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)
result = await captioner.convert(source="diagram.png")
print(result.result)
LMBasedImageToMermaid(lm_request_processor)
Bases: BaseImageToMermaid, UsesLM
LM-based implementation for converting an image into Mermaid diagram syntax.
This class leverages a language model (LM) pipeline to generate structured Mermaid syntax from image inputs and optional metadata. It uses prompt builders, LM invokers, and output parsers defined via a preset system to streamline model usage.
Inherits
BaseImageToMermaid: Base class defining the image-to-mermaid interface. UsesLM: Mixin providing shared logic for components using language models.
Attributes:
| Name | Type | Description |
|---|---|---|
lm_request_processor |
LMRequestProcessor
|
Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the image-to-mermaid pipeline. |
Initializes the LMBasedImageToMermaid instance with a language model request processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lm_request_processor
|
LMRequestProcessor
|
The processor handling prompt creation, LM invocation, and output parsing. |
required |
convert(source, **kwargs)
async
Convert image input into Mermaid syntax using the LM conversion flow.
Behavior
- Validates and loads image bytes from supported source formats.
- Builds Mermaid metadata payload from method keyword arguments.
- Filters prompt params to keys recognized by the prompt template.
- Sends the image as attachment to the LM request processor.
- Returns Mermaid syntax text and metadata wrapped in
TextResult.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
Image source to convert. Supported forms: 1. Raw image bytes. 2. Local file path. 3. URL string. 4. Base64 encoded image string. |
required |
**kwargs
|
Any
|
Mermaid generation options forwarded into metadata
and LM calls, including:
1. title (str, optional): Diagram title context.
2. description (str, optional): Additional scene or flow context.
3. event_emitter (Any, optional): Event emitter passed to LM request processing.
4. Any Mermaid schema fields recognized by
|
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
Mermaid syntax output with associated metadata. |
Example
converter = LMBasedImageToMermaid.from_preset()
result = await converter.convert(
source="workflow.png",
title="Ticket Escalation Flow",
description="Extract decision points and handoff states",
diagram_type="flowchart",
direction="LR",
event_emitter=my_event_emitter,
)
print(result.result)
from_preset(preset_name='default', lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs)
classmethod
Constructs an LMBasedImageToMermaid instance using a named preset configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
str | None
|
Name of the predefined preset configuration to use. Defaults to "default". |
'default'
|
lm_invoker_kwargs
|
dict | None
|
Keyword arguments passed to LM invoker
creation in the preset factory. Defaults to None. Valid keys follow
[ |
None
|
prompt_builder_kwargs
|
dict | None
|
Keyword arguments passed to prompt
builder creation in the preset factory. Defaults to None. Valid keys follow
|
None
|
**kwargs
|
Any
|
Additional kwargs for current-class |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Any additional kwargs forwarded to |
Returns:
| Name | Type | Description |
|---|---|---|
LMBasedImageToMermaid |
LMBasedImageToMermaid
|
An instance initialized with the preset's components. |
Example
from gllm_multimodal.modality_converter.image_to_text.image_to_mermaid import (
LMBasedImageToMermaid,
)
# Use default preset (Gemini flash-lite)
converter = LMBasedImageToMermaid.from_preset()
# Override model via lm_invoker_kwargs
converter = LMBasedImageToMermaid.from_preset(
lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)
result = await converter.convert(source="architecture.png")
print(result.result)
LMBasedImageToOcr(lm_request_processor, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE)
Bases: BaseImageToOcr, UsesLM
LM-based implementation for transcribing text from images and documents via OCR.
This class leverages a language model (LM) pipeline to transcribe visible text from image and
PDF sources into free text, reusing the source loading, attachment construction, and result
wrapping defined by BaseImageToOcr.
Inherits
BaseImageToOcr: Base class defining the image/document-to-OCR interface. UsesLM: Mixin providing shared logic for components using language models.
Attributes:
| Name | Type | Description |
|---|---|---|
lm_request_processor |
LMRequestProcessor
|
Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the OCR pipeline. |
include_bounding_boxes |
bool
|
Whether the LM output is parsed as the bounding box JSON
object and returned as |
bounding_box_range |
tuple[int, int]
|
The |
Usage
Use from_preset to easily instantiate with predefined configuration:
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr.lm_based_image_to_ocr import (
LMBasedImageToOcr,
)
converter = LMBasedImageToOcr.from_preset()
Initializes the LMBasedImageToOcr instance with a language model request processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lm_request_processor
|
LMRequestProcessor
|
The processor handling prompt creation, LM invocation, and output parsing. |
required |
include_bounding_boxes
|
bool
|
Whether the LM output is parsed as the bounding box JSON object. Defaults to False. |
False
|
bounding_box_range
|
tuple[int, int]
|
The |
DEFAULT_BOX_2D_RANGE
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
convert(source, **kwargs)
async
Run LM-based OCR transcription on a loaded image or document source.
This is a documented pass-through to BaseImageToOcr.convert: it does not reload media,
reconstruct attachments, or reimplement source validation and fallback handling. Those
behaviors, along with TextResult wrapping, are owned by the base class.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
The document or image source to transcribe. Supported forms: 1. Raw image or PDF bytes. 2. Local file path. 3. URL, S3 URI, or Google Drive URL. 4. Base64 encoded string. |
required |
**kwargs
|
Any
|
See Other Parameters below. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
multimodal_context |
list[Attachment | bytes | str]
|
Reference context sent after the document to help the LM read ambiguous characters, names, or terms. Context is never transcribed. Bytes are converted into Attachments, strings that resolve to an image or PDF source are loaded as Attachments, and other strings are sent as text. Defaults to []. |
event_emitter |
Any
|
Event emitter forwarded to LM request processing. |
hyperparameters |
dict[str, Any]
|
Hyperparameters forwarded to the LM invocation. |
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
OCR result with tag |
Raises:
| Type | Description |
|---|---|
TypeError
|
If source is not a str or bytes, or if |
ValueError
|
If source is an empty string. |
Example
converter = LMBasedImageToOcr.from_preset()
result = await converter.convert(source="manual.pdf")
print(result.result)
result = await converter.convert(
source="handwritten_form.png",
multimodal_context=["Product names: Glair, Catapa.", "reference_page.png"],
)
from_preset(preset_name=ImageOcrPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, include_bounding_boxes=False, bounding_box_range=DEFAULT_BOX_2D_RANGE, **kwargs)
classmethod
Initializes the LM-based OCR component using preset model configurations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
ImageOcrPreset | str | None
|
Preset name forwarded to the image-OCR
preset registry. Defaults to |
DEFAULT
|
lm_invoker_kwargs
|
dict[str, Any] | None
|
Keyword arguments passed to LM
invoker creation in the preset factory. Defaults to None. Valid keys follow
[ |
None
|
prompt_builder_kwargs
|
dict[str, Any] | None
|
Keyword arguments passed to
prompt builder creation in the preset factory. Defaults to None. Valid keys follow
|
None
|
include_bounding_boxes
|
bool
|
Whether the LM is additionally asked for the bounding box of every line of text. Defaults to False. |
False
|
bounding_box_range
|
tuple[int, int]
|
The |
DEFAULT_BOX_2D_RANGE
|
**kwargs
|
Any
|
Additional kwargs for current-class |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Any additional kwargs forwarded to |
Returns:
| Name | Type | Description |
|---|---|---|
LMBasedImageToOcr |
LMBasedImageToOcr
|
Initialized OCR component using the preset model. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
from gllm_multimodal.modality_converter.image_to_text.image_to_ocr import (
LMBasedImageToOcr,
ImageOcrPreset,
)
# Use the default preset (Gemini flash-lite).
ocr = LMBasedImageToOcr.from_preset()
# Override the model via lm_invoker_kwargs.
ocr = LMBasedImageToOcr.from_preset(
preset_name=ImageOcrPreset.DEFAULT,
lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)
result = await ocr.convert(source="invoice.png")
print(result.result)
# Also return the pixel bounding box of every line of text.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True)
result = await ocr.convert(source="invoice.png")
for detected_object in result.metadata.detected_objects:
print(detected_object.content, detected_object.to_bounding_box())
# Ask for bounding box coordinates on a different integer range.
ocr = LMBasedImageToOcr.from_preset(include_bounding_boxes=True, bounding_box_range=(0, 999))