Overview
Unified schema exports for gllm_multimodal.
This module re-exports all data models and result schemas used by modality converters and transformers. Importing from this module provides a convenient single entry point for all schema types.
Core multimodal types such as AudioTranscript, Keyframe, and
VideoSegment live in [gllm_core.schema.multimodal][]
and should be imported from there directly.
Exported schemas
| Schema | Description |
|---|---|
TextResult |
Base result for text extraction operations. |
CaptionResult |
Result of a captioning operation. |
Caption |
Input schema for captioning operations. |
OcrResult |
Structured result of an OCR operation. |
Mermaid |
Additional metadata for mermaid generation. |
Usage
from gllm_multimodal.schema import TextResult, CaptionResult, OcrResult
Caption
Bases: BaseModel
Result class for captioning operations (image, video, etc.).
This class provides a structured format for captioning results, supporting: - Multiple caption types (one-liner, detailed, domain-specific) - Caption count tracking - Metadata storage for processing details
Attributes:
| Name | Type | Description |
|---|---|---|
text_one_liner |
str
|
Brief, single-sentence summary of the content. Defaults to empty string if not provided. |
text_context |
str
|
Detailed, multi-sentence description of the content. Defaults to empty string if not provided. |
domain_knowledge |
str
|
Domain-specific interpretation or context. Defaults to empty string if not provided. |
number_of_captions |
int
|
Total number of distinct captions generated. Defaults to 0 if no captions are generated. |
media_metadata |
dict[str, Any]
|
Additional information about the media such as location. |
multimodal_context |
list[Attachment | str]
|
Optional list of external context
objects (files, bytes, or pre-processed inputs) or raw strings that can
enrich captioning results. Bytes are automatically converted into
Attachment objects via |
output_schema |
str
|
Output schema. Defaults to empty string if not provided. |
schema_description |
str
|
Schema description. Defaults to empty string if not provided. |
language |
str
|
Language of the captions. Defaults to "Indonesian" if not provided. |
handle_multimodal_context(multimodal_value)
classmethod
Normalize and validate multimodal_context.
This method ensures that the multimodal_context field is a list of
Attachment objects or strings. It handles multiple input cases:
- None -> returns an empty list
- list[bytes] -> converts each item into an Attachment via
Attachment.from_bytes - list[Attachment] -> keeps as-is
- list[str] -> keeps as-is if it's not a valid image/binary source, otherwise converts to Attachment.
- list[mixed] -> normalizes supported types
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
multimodal_value
|
Any
|
Input value provided to |
required |
Returns:
| Type | Description |
|---|---|
Any
|
list[Attachment | str]: A normalized list of |
handle_none_metadata(metadata_value)
classmethod
Handle None values for media_metadata by using empty dict.
handle_none_number_of_captions(caption_value)
classmethod
Handle None values for number_of_captions by using default.
handle_none_values(str_value)
classmethod
Handle None values by converting them to default values.
CaptionResult
Bases: Caption
Result of a captioning operation.
Attributes:
| Name | Type | Description |
|---|---|---|
captions |
str | list[str] | dict[str, Any]
|
The caption result. May be a single string, a list of captions, or a structured dictionary depending on the output format. |
Mermaid
Bases: BaseModel
Additional metadata for mermaid diagram generation.
Attributes:
| Name | Type | Description |
|---|---|---|
diagram_type |
str | None
|
Type of the diagram to be generated
(e.g. |
context |
str | None
|
Additional context or prompt used to generate
the mermaid diagram. Defaults to |
OcrResult
Bases: BaseModel
Structured result of an OCR operation.
Attributes:
| Name | Type | Description |
|---|---|---|
text |
str
|
Full concatenated text extracted from the document. This value is mirrored in TextResult.result for API consistency. |
lines |
list[str]
|
Engine-native line units when available (e.g. from specialized OCR backends). LM-based implementations typically leave this empty. Defaults to an empty list. |
confidence |
float | None
|
Overall confidence score of the extraction, expressed as a value between 0.0 and 1.0. Only populated by specialized OCR engines; LM-based implementations leave this as None. Defaults to None. |
page_count |
int
|
Number of pages processed. Populated by engines that support multi-page documents (e.g., Azure Document Intelligence). Defaults to 1. |
TextResult
Bases: BaseModel
Base class for all modality-to-text operation results.
This class provides the foundation for structured results from any modality conversion operation, including: - Image Captioning - OCR / Scene Text Detection - Audio Transcription - Video Captioning
Attributes:
| Name | Type | Description |
|---|---|---|
result |
str
|
The extracted or generated text from the source. This is the primary output of any modality conversion operation. May be empty if the operation fails or no text is found. |
tag |
str
|
A label identifying the type of conversion that produced
this result (e.g. |
metadata |
dict[str, Any] | BaseModel | None
|
Additional metadata from the conversion process. |