Skip to content

Overview

Unified schema exports for gllm_multimodal.

This module re-exports all data models and result schemas used by modality converters and transformers. Importing from this module provides a convenient single entry point for all schema types.

Core multimodal types such as AudioTranscript, Keyframe, and VideoSegment live in [gllm_core.schema.multimodal][] and should be imported from there directly.

Exported schemas

Schema Description
TextResult Base result for text extraction operations.
CaptionResult Result of a captioning operation.
Caption Input schema for captioning operations.
OcrResult Structured result of an OCR operation.
Mermaid Additional metadata for mermaid generation.

Usage

from gllm_multimodal.schema import TextResult, CaptionResult, OcrResult

Caption

Bases: BaseModel

Result class for captioning operations (image, video, etc.).

This class provides a structured format for captioning results, supporting: - Multiple caption types (one-liner, detailed, domain-specific) - Caption count tracking - Metadata storage for processing details

Attributes:

Name Type Description
text_one_liner str

Brief, single-sentence summary of the content. Defaults to empty string if not provided.

text_context str

Detailed, multi-sentence description of the content. Defaults to empty string if not provided.

domain_knowledge str

Domain-specific interpretation or context. Defaults to empty string if not provided.

number_of_captions int

Total number of distinct captions generated. Defaults to 0 if no captions are generated.

media_metadata dict[str, Any]

Additional information about the media such as location.

multimodal_context list[Attachment | str]

Optional list of external context objects (files, bytes, or pre-processed inputs) or raw strings that can enrich captioning results. Bytes are automatically converted into Attachment objects via Attachment.from_bytes.

output_schema str

Output schema. Defaults to empty string if not provided.

schema_description str

Schema description. Defaults to empty string if not provided.

language str

Language of the captions. Defaults to "Indonesian" if not provided.

handle_multimodal_context(multimodal_value) classmethod

Normalize and validate multimodal_context.

This method ensures that the multimodal_context field is a list of Attachment objects or strings. It handles multiple input cases:

  • None -> returns an empty list
  • list[bytes] -> converts each item into an Attachment via Attachment.from_bytes
  • list[Attachment] -> keeps as-is
  • list[str] -> keeps as-is if it's not a valid image/binary source, otherwise converts to Attachment.
  • list[mixed] -> normalizes supported types

Parameters:

Name Type Description Default
multimodal_value Any

Input value provided to multimodal_context.

required

Returns:

Type Description
Any

list[Attachment | str]: A normalized list of Attachment objects or strings.

handle_none_metadata(metadata_value) classmethod

Handle None values for media_metadata by using empty dict.

handle_none_number_of_captions(caption_value) classmethod

Handle None values for number_of_captions by using default.

handle_none_values(str_value) classmethod

Handle None values by converting them to default values.

CaptionResult

Bases: Caption

Result of a captioning operation.

Attributes:

Name Type Description
captions str | list[str] | dict[str, Any]

The caption result. May be a single string, a list of captions, or a structured dictionary depending on the output format.

Mermaid

Bases: BaseModel

Additional metadata for mermaid diagram generation.

Attributes:

Name Type Description
diagram_type str | None

Type of the diagram to be generated (e.g. "flowchart", "sequence"). Defaults to None.

context str | None

Additional context or prompt used to generate the mermaid diagram. Defaults to None.

OcrResult

Bases: BaseModel

Structured result of an OCR operation.

Attributes:

Name Type Description
text str

Full concatenated text extracted from the document. This value is mirrored in TextResult.result for API consistency.

lines list[str]

Engine-native line units when available (e.g. from specialized OCR backends). LM-based implementations typically leave this empty. Defaults to an empty list.

confidence float | None

Overall confidence score of the extraction, expressed as a value between 0.0 and 1.0. Only populated by specialized OCR engines; LM-based implementations leave this as None. Defaults to None.

page_count int

Number of pages processed. Populated by engines that support multi-page documents (e.g., Azure Document Intelligence). Defaults to 1.

TextResult

Bases: BaseModel

Base class for all modality-to-text operation results.

This class provides the foundation for structured results from any modality conversion operation, including: - Image Captioning - OCR / Scene Text Detection - Audio Transcription - Video Captioning

Attributes:

Name Type Description
result str

The extracted or generated text from the source. This is the primary output of any modality conversion operation. May be empty if the operation fails or no text is found.

tag str

A label identifying the type of conversion that produced this result (e.g. "captions", "ocr").

metadata dict[str, Any] | BaseModel | None

Additional metadata from the conversion process.