Skip to content

Overview

Video-to-text conversion module providing video captioning and transcription.

Submodules

HybridVideoToCaption(video_captioner=None, image_captioner=None, segmenter=None, transcriber=None, keyframe_extractor=None, pipeline_mode=PipelineMode.DIRECT_LM_CAPTION, audio_extractor_config=None, ocr_converter=None, segment_caption_strategy=SegmentCaptionStrategy.AUTO, video_summary_strategy=VideoSummaryStrategy.AUTO, caption_summarizer=None, **kwargs)

Bases: BaseVideoToCaption

Hybrid video captioning component with a pluggable flow registry.

Wraps an inner captioner (a BaseVideoToCaption implementation) and orchestrates optional segmenter, transcriber, and keyframe_extractor components into one of several named pipeline flows.

The active flow is selected by pipeline_mode, which maps to a registered handler function. Built-in handlers are pre-registered at module load time. New flows can be added without modifying this class via register_flow.

Attributes:

Name Type Description
video_captioner BaseVideoToCaption | None

Optional video captioner.

image_captioner BaseImageToCaption | None

Optional image captioner.

segmenter BaseSegmenter | None

Optional segmenter processor.

transcriber BaseAudioToText | None

Optional audio-to-text transcriber.

keyframe_extractor BaseKeyframeExtractor | None

Optional keyframe extractor processor.

ocr_converter BaseImageToOcr | None

Optional OCR converter run on each materialized keyframe. Ignored (with a warning) when no keyframe_extractor is configured.

caption_summarizer BaseLMInvoker | None

Optional text-only LM used to collapse keyframe captions into one segment caption under SegmentCaptionStrategy.SUMMARIZE_KEYFRAME_CAPTIONS, and segment captions into one video summary under VideoSummaryStrategy.SUMMARIZE_SEGMENT_CAPTIONS. Lazily built from a default model when not provided and either strategy is active.

pipeline_mode str

Active pipeline mode key.

segment_caption_strategy SegmentCaptionStrategy

How segment_caption is built when pipeline_mode is E2E_COMPREHENSIVE.

video_summary_strategy VideoSummaryStrategy

How video_summary is built when pipeline_mode is E2E_COMPREHENSIVE.

Example
Built-in DIRECT_LM_CAPTION mode
hybrid = HybridVideoToCaption(
    video_captioner=my_video_captioner,
    pipeline_mode=PipelineMode.DIRECT_LM_CAPTION,
)
result = await hybrid.convert("video.mp4", title="My Video")
E2E comprehensive (segment + transcript + keyframe + caption)
hybrid = HybridVideoToCaption(
    video_captioner=my_video_captioner,
    segmenter=my_segmenter,
    transcriber=my_transcriber,
    keyframe_extractor=my_keyframe_extractor,
    pipeline_mode=PipelineMode.E2E_COMPREHENSIVE,
)
result = await hybrid.convert("canal_tutorial.mp4")
# result.video_summary                 → whole-video summary
# result.segments[i].transcripts       → speech in that time window
# result.segments[i].segment_caption   → clip-level caption(s)
# result.segments[i].keyframes[j].caption → per-keyframe caption

See flow_e2e_comprehensive for a step-by-step walkthrough and a concrete VideoCaptionMetadata output example.

Register and use a custom flow
async def my_flow(ctx: FlowContext, attachment, caption_data, **kwargs):
    ...
    return await ctx.video_captioner.convert(attachment, **kwargs)

HybridVideoToCaption.register_flow(
    mode="my_flow",
    handler=my_flow,
    required_components=["segmenter"],
)

hybrid = HybridVideoToCaption(
    video_captioner=my_video_captioner,
    segmenter=my_segmenter,
    pipeline_mode="my_flow",
)
result = await hybrid.convert("video.mp4")

Initialise the hybrid captioner.

Parameters:

Name Type Description Default
video_captioner BaseVideoToCaption | None

The video captioner used to generate captions directly from video (or segment) attachments. Defaults to None.

None
image_captioner BaseImageToCaption | None

The image captioner used when captioning individual frames or images. Defaults to None.

None
segmenter BaseSegmenter | None

Segmenter processor. Defaults to None.

None
transcriber BaseAudioToText | None

Audio-to-text transcriber. Defaults to None.

None
keyframe_extractor BaseKeyframeExtractor | None

Frame extraction processor. Defaults to None.

None
pipeline_mode PipelineMode | str

The pipeline mode to use. Defaults to PipelineMode.DIRECT_LM_CAPTION.

DIRECT_LM_CAPTION
audio_extractor_config dict[str, Any] | None

Config forwarded to AudioExtractionProcessor.build (e.g. {"output_format": "mp3"} or {"audio_encoder": "lamemp3enc"}). Defaults to None (auto-select format, typically WAV).

None
ocr_converter BaseImageToOcr | None

OCR converter run on each keyframe materialized by keyframe_extractor. Its output is stored on Keyframe.detector_result["ocr"] and never fed to a captioner. Ignored (with a warning) when keyframe_extractor is not also provided. Defaults to None.

None
segment_caption_strategy SegmentCaptionStrategy | str

How segment_caption is built when pipeline_mode is E2E_COMPREHENSIVE. Ignored (with a warning) for other modes. Defaults to SegmentCaptionStrategy.AUTO.

AUTO
video_summary_strategy VideoSummaryStrategy | str

How video_summary is built when pipeline_mode is E2E_COMPREHENSIVE. Ignored (with a warning) for other modes. Defaults to VideoSummaryStrategy.AUTO.

AUTO
caption_summarizer BaseLMInvoker | None

Text-only LM used to collapse keyframe captions into one segment caption under SegmentCaptionStrategy.SUMMARIZE_KEYFRAME_CAPTIONS, and segment captions into one video summary under VideoSummaryStrategy.SUMMARIZE_SEGMENT_CAPTIONS. When not provided, a default LM is lazily built the first time either strategy is used. Defaults to None.

None
**kwargs Any

Forwarded to BaseVideoToCaption.

{}

Raises:

Type Description
ValueError

If pipeline_mode is not registered in the flow registry.

ValueError

If a component required by pipeline_mode is None.

ValueError

If segment_caption_strategy or video_summary_strategy is not a valid strategy value, or requires a component that was not provided.

Example
Explicit segment caption and video summary strategies
hybrid = HybridVideoToCaption(
    video_captioner=my_video_captioner,
    image_captioner=my_image_captioner,
    segmenter=my_segmenter,
    keyframe_extractor=my_keyframe_extractor,
    pipeline_mode=PipelineMode.E2E_COMPREHENSIVE,
    segment_caption_strategy=SegmentCaptionStrategy.CONCAT_KEYFRAME_CAPTIONS,
    video_summary_strategy=VideoSummaryStrategy.CONCAT_SEGMENT_CAPTIONS,
)
result = await hybrid.convert("lecture.mp4")

__init_subclass__(**kwargs)

Give every subclass its own independent registry copy.

convert(source, **kwargs) async

Convert video into caption output using the selected hybrid flow mode.

Behavior
  1. Validates and resolves source into a video attachment.
  2. Builds caption context fields from kwargs.
  3. Calculates video duration when available for downstream validation.
  4. Dispatches caption generation to the registered handler for pipeline_mode.
  5. Uses flow context components such as captioner, transcriber, segmenter, and keyframe tools.
  6. Returns video summary plus structured segments in TextResult.

Parameters:

Name Type Description Default
source str | bytes

Video source to caption. Supported forms: 1. Raw video bytes. 2. Local file path. 3. URL string. 4. Base64 encoded video string.

required
**kwargs Any

Runtime options forwarded to the active flow handler.

{}

Other Parameters:

Name Type Description
number_of_captions int

Requested number of generated captions.

title str

Short title context.

description str

Additional descriptive context.

domain_knowledge str

Domain hints for caption quality.

multimodal_context list[Attachment | str]

Supplemental context.

event_emitter Any

Event emitter passed into downstream components.

segment_caption_strategy SegmentCaptionStrategy | str

Overrides the constructor's segment_caption_strategy for this call only. Only applies to E2E_COMPREHENSIVE.

video_summary_strategy VideoSummaryStrategy | str

Overrides the constructor's video_summary_strategy for this call only. Only applies to E2E_COMPREHENSIVE.

prev_caption_window_size int

Number of preceding segment captions passed as context when captioning the next segment or its keyframes. Must be a non-negative integer. Defaults to DEFAULT_PREV_CAPTION_WINDOW_SIZE.

prev_keyframe_window_size int

Number of preceding keyframe images, and of their captions, passed as context when captioning the next keyframe. Lower it for models with a small image limit. Must be a non-negative integer. Defaults to DEFAULT_PREV_KEYFRAME_WINDOW_SIZE.

Returns:

Name Type Description
TextResult TextResult

Hybrid flow caption output with structured metadata.

Example
converter = HybridVideoToCaption.from_preset("e2e_lm_only")
result = await converter.convert(
    "video.mp4",
    title="Inspection footage",
    number_of_captions=6,
    domain_knowledge="industrial safety audit",
    prev_keyframe_window_size=3,
    multimodal_context=["Prioritize hazards and operator actions."],
    event_emitter=my_event_emitter,
)
print(result.result)

from_preset(preset_name=HybridPreset.E2E_LM_ONLY, **kwargs) classmethod

Create a HybridVideoToCaption using a preset configuration.

Delegates preset_name to the preset registry in preset.py to retrieve the appropriate component composition and pipeline mode, then wraps it in a HybridVideoToCaption. Extra arguments can be injected via kwargs.

Parameters:

Name Type Description Default
preset_name HybridPreset | str | None

Preset name forwarded to the registry. Defaults to HybridPreset.E2E_LM_ONLY.

E2E_LM_ONLY
**kwargs Any

See Other Parameters below.

{}

Other Parameters:

Name Type Description
video_captioner_kwargs dict | None

Overrides forwarded to the video captioner constructor (e.g. lm_invoker_kwargs, prompt_builder_kwargs). Popped before __init__. Defaults to None. Valid args reference: LMBasedVideoToCaption.from_preset. For nested lm_invoker_kwargs specifically, valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker]. For nested prompt_builder_kwargs specifically, valid keys follow PromptBuilder.

image_captioner_kwargs dict | None

Overrides forwarded to the keyframe image captioner constructor. Popped before __init__. Defaults to None. Valid args reference: LMBasedImageToCaption.from_preset.

segmenter_kwargs dict | None

Overrides forwarded to the preset segmenter constructor (detector=, config= for ShotBasedSegmenter on e2e_shot_middle_frame). Popped before __init__. Defaults to None. To inject FixedDurationSegmenter, pass segmenter=FixedDurationSegmenter(config={...}) via **kwargs instead. See build_media_toolkit and BaseSegmenter.

transcriber_kwargs dict | None

Overrides forwarded to the audio transcriber builder. Popped before __init__. Defaults to None. Pass model_id here to select transcript approach/provider; remaining keys are forwarded as kwargs to build_modality_converter. Where to see valid kwargs: 1. Routing logic and forwarding behavior: get_preset_hybrid_video_to_caption. 2. Generic builder contract and strategy keys (preset, lmrp_config, direct kwargs): build_modality_converter. 3. LM transcript preset kwargs for LM-based route: LMBasedAudioToTranscript.from_preset.

keyframe_extractor_kwargs dict | None

Overrides forwarded to the keyframe extractor constructor. Popped before __init__. Defaults to None. Valid args should follow media-toolkit factory/class constructor signatures, see build_media_toolkit and BaseKeyframeExtractor.

ocr_converter_kwargs dict | None

Overrides forwarded to the OCR converter constructor. Popped before __init__. Defaults to None. Only consumed by e2e_ocr_driven; other presets ignore it with a warning. Valid args reference: LMBasedImageToOcr.from_preset.

caption_summarizer_kwargs dict | None

Overrides forwarded to the caption_summarizer LM invoker builder. Popped before __init__. Injected for every preset regardless of segment_caption_strategy, since any E2E preset can be switched to SUMMARIZE_KEYFRAME_CAPTIONS via segment_caption_strategy=. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker]: model_id, credentials, config.

pipeline_mode PipelineMode | str

Override the active pipeline mode set by the preset. Merged into __init__ after preset resolution.

**kwargs Any

Any other __init__ parameter (e.g. video_captioner, transcriber, segmenter) to replace a preset component entirely.

Returns:

Name Type Description
HybridVideoToCaption 'HybridVideoToCaption'

A fully initialised hybrid captioner.

Example
from gllm_multimodal.modality_converter.video_to_text.video_to_caption import (
    HybridVideoToCaption,
    HybridPreset,
)

captioner = HybridVideoToCaption.from_preset(
    preset_name=HybridPreset.E2E_COMPREHENSIVE,
    video_captioner_kwargs={
        "lm_invoker_kwargs": {"model_id": "google/gemini-3.1-flash-lite"},
    },
    transcriber_kwargs={
        "api_key": "sk-...",
        "model": "openai/whisper-large-v3",
    },
    segmenter_kwargs={"config": {"threshold": 30.0, "min_shot_duration": 1.5}},
)
result = await captioner.convert(source="video.mp4")
print(result.result)
Example
Text-detection keyframes with OCR (e2e_ocr_driven)
captioner = HybridVideoToCaption.from_preset(
    preset_name=HybridPreset.E2E_OCR_DRIVEN,
)
result = await captioner.convert(source="slides.mp4")
# Each keyframe carries both a visual caption and the raw on-screen text:
# result.metadata.segments[i].keyframes[j].caption
# result.metadata.segments[i].keyframes[j].detector_result["ocr"]["content"]

register_flow(mode, handler, required_components=None) classmethod

Register a new pipeline flow.

After registration the mode string can be passed as pipeline_mode to any new HybridVideoToCaption instance.

Parameters:

Name Type Description Default
mode str

Unique key for this flow (e.g. "my_custom_flow"). Can be a PipelineMode value or any arbitrary string.

required
handler FlowHandler

An async callable with the signature:

async def handler(
    ctx: FlowContext,
    video_attachment: Attachment,
    caption_data: Caption,
    **kwargs: Any,
) -> VideoCaptionMetadata | None: ...
required
required_components list[str | tuple[str, ...]] | None

Names of HybridVideoToCaption attributes (e.g. ["segmenter", "transcriber"]) that must not be None when this mode is selected. Validation happens at __init__ time. Use a tuple (e.g. [("captioner", "transcriber")]) to specify that at least one of the components in the tuple is required. Defaults to [].

None
Example
Register a custom flow
async def my_flow(
    ctx: FlowContext, attachment, caption_data, **kwargs
) -> VideoCaptionMetadata | None:
    result = await ctx.video_captioner.convert(attachment, **kwargs)
    return VideoCaptionMetadata(
        video_summary=result.result,
        segments=[],
    )

HybridVideoToCaption.register_flow(
    mode="my_flow",
    handler=my_flow,
    required_components=["segmenter"],
)

LMBasedVideoToCaption(lm_request_processor, transform=None, max_retries=2, **kwargs)

Bases: BaseVideoToCaption, UsesLM

Video captioning implementation using Language Models.

This class implements the VideoToCaption interface using LMs for generating natural language captions.

Initialize the LM based video captioning component.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

Language model request processor instance that supports multimodal inputs.

required
transform list[MediaToolkit] | None

Processor(s) to transform the video attachment.

None
max_retries int

Maximum number of retries for LM invocations. Defaults to 2.

2
**kwargs Any

Additional keyword arguments to pass to the parent constructor.

{}

convert(source, **kwargs) async

Convert video into caption output using the LM-based caption implementation.

Behavior
  1. Validates source input and resolves it to a video attachment.
  2. Builds caption context from incoming keyword arguments.
  3. Passes video plus multimodal context into LM request processing.
  4. Validates LM response schema and segment timing constraints.
  5. Retries LM generation up to max_retries on recoverable validation issues.
  6. Returns structured caption metadata and summary in TextResult.

Parameters:

Name Type Description Default
source str | bytes

Video source to caption. Supported forms: 1. Raw video bytes. 2. Local file path. 3. URL string. 4. Base64 encoded video string. 5. Attachment-compatible source supported by media loader.

required
**kwargs Any

Captioning options forwarded to runtime components.

{}

Other Parameters:

Name Type Description
number_of_captions int

Requested number of caption segments.

title str

One-line title or summary hint.

description str

Additional narrative context.

domain_knowledge str

Domain-specific guidance for captioning.

multimodal_context list[Attachment | str]

Extra context attachments.

delete_multimodal_context bool

Whether multimodal context is removed after processing.

event_emitter Any

Event emitter passed to LM request processing.

Returns:

Name Type Description
TextResult TextResult

Video summary text with segment-level caption metadata.

Example
from gllm_inference.schema import Attachment

converter = LMBasedVideoToCaption.from_preset()
result = await converter.convert(
    "video.mp4",
    title="Product demo",
    number_of_captions=5,
    multimodal_context=[
        "Focus on user actions and UI feedback.",
        Attachment.from_file("release_notes.png"),
    ],
    delete_multimodal_context=False,
    event_emitter=my_event_emitter,
)
print(result.result)

from_preset(preset_name='default', lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs) classmethod

Initialize the LM based video captioning component using preset model configurations.

Parameters:

Name Type Description Default
preset_name str | None

Preset name forwarded to video-caption preset registry. Defaults to "default".

'default'
lm_invoker_kwargs dict | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
**kwargs Any

Additional kwargs for current-class __init__ parameters. Preset defaults are merged first, then from_lm_components forwards kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs merged into preset kwargs and forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_video_to_caption. 2. Final converter constructor parameters: LMBasedVideoToCaption.

Returns:

Name Type Description
LMBasedVideoToCaption LMBasedVideoToCaption

Initialized video captioning component using preset model.

Example
from gllm_multimodal.modality_converter.video_to_text.video_to_caption import (
    LMBasedVideoToCaption,
)

# Use default preset (Gemini flash-lite)
captioner = LMBasedVideoToCaption.from_preset()

# Override model via lm_invoker_kwargs
captioner = LMBasedVideoToCaption.from_preset(
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await captioner.convert(source="video.mp4")
print(result.result)