Overview
Video-to-caption modules for generating descriptive captions from videos.
This package provides concrete implementations for captioning video content using language models.
Submodules
hybrid_video_to_caption-- Hybrid video captioning with keyframe-based pipeline.
Exported Classes
LMBasedVideoToCaption-- LM-based video captioning.
HybridVideoToCaption(video_captioner=None, image_captioner=None, segmenter=None, transcriber=None, keyframe_extractor=None, pipeline_mode=PipelineMode.DIRECT_LM_CAPTION, audio_extractor_config=None, ocr_converter=None, segment_caption_strategy=SegmentCaptionStrategy.AUTO, video_summary_strategy=VideoSummaryStrategy.AUTO, caption_summarizer=None, **kwargs)
Bases: BaseVideoToCaption
Hybrid video captioning component with a pluggable flow registry.
Wraps an inner captioner (a BaseVideoToCaption implementation)
and orchestrates optional segmenter, transcriber, and
keyframe_extractor components into one of several named pipeline flows.
The active flow is selected by pipeline_mode, which maps to a registered
handler function. Built-in handlers are pre-registered at module load time.
New flows can be added without modifying this class via register_flow.
Attributes:
| Name | Type | Description |
|---|---|---|
video_captioner |
BaseVideoToCaption | None
|
Optional video captioner. |
image_captioner |
BaseImageToCaption | None
|
Optional image captioner. |
segmenter |
BaseSegmenter | None
|
Optional segmenter processor. |
transcriber |
BaseAudioToText | None
|
Optional audio-to-text transcriber. |
keyframe_extractor |
BaseKeyframeExtractor | None
|
Optional keyframe extractor processor. |
ocr_converter |
BaseImageToOcr | None
|
Optional OCR converter run on each materialized
keyframe. Ignored (with a warning) when no |
caption_summarizer |
BaseLMInvoker | None
|
Optional text-only LM used to collapse keyframe
captions into one segment caption under |
pipeline_mode |
str
|
Active pipeline mode key. |
segment_caption_strategy |
SegmentCaptionStrategy
|
How |
video_summary_strategy |
VideoSummaryStrategy
|
How
|
Example
Built-in DIRECT_LM_CAPTION mode
hybrid = HybridVideoToCaption(
video_captioner=my_video_captioner,
pipeline_mode=PipelineMode.DIRECT_LM_CAPTION,
)
result = await hybrid.convert("video.mp4", title="My Video")
E2E comprehensive (segment + transcript + keyframe + caption)
hybrid = HybridVideoToCaption(
video_captioner=my_video_captioner,
segmenter=my_segmenter,
transcriber=my_transcriber,
keyframe_extractor=my_keyframe_extractor,
pipeline_mode=PipelineMode.E2E_COMPREHENSIVE,
)
result = await hybrid.convert("canal_tutorial.mp4")
# result.video_summary → whole-video summary
# result.segments[i].transcripts → speech in that time window
# result.segments[i].segment_caption → clip-level caption(s)
# result.segments[i].keyframes[j].caption → per-keyframe caption
See flow_e2e_comprehensive
for a step-by-step walkthrough and a concrete VideoCaptionMetadata output example.
Register and use a custom flow
async def my_flow(ctx: FlowContext, attachment, caption_data, **kwargs):
...
return await ctx.video_captioner.convert(attachment, **kwargs)
HybridVideoToCaption.register_flow(
mode="my_flow",
handler=my_flow,
required_components=["segmenter"],
)
hybrid = HybridVideoToCaption(
video_captioner=my_video_captioner,
segmenter=my_segmenter,
pipeline_mode="my_flow",
)
result = await hybrid.convert("video.mp4")
Initialise the hybrid captioner.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
video_captioner
|
BaseVideoToCaption | None
|
The video captioner used to generate captions directly from video (or segment) attachments. Defaults to None. |
None
|
image_captioner
|
BaseImageToCaption | None
|
The image captioner used when captioning individual frames or images. Defaults to None. |
None
|
segmenter
|
BaseSegmenter | None
|
Segmenter processor. Defaults to None. |
None
|
transcriber
|
BaseAudioToText | None
|
Audio-to-text transcriber. Defaults to None. |
None
|
keyframe_extractor
|
BaseKeyframeExtractor | None
|
Frame extraction processor. Defaults to None. |
None
|
pipeline_mode
|
PipelineMode | str
|
The pipeline mode to use. Defaults to PipelineMode.DIRECT_LM_CAPTION. |
DIRECT_LM_CAPTION
|
audio_extractor_config
|
dict[str, Any] | None
|
Config forwarded to
|
None
|
ocr_converter
|
BaseImageToOcr | None
|
OCR converter run on each keyframe
materialized by |
None
|
segment_caption_strategy
|
SegmentCaptionStrategy | str
|
How
|
AUTO
|
video_summary_strategy
|
VideoSummaryStrategy | str
|
How
|
AUTO
|
caption_summarizer
|
BaseLMInvoker | None
|
Text-only LM used to collapse
keyframe captions into one segment caption under
|
None
|
**kwargs
|
Any
|
Forwarded to |
{}
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
ValueError
|
If a component required by |
ValueError
|
If |
Example
Explicit segment caption and video summary strategies
hybrid = HybridVideoToCaption(
video_captioner=my_video_captioner,
image_captioner=my_image_captioner,
segmenter=my_segmenter,
keyframe_extractor=my_keyframe_extractor,
pipeline_mode=PipelineMode.E2E_COMPREHENSIVE,
segment_caption_strategy=SegmentCaptionStrategy.CONCAT_KEYFRAME_CAPTIONS,
video_summary_strategy=VideoSummaryStrategy.CONCAT_SEGMENT_CAPTIONS,
)
result = await hybrid.convert("lecture.mp4")
__init_subclass__(**kwargs)
Give every subclass its own independent registry copy.
convert(source, **kwargs)
async
Convert video into caption output using the selected hybrid flow mode.
Behavior
- Validates and resolves source into a video attachment.
- Builds caption context fields from kwargs.
- Calculates video duration when available for downstream validation.
- Dispatches caption generation to the registered handler for
pipeline_mode. - Uses flow context components such as captioner, transcriber, segmenter, and keyframe tools.
- Returns video summary plus structured segments in
TextResult.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
Video source to caption. Supported forms: 1. Raw video bytes. 2. Local file path. 3. URL string. 4. Base64 encoded video string. |
required |
**kwargs
|
Any
|
Runtime options forwarded to the active flow handler. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
number_of_captions |
int
|
Requested number of generated captions. |
title |
str
|
Short title context. |
description |
str
|
Additional descriptive context. |
domain_knowledge |
str
|
Domain hints for caption quality. |
multimodal_context |
list[Attachment | str]
|
Supplemental context. |
event_emitter |
Any
|
Event emitter passed into downstream components. |
segment_caption_strategy |
SegmentCaptionStrategy | str
|
Overrides the
constructor's |
video_summary_strategy |
VideoSummaryStrategy | str
|
Overrides the
constructor's |
prev_caption_window_size |
int
|
Number of preceding segment captions passed as
context when captioning the next segment or its keyframes. Must be a non-negative
integer. Defaults to |
prev_keyframe_window_size |
int
|
Number of preceding keyframe images, and of their
captions, passed as context when captioning the next keyframe. Lower it for models with a
small image limit. Must be a non-negative integer. Defaults to
|
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
Hybrid flow caption output with structured metadata. |
Example
converter = HybridVideoToCaption.from_preset("e2e_lm_only")
result = await converter.convert(
"video.mp4",
title="Inspection footage",
number_of_captions=6,
domain_knowledge="industrial safety audit",
prev_keyframe_window_size=3,
multimodal_context=["Prioritize hazards and operator actions."],
event_emitter=my_event_emitter,
)
print(result.result)
from_preset(preset_name=HybridPreset.E2E_LM_ONLY, **kwargs)
classmethod
Create a HybridVideoToCaption using a preset configuration.
Delegates preset_name to the preset registry in preset.py to retrieve
the appropriate component composition and pipeline mode, then wraps it in
a HybridVideoToCaption. Extra arguments can be injected via kwargs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
HybridPreset | str | None
|
Preset name forwarded to the registry.
Defaults to |
E2E_LM_ONLY
|
**kwargs
|
Any
|
See Other Parameters below. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
video_captioner_kwargs |
dict | None
|
Overrides forwarded to the video
captioner constructor (e.g. |
image_captioner_kwargs |
dict | None
|
Overrides forwarded to the keyframe
image captioner constructor. Popped before |
segmenter_kwargs |
dict | None
|
Overrides forwarded to the preset
segmenter constructor ( |
transcriber_kwargs |
dict | None
|
Overrides forwarded to the audio
transcriber builder. Popped before |
keyframe_extractor_kwargs |
dict | None
|
Overrides forwarded to the keyframe
extractor constructor. Popped before |
ocr_converter_kwargs |
dict | None
|
Overrides forwarded to the OCR
converter constructor. Popped before |
caption_summarizer_kwargs |
dict | None
|
Overrides forwarded to the
|
pipeline_mode |
PipelineMode | str
|
Override the active pipeline mode set
by the preset. Merged into |
**kwargs |
Any
|
Any other |
Returns:
| Name | Type | Description |
|---|---|---|
HybridVideoToCaption |
'HybridVideoToCaption'
|
A fully initialised hybrid captioner. |
Example
from gllm_multimodal.modality_converter.video_to_text.video_to_caption import (
HybridVideoToCaption,
HybridPreset,
)
captioner = HybridVideoToCaption.from_preset(
preset_name=HybridPreset.E2E_COMPREHENSIVE,
video_captioner_kwargs={
"lm_invoker_kwargs": {"model_id": "google/gemini-3.1-flash-lite"},
},
transcriber_kwargs={
"api_key": "sk-...",
"model": "openai/whisper-large-v3",
},
segmenter_kwargs={"config": {"threshold": 30.0, "min_shot_duration": 1.5}},
)
result = await captioner.convert(source="video.mp4")
print(result.result)
Example
Text-detection keyframes with OCR (e2e_ocr_driven)
captioner = HybridVideoToCaption.from_preset(
preset_name=HybridPreset.E2E_OCR_DRIVEN,
)
result = await captioner.convert(source="slides.mp4")
# Each keyframe carries both a visual caption and the raw on-screen text:
# result.metadata.segments[i].keyframes[j].caption
# result.metadata.segments[i].keyframes[j].detector_result["ocr"]["content"]
register_flow(mode, handler, required_components=None)
classmethod
Register a new pipeline flow.
After registration the mode string can be passed as pipeline_mode
to any new HybridVideoToCaption instance.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mode
|
str
|
Unique key for this flow (e.g. |
required |
handler
|
FlowHandler
|
An async callable with the signature: |
required |
required_components
|
list[str | tuple[str, ...]] | None
|
Names of |
None
|
Example
Register a custom flow
async def my_flow(
ctx: FlowContext, attachment, caption_data, **kwargs
) -> VideoCaptionMetadata | None:
result = await ctx.video_captioner.convert(attachment, **kwargs)
return VideoCaptionMetadata(
video_summary=result.result,
segments=[],
)
HybridVideoToCaption.register_flow(
mode="my_flow",
handler=my_flow,
required_components=["segmenter"],
)
LMBasedVideoToCaption(lm_request_processor, transform=None, max_retries=2, **kwargs)
Bases: BaseVideoToCaption, UsesLM
Video captioning implementation using Language Models.
This class implements the VideoToCaption interface using LMs for generating natural language captions.
Initialize the LM based video captioning component.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lm_request_processor
|
LMRequestProcessor
|
Language model request processor instance that supports multimodal inputs. |
required |
transform
|
list[MediaToolkit] | None
|
Processor(s) to transform the video attachment. |
None
|
max_retries
|
int
|
Maximum number of retries for LM invocations. Defaults to 2. |
2
|
**kwargs
|
Any
|
Additional keyword arguments to pass to the parent constructor. |
{}
|
convert(source, **kwargs)
async
Convert video into caption output using the LM-based caption implementation.
Behavior
- Validates source input and resolves it to a video attachment.
- Builds caption context from incoming keyword arguments.
- Passes video plus multimodal context into LM request processing.
- Validates LM response schema and segment timing constraints.
- Retries LM generation up to
max_retrieson recoverable validation issues. - Returns structured caption metadata and summary in
TextResult.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
Video source to caption. Supported forms: 1. Raw video bytes. 2. Local file path. 3. URL string. 4. Base64 encoded video string. 5. Attachment-compatible source supported by media loader. |
required |
**kwargs
|
Any
|
Captioning options forwarded to runtime components. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
number_of_captions |
int
|
Requested number of caption segments. |
title |
str
|
One-line title or summary hint. |
description |
str
|
Additional narrative context. |
domain_knowledge |
str
|
Domain-specific guidance for captioning. |
multimodal_context |
list[Attachment | str]
|
Extra context attachments. |
delete_multimodal_context |
bool
|
Whether multimodal context is removed after processing. |
event_emitter |
Any
|
Event emitter passed to LM request processing. |
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
Video summary text with segment-level caption metadata. |
Example
from gllm_inference.schema import Attachment
converter = LMBasedVideoToCaption.from_preset()
result = await converter.convert(
"video.mp4",
title="Product demo",
number_of_captions=5,
multimodal_context=[
"Focus on user actions and UI feedback.",
Attachment.from_file("release_notes.png"),
],
delete_multimodal_context=False,
event_emitter=my_event_emitter,
)
print(result.result)
from_preset(preset_name='default', lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs)
classmethod
Initialize the LM based video captioning component using preset model configurations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
str | None
|
Preset name forwarded to video-caption preset registry.
Defaults to |
'default'
|
lm_invoker_kwargs
|
dict | None
|
Keyword arguments passed to LM invoker
creation in the preset factory. Defaults to None. Valid keys follow
[ |
None
|
prompt_builder_kwargs
|
dict | None
|
Keyword arguments passed to prompt
builder creation in the preset factory. Defaults to None. Valid keys follow
|
None
|
**kwargs
|
Any
|
Additional kwargs for current-class |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Any additional kwargs merged into preset kwargs and forwarded to
|
Returns:
| Name | Type | Description |
|---|---|---|
LMBasedVideoToCaption |
LMBasedVideoToCaption
|
Initialized video captioning component using preset model. |
Example
from gllm_multimodal.modality_converter.video_to_text.video_to_caption import (
LMBasedVideoToCaption,
)
# Use default preset (Gemini flash-lite)
captioner = LMBasedVideoToCaption.from_preset()
# Override model via lm_invoker_kwargs
captioner = LMBasedVideoToCaption.from_preset(
lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)
result = await captioner.convert(source="video.mp4")
print(result.result)