Skip to content

Overview

Keyframe extractor toolkit for video processing.

This module provides classes for extracting keyframes from video streams based on uniform sampling or content-aware selection strategies.

Exported Classes

BaseKeyframeExtractor()

Bases: CompositeMediaMixin, MediaToolkit[Attachment, list[Attachment]], ABC

Abstract base class for keyframe extraction from video attachments.

Keyframe extractors identify representative frames (extract) and optionally materialize them into image attachments (materialize).

extract(attachment) async

Compute keyframe plans for one attachment without materializing frame bytes.

Guarantees mimetype validation before calling _extract.

Parameters:

Name Type Description Default
attachment Attachment

The video attachment to analyze.

required

Returns:

Type Description
list[Keyframe]

list[Keyframe]: Keyframe plans containing at least time_offset.

materialize(attachment, keyframe, keyframe_index=0) async

Materialize one keyframe plan into an image attachment.

Guarantees mimetype validation before calling _materialize.

Parameters:

Name Type Description Default
attachment Attachment

Source video attachment.

required
keyframe Keyframe

Keyframe plan with time_offset.

required
keyframe_index int

Zero-based index within the keyframe plan list. Defaults to 0.

0

Returns:

Name Type Description
Attachment Attachment

Materialized keyframe attachment with Keyframe metadata.

output_validator(attachment)

Validate that each output attachment has Keyframe-compatible metadata dict.

Parameters:

Name Type Description Default
attachment Attachment | list[Attachment]

The attachment(s) to validate.

required

Returns:

Type Description
Attachment | list[Attachment]

Attachment | list[Attachment]: The validated attachment(s).

Raises:

Type Description
TypeError

If any attachment metadata is not a dictionary.

ValueError

If any metadata dictionary cannot be parsed as Keyframe.

LDDRKeyframeExtractor(frame_budget=8, min_tokens=256, max_tokens=1024, tau=1.0, sample_fps=1.0, query=None, em_invoker=None, deinterlace=DeinterlaceMode.AUTO, decode_long_edge=DEFAULT_DECODE_LONG_EDGE)

Bases: BaseKeyframeExtractor

Extract keyframes using Linear-DPP + dynamic-resolution allocation.

The extractor performs: 1. Frame decoding at sample_fps. 2. Feature extraction for each sampled frame. 3. Greedy Linear-DPP frame selection. 4. Group-DPP-inspired importance scoring. 5. Token-budget allocation for dynamic frame resolution.

Initialize the LDDR keyframe extractor.

Parameters:

Name Type Description Default
frame_budget int

Maximum number of keyframe candidates. Defaults to 8.

8
min_tokens int

Minimum per-frame token allocation. Defaults to 256.

256
max_tokens int

Maximum per-frame token allocation. Defaults to 1024.

1024
tau float

Query-density prior exponent. Defaults to 1.0.

1.0
sample_fps float

Sampling FPS used before selection. Any positive rate is accepted. The nested FrameSamplingProcessor only supports whole FPS values, so rates >= 1 are rounded to the nearest int and rates below 1 are realized by sampling at 1 fps and keeping one frame per 1 / sample_fps seconds. Defaults to 1.0.

1.0
query str | None

Optional query text for query-aware LDDR selection. Defaults to None.

None
em_invoker BaseEMInvoker | None

A gllm-inference EM invoker instance. When None, keyframe selection falls back to deterministic pseudo-embeddings and is NOT content-aware (intended for tests only). A warning is logged on this path. Defaults to None.

None
deinterlace DeinterlaceMode | bool | str

Deinterlace policy for the nested frame extraction (off / force / auto; True→force, False→off). Defaults to auto so progressive sources skip the yadif pass.

AUTO
decode_long_edge int | None

Downscale decoded frames so the longest edge is at most this many pixels (embeddings do not need native resolution). Original dimensions are retained for token-to-resolution mapping. None keeps native resolution. Defaults to 512.

DEFAULT_DECODE_LONG_EDGE

Raises:

Type Description
ValueError

If constructor arguments are invalid.

TextTransitionConfig

Bases: BaseModel

Tuning parameters for the settled text-transition state machine.

Attributes:

Name Type Description
persist_frames int

Number of most recent text masks kept for persistence voting.

persist_min int

Minimum number of those masks in which a pixel must be text to count as stable text. Cannot exceed persist_frames.

activity_tau float

Time constant, in seconds, of the decaying per-pixel motion map.

pixel_delta float

Minimum absolute grayscale change for a pixel to count as changed.

motion_rate float

Motion-map level at or above which a pixel is treated as moving and excluded from appearance and disappearance measurements.

appear_threshold float

Frame fraction of newly appeared stable text that starts a transition.

vanish_threshold float

Frame fraction of disappeared reference text that starts a transition.

settle_epsilon float

Maximum frame-to-frame stable-text change fraction treated as settled.

settle_frames int

Consecutive settled observations required to commit a transition.

max_transition float

Maximum transition duration, in seconds, before it is committed regardless of settling.

max_gap float

Seconds since the last commit after which persistently moving text is committed. 0 disables this path.

motion_text_floor float

Frame fraction of stable text inside moving regions that marks an observation as text-in-motion.

motion_text_runs int

Consecutive text-in-motion observations required before the max_gap commit path applies.

min_text_fraction float

Minimum stable-text frame fraction for a committed sample to be emitted as a keyframe.

validate_persistence_window()

Reject a persistence vote larger than its window.

Returns:

Name Type Description
TextTransitionConfig TextTransitionConfig

This validated configuration.

Raises:

Type Description
ValueError

If persist_min exceeds persist_frames.

TextTransitionKeyframeExtractor(text_detector=DEFAULT_TEXT_DETECTOR, sample_fps=DEFAULT_SAMPLE_FPS, max_width=DEFAULT_MAX_WIDTH, max_height=DEFAULT_MAX_HEIGHT, transition_config=None, text_detector_kwargs=None)

Bases: BaseKeyframeExtractor

Select keyframes where on-screen text settles after a transition.

Text masks come from any detector implementing BaseTextDetector. Detectors that return boxes are rasterized into masks.

The extractor exposes three public entry points:

  • extract returns Keyframe plans (time offsets) without decoding output frames.
  • process plans keyframes and materializes them as lossless PNG attachments in one frame-extraction call.
  • materialize decodes one previously computed plan without rerunning analysis.

Usage Example

Keyframe images (process):

from gllm_inference.schema import Attachment
from gllm_multimodal.builder.media_toolkit_builder import build_media_toolkit

extractor = build_media_toolkit("TextTransitionKeyframeExtractor", sample_fps=2.0)

video = Attachment.from_path("/path/to/lecture.mp4")
keyframes = await extractor.process(video)
for keyframe in keyframes:
    print(keyframe.filename, keyframe.metadata["time_offset"])

Keyframe plans (extract + materialize):

from gllm_inference.schema import Attachment
from gllm_multimodal.builder.media_toolkit_builder import build_media_toolkit

extractor = build_media_toolkit("TextTransitionKeyframeExtractor")

video = Attachment.from_path("/path/to/lecture.mp4")
plans = await extractor.extract(video)
print([plan.time_offset for plan in plans])

first_frame = await extractor.materialize(video, plans[0], keyframe_index=0)

Custom detector and tuning:

from gllm_multimodal.builder.media_toolkit_builder import build_media_toolkit

extractor = build_media_toolkit(
    "TextTransitionKeyframeExtractor",
    text_detector="PPOCRTextDetector",
    text_detector_kwargs={"source": "ppocrv6_medium", "cache_dir": "~/.cache/gllm/ppocr"},
    transition_config={"settle_frames": 4, "max_gap": 30.0},
)

A detector built from a class name is owned by the extractor and released by close or by leaving a with block. A detector instance passed in stays owned by the caller. Build the extractor once and reuse it across videos so the detector session stays warm.

Attributes:

Name Type Description
text_detector BaseTextDetector

Synchronous text detector returning a mask or boxes.

sample_fps float

Analysis-frame sampling rate in frames per second.

max_width int

Maximum analysis-frame width, used only when the detector has no preferred input size.

max_height int

Maximum analysis-frame height, used only when the detector has no preferred input size.

transition_config TextTransitionConfig

Text-transition state-machine settings.

Initialize detector-driven text selection.

A detector given by class name is built through the MediaToolkit registry, so the extractor can be constructed from JSON-serializable specs.

Parameters:

Name Type Description Default
text_detector BaseTextDetector | str

Detector instance, or the registered class name of a BaseTextDetector. Defaults to "PPOCRTextDetector".

DEFAULT_TEXT_DETECTOR
sample_fps float

Analysis sampling rate in frames per second. Defaults to 1.0.

DEFAULT_SAMPLE_FPS
max_width int

Maximum analysis width when the detector has no preferred input size. Defaults to 1280.

DEFAULT_MAX_WIDTH
max_height int

Maximum analysis height when the detector has no preferred input size. Defaults to 1280.

DEFAULT_MAX_HEIGHT
transition_config TextTransitionConfig | dict[str, Any] | None

Transition settings, or a mapping validated into TextTransitionConfig. Defaults to None, which uses the default settings.

None
text_detector_kwargs dict[str, Any] | None

Constructor keyword arguments used when text_detector is a class name. Defaults to None.

None

Raises:

Type Description
TypeError

If the resolved detector is not a BaseTextDetector or transition_config has an unsupported type.

ValueError

If text_detector names an unregistered class, text_detector_kwargs is combined with a detector instance, or sample_fps, max_width, max_height, or transition_config is invalid.

__enter__()

Return this extractor for synchronous context management.

Returns:

Name Type Description
TextTransitionKeyframeExtractor TextTransitionKeyframeExtractor

This extractor instance.

__exit__(exc_type, exc_value, traceback)

Release the detector if this extractor built it.

Parameters:

Name Type Description Default
exc_type object

Exception type supplied by the context manager.

required
exc_value object

Exception value supplied by the context manager.

required
traceback object

Traceback supplied by the context manager.

required

close()

Release the detector if this extractor built it from a class name.

A caller-supplied detector instance is left open. Safe to call multiple times.

UniformKeyframeExtractor(num_frames=DEFAULT_NUM_FRAMES, deinterlace=DeinterlaceMode.AUTO)

Bases: BaseKeyframeExtractor

Extract keyframes at uniformly spaced timestamps within a clip.

Placement uses midpoints of equal duration bins: time_offset = duration * (i + 0.5) / num_frames. When num_frames=1, the single keyframe is the middle frame.

Materialization uses nested FrameExtractionProcessor (CompositeMediaMixin) with a configurable deinterlace policy (default auto: yadif runs only when ffprobe reports an interlaced field order, so progressive sources pass through unfiltered).

Attributes:

Name Type Description
num_frames int

Number of keyframes to extract per clip.

deinterlace DeinterlaceMode

Deinterlace policy forwarded to the nested frame extraction call.

Initialize the uniform keyframe extractor.

Parameters:

Name Type Description Default
num_frames int

Number of frames to sample uniformly. Defaults to 1 (middle frame).

DEFAULT_NUM_FRAMES
deinterlace DeinterlaceMode | bool | str

Deinterlace policy for the nested frame extraction (off / force / auto; True→force, False→off). Defaults to auto.

AUTO

Raises:

Type Description
ValueError

If num_frames is not positive or deinterlace cannot be coerced.