Overview
Keyframe extractor toolkit for video processing.
This module provides classes for extracting keyframes from video streams based on uniform sampling or content-aware selection strategies.
Exported Classes
BaseKeyframeExtractor-- Abstract base for keyframe extraction.UniformKeyframeExtractor-- Uniformly spaced keyframes (middle frame whennum_frames=1).LDDRKeyframeExtractor-- Linear-DPP dynamic-resolution keyframe selection.TextTransitionKeyframeExtractor-- Settled text-transition keyframes from any text detector.TextTransitionConfig-- Tuning parameters for text-transition keyframe extraction.
BaseKeyframeExtractor()
Bases: CompositeMediaMixin, MediaToolkit[Attachment, list[Attachment]], ABC
Abstract base class for keyframe extraction from video attachments.
Keyframe extractors identify representative frames (extract) and optionally
materialize them into image attachments (materialize).
extract(attachment)
async
Compute keyframe plans for one attachment without materializing frame bytes.
Guarantees mimetype validation before calling _extract.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
The video attachment to analyze. |
required |
Returns:
| Type | Description |
|---|---|
list[Keyframe]
|
list[Keyframe]: Keyframe plans containing at least |
materialize(attachment, keyframe, keyframe_index=0)
async
Materialize one keyframe plan into an image attachment.
Guarantees mimetype validation before calling _materialize.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video attachment. |
required |
keyframe
|
Keyframe
|
Keyframe plan with |
required |
keyframe_index
|
int
|
Zero-based index within the keyframe plan list. Defaults to 0. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Materialized keyframe attachment with |
output_validator(attachment)
Validate that each output attachment has Keyframe-compatible metadata dict.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment | list[Attachment]
|
The attachment(s) to validate. |
required |
Returns:
| Type | Description |
|---|---|
Attachment | list[Attachment]
|
Attachment | list[Attachment]: The validated attachment(s). |
Raises:
| Type | Description |
|---|---|
TypeError
|
If any attachment metadata is not a dictionary. |
ValueError
|
If any metadata dictionary cannot be parsed as |
LDDRKeyframeExtractor(frame_budget=8, min_tokens=256, max_tokens=1024, tau=1.0, sample_fps=1.0, query=None, em_invoker=None, deinterlace=DeinterlaceMode.AUTO, decode_long_edge=DEFAULT_DECODE_LONG_EDGE)
Bases: BaseKeyframeExtractor
Extract keyframes using Linear-DPP + dynamic-resolution allocation.
The extractor performs:
1. Frame decoding at sample_fps.
2. Feature extraction for each sampled frame.
3. Greedy Linear-DPP frame selection.
4. Group-DPP-inspired importance scoring.
5. Token-budget allocation for dynamic frame resolution.
Initialize the LDDR keyframe extractor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
frame_budget
|
int
|
Maximum number of keyframe candidates. Defaults to 8. |
8
|
min_tokens
|
int
|
Minimum per-frame token allocation. Defaults to 256. |
256
|
max_tokens
|
int
|
Maximum per-frame token allocation. Defaults to 1024. |
1024
|
tau
|
float
|
Query-density prior exponent. Defaults to 1.0. |
1.0
|
sample_fps
|
float
|
Sampling FPS used before selection.
Any positive rate is accepted. The nested
|
1.0
|
query
|
str | None
|
Optional query text for query-aware LDDR selection. Defaults to None. |
None
|
em_invoker
|
BaseEMInvoker | None
|
A gllm-inference EM invoker instance. When None, keyframe selection falls back to deterministic pseudo-embeddings and is NOT content-aware (intended for tests only). A warning is logged on this path. Defaults to None. |
None
|
deinterlace
|
DeinterlaceMode | bool | str
|
Deinterlace
policy for the nested frame extraction ( |
AUTO
|
decode_long_edge
|
int | None
|
Downscale decoded frames so
the longest edge is at most this many pixels (embeddings do not
need native resolution). Original dimensions are retained for
token-to-resolution mapping. |
DEFAULT_DECODE_LONG_EDGE
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If constructor arguments are invalid. |
TextTransitionConfig
Bases: BaseModel
Tuning parameters for the settled text-transition state machine.
Attributes:
| Name | Type | Description |
|---|---|---|
persist_frames |
int
|
Number of most recent text masks kept for persistence voting. |
persist_min |
int
|
Minimum number of those masks in which a pixel must be text to
count as stable text. Cannot exceed |
activity_tau |
float
|
Time constant, in seconds, of the decaying per-pixel motion map. |
pixel_delta |
float
|
Minimum absolute grayscale change for a pixel to count as changed. |
motion_rate |
float
|
Motion-map level at or above which a pixel is treated as moving and excluded from appearance and disappearance measurements. |
appear_threshold |
float
|
Frame fraction of newly appeared stable text that starts a transition. |
vanish_threshold |
float
|
Frame fraction of disappeared reference text that starts a transition. |
settle_epsilon |
float
|
Maximum frame-to-frame stable-text change fraction treated as settled. |
settle_frames |
int
|
Consecutive settled observations required to commit a transition. |
max_transition |
float
|
Maximum transition duration, in seconds, before it is committed regardless of settling. |
max_gap |
float
|
Seconds since the last commit after which persistently moving text
is committed. |
motion_text_floor |
float
|
Frame fraction of stable text inside moving regions that marks an observation as text-in-motion. |
motion_text_runs |
int
|
Consecutive text-in-motion observations required before the
|
min_text_fraction |
float
|
Minimum stable-text frame fraction for a committed sample to be emitted as a keyframe. |
validate_persistence_window()
Reject a persistence vote larger than its window.
Returns:
| Name | Type | Description |
|---|---|---|
TextTransitionConfig |
TextTransitionConfig
|
This validated configuration. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
TextTransitionKeyframeExtractor(text_detector=DEFAULT_TEXT_DETECTOR, sample_fps=DEFAULT_SAMPLE_FPS, max_width=DEFAULT_MAX_WIDTH, max_height=DEFAULT_MAX_HEIGHT, transition_config=None, text_detector_kwargs=None)
Bases: BaseKeyframeExtractor
Select keyframes where on-screen text settles after a transition.
Text masks come from any detector implementing
BaseTextDetector.
Detectors that return boxes are rasterized into masks.
The extractor exposes three public entry points:
extractreturnsKeyframeplans (time offsets) without decoding output frames.processplans keyframes and materializes them as lossless PNG attachments in one frame-extraction call.materializedecodes one previously computed plan without rerunning analysis.
Usage Example
Keyframe images (process):
from gllm_inference.schema import Attachment
from gllm_multimodal.builder.media_toolkit_builder import build_media_toolkit
extractor = build_media_toolkit("TextTransitionKeyframeExtractor", sample_fps=2.0)
video = Attachment.from_path("/path/to/lecture.mp4")
keyframes = await extractor.process(video)
for keyframe in keyframes:
print(keyframe.filename, keyframe.metadata["time_offset"])
Keyframe plans (extract + materialize):
from gllm_inference.schema import Attachment
from gllm_multimodal.builder.media_toolkit_builder import build_media_toolkit
extractor = build_media_toolkit("TextTransitionKeyframeExtractor")
video = Attachment.from_path("/path/to/lecture.mp4")
plans = await extractor.extract(video)
print([plan.time_offset for plan in plans])
first_frame = await extractor.materialize(video, plans[0], keyframe_index=0)
Custom detector and tuning:
from gllm_multimodal.builder.media_toolkit_builder import build_media_toolkit
extractor = build_media_toolkit(
"TextTransitionKeyframeExtractor",
text_detector="PPOCRTextDetector",
text_detector_kwargs={"source": "ppocrv6_medium", "cache_dir": "~/.cache/gllm/ppocr"},
transition_config={"settle_frames": 4, "max_gap": 30.0},
)
A detector built from a class name is owned by the extractor and released by
close
or by leaving a with block. A detector instance passed in stays owned by the caller.
Build the extractor once and reuse it across videos so the detector session stays warm.
Attributes:
| Name | Type | Description |
|---|---|---|
text_detector |
BaseTextDetector
|
Synchronous text detector returning a mask or boxes. |
sample_fps |
float
|
Analysis-frame sampling rate in frames per second. |
max_width |
int
|
Maximum analysis-frame width, used only when the detector has no preferred input size. |
max_height |
int
|
Maximum analysis-frame height, used only when the detector has no preferred input size. |
transition_config |
TextTransitionConfig
|
Text-transition state-machine settings. |
Initialize detector-driven text selection.
A detector given by class name is built through the MediaToolkit registry, so the
extractor can be constructed from JSON-serializable specs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text_detector
|
BaseTextDetector | str
|
Detector instance, or the registered
class name of a |
DEFAULT_TEXT_DETECTOR
|
sample_fps
|
float
|
Analysis sampling rate in frames per second. Defaults to 1.0. |
DEFAULT_SAMPLE_FPS
|
max_width
|
int
|
Maximum analysis width when the detector has no preferred input size. Defaults to 1280. |
DEFAULT_MAX_WIDTH
|
max_height
|
int
|
Maximum analysis height when the detector has no preferred input size. Defaults to 1280. |
DEFAULT_MAX_HEIGHT
|
transition_config
|
TextTransitionConfig | dict[str, Any] | None
|
Transition settings, or a mapping validated into |
None
|
text_detector_kwargs
|
dict[str, Any] | None
|
Constructor keyword arguments
used when |
None
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If the resolved detector is not a |
ValueError
|
If |
__enter__()
Return this extractor for synchronous context management.
Returns:
| Name | Type | Description |
|---|---|---|
TextTransitionKeyframeExtractor |
TextTransitionKeyframeExtractor
|
This extractor instance. |
__exit__(exc_type, exc_value, traceback)
Release the detector if this extractor built it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
exc_type
|
object
|
Exception type supplied by the context manager. |
required |
exc_value
|
object
|
Exception value supplied by the context manager. |
required |
traceback
|
object
|
Traceback supplied by the context manager. |
required |
close()
Release the detector if this extractor built it from a class name.
A caller-supplied detector instance is left open. Safe to call multiple times.
UniformKeyframeExtractor(num_frames=DEFAULT_NUM_FRAMES, deinterlace=DeinterlaceMode.AUTO)
Bases: BaseKeyframeExtractor
Extract keyframes at uniformly spaced timestamps within a clip.
Placement uses midpoints of equal duration bins:
time_offset = duration * (i + 0.5) / num_frames.
When num_frames=1, the single keyframe is the middle frame.
Materialization uses nested FrameExtractionProcessor (CompositeMediaMixin)
with a configurable deinterlace policy (default auto: yadif runs only
when ffprobe reports an interlaced field order, so progressive sources pass
through unfiltered).
Attributes:
| Name | Type | Description |
|---|---|---|
num_frames |
int
|
Number of keyframes to extract per clip. |
deinterlace |
DeinterlaceMode
|
Deinterlace policy forwarded to the nested frame extraction call. |
Initialize the uniform keyframe extractor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_frames
|
int
|
Number of frames to sample uniformly. Defaults to 1 (middle frame). |
DEFAULT_NUM_FRAMES
|
deinterlace
|
DeinterlaceMode | bool | str
|
Deinterlace
policy for the nested frame extraction ( |
AUTO
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |