Overview
Media toolkit processor public API and loading entrypoint.
This package exposes processor families and shared backend base classes used by
builder/factory flows. Importing from here ensures family classes and concrete
backend implementations are loaded so subclass registration is populated in
MediaToolkit.registry and each family's backend map.
Design overview
- Family package layout:
<family>/base.pyfor the abstract family class and<family>/<backend>_backend.pyfor concrete implementations. - Construction:
MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer"). - Composite components (segmenters/keyframe extractors):
cache nested processors by class name + backend and can switch backend by
changing
component.backendbefore callingget_processor(...).
Contributor quick-start for a new processor family
- Create
processor/<new_family>/base.pysubclassingBackendSelectableProcessor. - Add backend implementations declaring
BACKEND(+IS_DEFAULTif needed). - Re-export in
processor/<new_family>/__init__.pyand this module. - Add tests for family registration and factory creation.
AudioExtractionProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for extracting audio tracks from video attachments.
This class serves as a unified entry point for audio extraction operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.
Why use this base class?
- Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
- Simplicity: No need to handle fallback logic or conditional imports yourself.
- Future-proofing: New backends can be added to the library without requiring changes to your application code.
Usage Example
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment
# Instantiates the best available backend automatically
processor = AudioExtractionProcessor.build()
attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment
# Explicitly force the ffmpeg backend
processor = AudioExtractionProcessor.build(backend="ffmpeg")
attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment
# Explicitly force the moviepy backend
processor = AudioExtractionProcessor.build(backend="moviepy")
attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
BackendSelectableProcessor()
Bases: MediaToolkit[T_in, T_out], ABC
Abstract base for a processor family with pluggable backends.
Each family owns its backend registry. Concrete backend classes are
auto-registered via __init_subclass__ by declaring:
BACKEND: backend key, e.g.gstreamerorffmpegIS_DEFAULT: whether this backend is the family default
Why this exists
It lets callers construct by stable family class name while deferring runtime selection of backend implementation.
Minimal contributor pattern
# base.py
class ImageTilingProcessor(BackendSelectableProcessor):
pass
# pil_backend.py
class PilImageTilingProcessor(ImageTilingProcessor):
BACKEND = "pil"
# cv2_backend.py
class Cv2ImageTilingProcessor(ImageTilingProcessor):
BACKEND = "cv2"
Then callers can use
MediaToolkit.build("ImageTilingProcessor", backend="pil").
__init_subclass__(**kwargs)
Automatically register backend classes into their family registry.
This hook runs at class definition/import time and maintains per-family backend mappings used by backend-selectable builds.
Registration flow
- If
clsis the family base itself (e.g.VideoClipProcessor), reset_backendsand_default_backendfor that family. - If
clsis abstract, skip registration. - Otherwise treat
clsas a concrete backend implementation: - require class variable
BACKEND(e.g."gstreamer"), - reject duplicate backend keys within the same family,
- add mapping
family_base._backends[BACKEND] = cls.
Default backend resolution
- If
IS_DEFAULT=Trueon a concrete class, that backend becomes the family's default, even if a first-registered fallback already exists. - If no explicit default exists yet, the first registered backend is used as fallback default.
- Multiple
IS_DEFAULT=Truedeclarations in one family raiseValueError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
**kwargs
|
Any
|
Extra class declaration keyword arguments forwarded to
parent |
{}
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If a concrete backend class omits |
ValueError
|
If duplicate backend keys are registered in one family. |
ValueError
|
If more than one backend declares |
build(backend=None, **kwargs)
classmethod
Build a backend implementation for this processor family.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
backend
|
str | None
|
Backend key. Uses the family default when omitted. |
None
|
**kwargs
|
Any
|
Constructor kwargs forwarded to the backend class. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
BackendSelectableProcessor |
BackendSelectableProcessor
|
Instantiated backend processor. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the backend is unknown or no default is configured. |
Example
For a family class AudioExtractionProcessor,
AudioExtractionProcessor.build(backend="gstreamer") returns the
registered GStreamer implementation class instance.
<GstAudioExtractionProcessor instance>
build_from_registry(backend=None, **kwargs)
classmethod
Build a family backend or instantiate a concrete backend class.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
backend
|
str | MediaBackend | None
|
Backend key for family resolution. Defaults to None. |
None
|
**kwargs
|
Any
|
Constructor kwargs forwarded to the backend class. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
BackendSelectableProcessor |
BackendSelectableProcessor
|
Instantiated processor. |
Example
processor = AudioExtractionProcessor.build_from_registry(backend="gstreamer")
<GstAudioExtractionProcessor instance>
get_install_hint(backend=None)
classmethod
Return the installation hint for a backend.
Reads INSTALL_HINT from the registered backend class. If the backend
is not registered or carries no hint, returns a generic fallback.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
backend
|
str | None
|
Backend key to look up. Defaults to
|
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Human-readable install hint for the backend, or a generic fallback. |
is_family_base(processor_cls)
classmethod
Return whether processor_cls is a family abstract base.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
processor_cls
|
type[BackendSelectableProcessor]
|
Candidate processor class. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
|
Example
assert BackendSelectableProcessor.is_family_base(AudioExtractionProcessor) is True
True
list_available_backends()
classmethod
Return registered backend keys for backend-selectable family bases.
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Backend keys available for |
Example
AudioExtractionProcessor.list_available_backends()
# e.g. ["gstreamer", "ffmpeg", "moviepy"]
["gstreamer", "ffmpeg", "moviepy"]
list_backends()
classmethod
Return registered backend keys for this processor family.
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Backend keys available for |
Example
backends = AudioExtractionProcessor.list_backends()
["gstreamer", "ffmpeg", "moviepy"]
BaseFFmpegProcessor(config=None)
Bases: MediaToolkit[Attachment, Attachment], Generic[TConfig]
Shared helpers for processors that shell out to the ffmpeg binary.
Unlike BaseGstreamerProcessor,
this base does not require ffmpeg at construction time — availability is
checked when a command is run. That keeps import / build paths light when
ffmpeg is only needed for optional I/O.
Attributes:
| Name | Type | Description |
|---|---|---|
INSTALL_HINT |
str
|
Human-readable install guidance for missing ffmpeg. |
config |
TConfig
|
Coerced stable configuration from |
Initialize the FFmpeg helper base and coerce constructor config.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | BaseModel | None
|
Stable config
dict, family/backend model, or |
None
|
config_model()
classmethod
Return the stable configuration model for this processor.
Concrete backends must override with their FFmpeg config type.
Returns:
| Type | Description |
|---|---|
type[TConfig]
|
type[TConfig]: The configuration class used at construction time. |
BaseGstreamerProcessor(config=None, *, enable_video=True, enable_audio=True)
Bases: MediaToolkit[Attachment, Attachment], Generic[TConfig]
Abstract base class for GStreamer-powered processors.
Subclasses must implement only _execute_pipeline. All common
concerns — availability checks, Conda environment configuration, GStreamer
initialisation, encoder/element selection, the standard EOS/error bus loop,
and temporary-file cleanup — are handled here.
Parameterise with the concrete config model, e.g.
BaseGstreamerProcessor[GstFrameDecodeConfig].
Typical subclass skeleton::
class MyGstProcessor(BaseGstreamerProcessor[GstBaseConfig]):
def __init__(self, my_param, config=None):
super().__init__(config=config)
self.my_param = my_param
async def _process(self, attachment: Attachment) -> Attachment:
return await self._process_single(attachment)
async def _execute_pipeline(self, input_path, output_path):
# build & run YOUR GStreamer pipeline here
...
Attributes:
| Name | Type | Description |
|---|---|---|
logger |
Logger bound to the concrete subclass name. |
|
config |
TConfig
|
Runtime configuration for this processor family. |
video_encoder_info |
EncoderFormatInfo | None
|
Selected video encoder
format info, or |
audio_encoder_info |
EncoderFormatInfo | None
|
Selected audio encoder
format info, or |
Verify GStreamer availability, configure the environment, and select encoders.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | GstBaseConfig | BaseModel | None
|
Optional configuration.
A plain |
None
|
enable_video
|
bool
|
When |
True
|
enable_audio
|
bool
|
When |
True
|
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If GStreamer is not installed or cannot be initialised. |
RuntimeError
|
If no suitable video encoder is found in the registry. |
TypeError
|
If |
config_model()
classmethod
Return the stable configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[TConfig]
|
type[TConfig]: The configuration class used at construction time. |
DeinterlaceConfig
Bases: BaseModel
Backend-agnostic stable configuration for deinterlace processors.
Only knobs that apply across FFmpeg / GStreamer (and future engines) belong here. Engine-specific fields (yadif mode, x264 CRF, GST element props, …) live on backend subclasses that inherit this model.
Attributes:
| Name | Type | Description |
|---|---|---|
strip_audio |
bool
|
Drop audio when producing progressive output. Defaults to True. |
DeinterlaceMode
Bases: StrEnum
How frame extraction applies deinterlace (e.g. FFmpeg yadif).
Attributes:
| Name | Type | Description |
|---|---|---|
OFF |
Never deinterlace. |
|
FORCE |
Always deinterlace (ignore field-order metadata). |
|
AUTO |
Deinterlace only when ffprobe reports an interlaced field order. |
DeinterlaceProcessConfig
Bases: ProcessorProcessConfig
Backend-agnostic per-invocation overrides for deinterlace.
None fields fall back to the stable constructor config. Backend
process-config models inherit this class and add engine-specific overrides.
DeinterlaceProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for deinterlacing video attachments.
Why use this base class?
- Portability: Swap FFmpeg vs GStreamer without changing call sites.
- I/O boundary: Deinterlace is media transform I/O, not a segmenter algorithm.
- Tunable: Shared knobs on
DeinterlaceConfig; backends extend for engine-specific settings via inheritance.
Usage
from gllm_multimodal.media_toolkit.processor.deinterlace_processor import (
DeinterlaceConfig,
DeinterlaceProcessor,
)
processor = DeinterlaceProcessor.build(
backend="ffmpeg",
config=DeinterlaceConfig(strip_audio=True),
)
progressive = await processor.process(video_attachment)
is_interlaced(attachment)
staticmethod
Return whether ffprobe reports an interlaced field order for an attachment.
Writes attachment.data to a temporary file, probes field_order, then
cleans up. Returns False when ffprobe is unavailable, the probe fails,
or the field order is progressive or unknown (fail-open: unknown is
treated as progressive so callers skip the deinterlace pass; a warning
is logged whenever the probe cannot determine interlacing).
Callers that need fail-safe behavior (deinterlace when unknown) should check ffprobe availability separately instead of relying on this gate.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Video attachment to probe. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
|
is_interlaced_path(video_path)
staticmethod
Return whether ffprobe reports an interlaced field order for a local path.
Public alias of
is_interlaced_path
kept on the family base so frame-extraction and other families do not
reach into a private member.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
video_path
|
str
|
Path to a local video file. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
|
FFmpegBaseConfig
Bases: BaseModel
Minimal shared configuration for FFmpeg CLI processors.
Family backends extend this (or their own shared family config) with engine-specific fields. Kept intentionally thin — FFmpeg leaves share process helpers more than stable knobs.
FFmpegDeinterlaceConfig
Bases: DeinterlaceConfig
FFmpeg-specific stable config (inherits shared DeinterlaceConfig).
Attributes:
| Name | Type | Description |
|---|---|---|
yadif_mode |
int
|
Yadif mode (0=send_frame, 1=send_field, 2=send_frame_nospatial, 3=send_field_nospatial). Defaults to 0. |
yadif_parity |
int
|
Field parity (-1=auto, 0=tff, 1=bff). Defaults to -1. |
yadif_deint |
int
|
Deinterlace all frames (0) or only flagged (1). Defaults to 0. |
video_codec |
str
|
Video encoder name. Defaults to |
preset |
str
|
x264 preset. Defaults to |
crf |
int
|
Constant rate factor (0–51). Defaults to 23. |
filter_override |
str | None
|
Full |
validate_non_empty(value)
classmethod
Reject blank codec/preset strings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
Candidate string. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Validated string. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
yadif_filter()
Build the -vf filter string.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Filter graph for FFmpeg |
FFmpegDeinterlaceProcessConfig
Bases: DeinterlaceProcessConfig
Per-invocation FFmpeg overrides (inherits shared DeinterlaceProcessConfig).
Any field left as None falls back to the stable constructor config.
FFmpegDeinterlaceProcessor(config=None)
Bases: BaseFFmpegProcessor[FFmpegDeinterlaceConfig], DeinterlaceProcessor
Deinterlace video with FFmpeg yadif.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
FFmpegDeinterlaceConfig
|
Stable constructor configuration. |
Initialize the FFmpeg deinterlace processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | DeinterlaceConfig | FFmpegDeinterlaceConfig | None
|
Shared |
None
|
config_model()
classmethod
Return the stable configuration model.
deinterlace_video(video_path, output_path=None, *, params=None)
Deinterlace a video path with FFmpeg yadif into a progressive file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
video_path
|
str
|
Source video path. |
required |
output_path
|
str | None
|
Destination path. When None, a
temporary |
None
|
params
|
FFmpegDeinterlaceConfig | None
|
Effective options for
this call. Defaults to |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Path to the progressive output video. |
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
If |
RuntimeError
|
If FFmpeg deinterlace fails. |
process(attachment, **kwargs)
async
Deinterlace one video attachment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Input interlaced (or progressive) video. |
required |
**kwargs
|
Any
|
May include |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Progressive video attachment. |
process_config_model()
classmethod
Return the per-invocation configuration model.
FFmpegFrameDecodeConfig
Bases: FrameDecodeFieldsMixin, FFmpegBaseConfig
FFmpeg-specific stable configuration for dense frame decoding.
FFmpegFrameDecodeProcessConfig
FFmpegFrameDecodeProcessor(config=None)
Bases: BaseFFmpegProcessor[FFmpegFrameDecodeConfig], FrameDecodeProcessor
Decode every video frame to PNG attachments with the FFmpeg CLI.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
FFmpegFrameDecodeConfig
|
Stable constructor configuration. |
Initialize the FFmpeg frame decode processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | FrameDecodeConfig | FFmpegFrameDecodeConfig | None
|
Shared or FFmpeg-specific config. Defaults to None. |
None
|
build_filter(*, sample_fps, target_width)
Build the -vf filter chain for the requested knobs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sample_fps
|
int | None
|
Downsample rate, if any. |
required |
target_width
|
int | None
|
Downscale width, if any. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: Comma-joined filter chain, or |
config_model()
classmethod
Return the stable configuration model.
process(attachment, **kwargs)
async
Decode every frame of one video attachment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
**kwargs
|
Any
|
May include |
{}
|
Returns:
| Type | Description |
|---|---|
list[Attachment]
|
list[Attachment]: PNG image attachments in decode order. |
process_config_model()
classmethod
Return the per-invocation configuration model.
FFmpegFrameExtractionConfig
Bases: FrameExtractionConfig
FFmpeg-specific stable config (inherits shared FrameExtractionConfig).
Yadif fields apply when deinterlace is force or auto (and the
probe selects yadif for auto).
Attributes:
| Name | Type | Description |
|---|---|---|
yadif_mode |
int
|
Yadif mode. Defaults to 0. |
yadif_parity |
int
|
Field parity (-1=auto). Defaults to -1. |
yadif_deint |
int
|
Deinterlace all (0) or flagged-only (1). Defaults to 0. |
filter_override |
str | None
|
Full |
yadif_filter()
Build the -vf filter string for deinterlaced extract.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Filter graph for FFmpeg |
FFmpegFrameExtractionProcessConfig
Bases: FrameExtractionProcessConfig
Per-invocation FFmpeg overrides (inherits shared process config).
FFmpegFrameExtractionProcessor(config=None)
Bases: BaseFFmpegProcessor[FFmpegFrameExtractionConfig], FrameExtractionProcessor
Extract image frames with FFmpeg (optional yadif).
Attributes:
| Name | Type | Description |
|---|---|---|
config |
FFmpegFrameExtractionConfig
|
Stable constructor configuration. |
Initialize the FFmpeg frame extraction processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | FrameExtractionConfig | FFmpegFrameExtractionConfig | None
|
Shared or FFmpeg-specific config. Defaults to None. |
None
|
config_model()
classmethod
Return the stable configuration model.
extract_frames(video_path, timestamps, *, params=None)
Extract encoded image bytes at each timestamp from a video path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
video_path
|
str
|
Source video path. |
required |
timestamps
|
list[float]
|
Non-empty non-negative timestamps. |
required |
params
|
FFmpegFrameExtractionConfig | None
|
Effective
options. Defaults to |
None
|
Returns:
| Type | Description |
|---|---|
list[bytes]
|
list[bytes]: Encoded frames in the same order as |
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
If ffmpeg is unavailable. |
RuntimeError
|
If any extraction fails. |
ValueError
|
If timestamps are invalid. |
process(attachment, **kwargs)
async
Extract frames from one video attachment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
**kwargs
|
Any
|
May include |
{}
|
Returns:
| Type | Description |
|---|---|
list[Attachment]
|
list[Attachment]: One image attachment per timestamp. |
process_config_model()
classmethod
Return the per-invocation configuration model.
set_timestamps(timestamps)
Update constructor-level default timestamps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamps
|
list[float]
|
Non-empty non-negative timestamps. |
required |
stream_frames(attachment, sample_fps, deinterlace=None)
async
Stream sampled BGR frames from a single FFmpeg process.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
sample_fps
|
float
|
Positive sampling rate in frames per second. |
required |
deinterlace
|
DeinterlaceMode | None
|
Deinterlace override. Defaults to None. |
None
|
Yields:
| Type | Description |
|---|---|
AsyncIterator[tuple[float, ndarray]]
|
tuple[float, np.ndarray]: Timestamp and BGR |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
FileNotFoundError
|
If ffmpeg or ffprobe is unavailable. |
RuntimeError
|
If probing or decoding fails. |
FrameDecodeConfig
FrameDecodeProcessConfig
Bases: ProcessorProcessConfig
Backend-agnostic per-invocation frame decode config.
Attributes:
| Name | Type | Description |
|---|---|---|
sample_fps |
int | None
|
Optional rate override. |
target_width |
int | None
|
Optional width override. |
FrameDecodeProcessor()
Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC
Family base for dense video-frame decoding.
Why use this base class?
- Portability: Swap FFmpeg vs GStreamer without changing call sites.
- FIPS choice:
backend="gstreamer"decodes with system plugins (no bundled-FFmpeg wheels);backend="ffmpeg"shells out to the systemffmpegbinary. - Separation: Frame decoding stays here; frame scoring (shot detection, keyframes) lives in segmenters/extractors.
Usage
from gllm_multimodal.media_toolkit.processor.frame_decode_processor import (
FrameDecodeProcessor,
)
processor = FrameDecodeProcessor.build(backend="gstreamer")
frames = await processor.process(video_attachment)
frame_metadata(frame_index, fps)
Build the metadata dict attached to every decoded frame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
frame_index
|
int
|
Zero-based position in decode order. |
required |
fps
|
float | None
|
Effective sampling rate, if known. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
dict[str, Any]: |
iter_rgb_frames_sync(attachment, *, sample_fps=None, target_width=None, as_rgb=True)
Yield decoded frames after the backend writes the full PNG sequence.
The decoder subprocess/pipeline completes first, so peak temp-disk
usage is every sampled PNG at once. After that, this generator opens
each file in order, yields the payload, and deletes the PNG so Python
RAM stays O(1) in frames. Prefer this over process when callers
only need a scored stream and can tolerate the peak-disk cost.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
sample_fps
|
int | None
|
Override downsample rate. Defaults to the constructor config. |
None
|
target_width
|
int | None
|
Override downscale width. Defaults to the constructor config. |
None
|
as_rgb
|
bool
|
When True (default), yield RGB arrays.
When False, yield raw PNG bytes (used by |
True
|
Yields:
| Type | Description |
|---|---|
tuple[Any, dict[str, Any]]
|
tuple[Any, dict[str, Any]]: |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If decoding fails or yields no frames. |
FileNotFoundError
|
If a required decoder binary is missing. |
FrameExtractionConfig
Bases: BaseModel
Backend-agnostic stable configuration for frame extraction.
Yadif knobs are shared so FFmpeg can inline -vf yadif and GStreamer can
forward the same values to
DeinterlaceProcessor.
Attributes:
| Name | Type | Description |
|---|---|---|
output_format |
str
|
Image encode format (JPEG/PNG). Defaults to JPEG. |
deinterlace |
DeinterlaceMode
|
Deinterlace policy ( |
yadif_mode |
int
|
Yadif mode (0=send_frame, 1=send_field, 2=send_frame_nospatial, 3=send_field_nospatial). Defaults to 0. |
yadif_parity |
int
|
Field parity (-1=auto, 0=tff, 1=bff). Defaults to -1. |
yadif_deint |
int
|
Deinterlace all frames (0) or flagged-only (1). Defaults to 0. |
default_timestamps |
list[float] | None
|
Used when |
validate_default_timestamps(value)
classmethod
Validate optional constructor-level timestamps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
list[float] | None
|
Candidate timestamps. |
required |
Returns:
| Type | Description |
|---|---|
list[float] | None
|
list[float] | None: Validated timestamps or None. |
validate_deinterlace_mode(value)
classmethod
Coerce bool / string deinterlace values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
Candidate mode. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
DeinterlaceMode |
DeinterlaceMode
|
Normalized mode. |
validate_output_format(value)
classmethod
Reject blank output format strings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
Candidate format. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Normalized upper-case format. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If blank. |
yadif_filter()
Build the shared yadif=mode:parity:deint filter description.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Yadif filter string consumed by FFmpeg |
FrameExtractionProcessConfig
Bases: ProcessorProcessConfig
Backend-agnostic per-invocation frame extraction config.
Attributes:
| Name | Type | Description |
|---|---|---|
timestamps |
list[float]
|
Required non-empty timestamps in seconds. |
output_format |
str | None
|
Optional format override. |
deinterlace |
DeinterlaceMode | None
|
Optional deinterlace override. |
yadif_mode |
int | None
|
Optional yadif mode override. |
yadif_parity |
int | None
|
Optional field-parity override. |
yadif_deint |
int | None
|
Optional flagged-only override. |
validate_optional_deinterlace_mode(value)
classmethod
Coerce optional bool / string deinterlace overrides.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
Candidate mode or None. |
required |
Returns:
| Type | Description |
|---|---|
DeinterlaceMode | None
|
DeinterlaceMode | None: Normalized mode or None. |
validate_optional_output_format(value)
classmethod
Normalize optional format override.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str | None
|
Candidate format. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: Upper-case format or None. |
validate_process_timestamps(value)
classmethod
Validate per-call timestamps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
list[float]
|
Candidate timestamps. |
required |
Returns:
| Type | Description |
|---|---|
list[float]
|
list[float]: Validated timestamps. |
FrameExtractionProcessor()
Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC
Family base for extracting image frames at timestamps.
Why use this base class?
- Portability: Swap FFmpeg vs GStreamer without changing call sites.
- Batch I/O: Extract many keyframes in one
processcall. - Separation: Keyframe planning stays in extractors; decode is here.
Usage
from gllm_multimodal.media_toolkit.processor.frame_extraction_processor import (
FrameExtractionProcessConfig,
FrameExtractionProcessor,
)
processor = FrameExtractionProcessor.build()
frames = await processor.process(
video_attachment,
process_config=FrameExtractionProcessConfig(timestamps=[1.5, 4.0]),
)
set_timestamps(timestamps)
abstractmethod
Configure default timestamps used when process_config is omitted.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamps
|
list[float]
|
Non-empty list of non-negative seconds. |
required |
stream_frames(attachment, sample_fps, deinterlace=None)
abstractmethod
Stream frames sampled at a fixed rate from a single decode pass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
sample_fps
|
float
|
Positive sampling rate in frames per second. |
required |
deinterlace
|
DeinterlaceMode | None
|
Deinterlace override. Defaults to None. |
None
|
Returns:
| Type | Description |
|---|---|
AsyncIterator[tuple[float, ndarray]]
|
AsyncIterator[tuple[float, np.ndarray]]: Timestamps and BGR |
FrameSamplingProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for resampling video attachments to a target frame rate.
This class serves as a unified entry point for frame sampling operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.
Why use this base class?
- Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
- Simplicity: No need to handle fallback logic or conditional imports yourself.
- Future-proofing: New backends can be added to the library without requiring changes to your application code.
Usage Example
from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import (
FrameSamplingProcessor,
GstFrameSamplingConfig,
)
from gllm_inference.schema import Attachment
# Instantiates the best available backend automatically
processor = FrameSamplingProcessor.build(
config=GstFrameSamplingConfig(default_target_fps=2)
)
attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import FrameSamplingProcessor
from gllm_inference.schema import Attachment
# Explicitly force the ffmpeg backend
processor = FrameSamplingProcessor.build(backend="ffmpeg")
attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)
GstAudioExtractionConfig
Bases: GstBaseConfig
Configuration for GstAudioExtractionProcessor.
Attributes:
| Name | Type | Description |
|---|---|---|
sample_rate |
int | None
|
Force output sample rate in Hz (e.g. |
channels |
int | None
|
Force output channel count (e.g. |
output_format |
str | None
|
Preferred container/extension to extract to
( |
audio_encoder |
str | None
|
Pin a specific GStreamer audio encoder element
(e.g. |
GstAudioExtractionProcessConfig
Bases: ProcessorProcessConfig
Per-invocation configuration for [process][gllm_multimodal.media_toolkit.media_toolkit.MediaToolkit.process].
Attributes:
| Name | Type | Description |
|---|---|---|
sample_rate |
int | None
|
Force output sample rate in Hz. |
channels |
int | None
|
Force output channel count. |
GstAudioExtractionProcessor(config=None)
Bases: BaseGstreamerProcessor[GstAudioExtractionConfig], AudioExtractionProcessor
Extracts the audio track from a video Attachment using GStreamer.
The processor demuxes the input video, re-encodes (or passes through) the
audio into the best available format, and returns the result as an
Attachment whose mime_type reflects the audio container.
If the video has no audio track the original Attachment is returned
unchanged, so callers do not need to handle None.
Attributes:
| Name | Type | Description |
|---|---|---|
audio_encoder_info |
EncoderFormatInfo | None
|
Selected audio encoder format info. |
config |
GstAudioExtractionConfig
|
Runtime configuration. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If GStreamer is unavailable or no suitable audio encoder is found. |
ValueError
|
If |
Example
processor = GstAudioExtractionProcessor(config={"output_format": "mp3"})
audio_attachment = await processor.process(video_attachment)
# audio_attachment.mime_type == "audio/mpeg"
# audio_attachment.filename == "audio_my_video.mp3"
Initialise GStreamer and select the audio encoder / output format.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | GstAudioExtractionConfig | None
|
Optional
configuration. A plain |
None
|
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If GStreamer is unavailable or no audio encoder is found. |
ValueError
|
If |
config_model()
classmethod
Return the stable configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstAudioExtractionConfig]
|
type[GstAudioExtractionConfig]: The configuration class used at |
type[GstAudioExtractionConfig]
|
construction time for stable (per-instance) settings. |
process(attachment, **kwargs)
async
Extract the audio track of one video attachment.
This public entrypoint keeps configuration ergonomics simple for callers:
stable defaults can be supplied in the constructor, while per-call
overrides (for sample rate/channels) can be passed through
process_config in kwargs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Input video attachment whose audio should be extracted. |
required |
**kwargs
|
Any
|
Additional runtime options, typically
|
{}
|
Notes
- Delegates shared mimetype validation and dispatch to
MediaToolkit.process.
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Extracted audio attachment, or the original attachment when |
Attachment
|
no audio stream is detected. |
Example
processor = GstAudioExtractionProcessor(
config={"output_format": "mp3"},
)
audio_attachment = await processor.process(
attachment=video_attachment,
process_config={"sample_rate": 16000, "channels": 1},
)
process_config_model()
classmethod
Return the per-invocation configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstAudioExtractionProcessConfig]
|
type[GstAudioExtractionProcessConfig]: The configuration class |
type[GstAudioExtractionProcessConfig]
|
accepted by |
type[GstAudioExtractionProcessConfig]
|
|
type[GstAudioExtractionProcessConfig]
|
for per-call overrides. |
GstBaseConfig
Bases: BaseModel
Minimal shared configuration for all GStreamer-based processors.
Attributes:
| Name | Type | Description |
|---|---|---|
timeout |
int
|
Maximum seconds a GStreamer pipeline may run before being forcibly terminated. Defaults to 300 seconds. |
audio_passthrough |
bool
|
When |
video_encoder |
str | None
|
Pin a specific GStreamer video encoder
element name (e.g. |
audio_encoder |
str | None
|
Pin a specific GStreamer audio encoder
element name (e.g. |
GstDeinterlaceConfig
Bases: DeinterlaceConfig, GstBaseConfig
GStreamer-specific stable config (inherits shared DeinterlaceConfig).
Yadif / encode defaults mirror
FFmpegDeinterlaceConfig
so both backends produce comparable progressive output.
Attributes:
| Name | Type | Description |
|---|---|---|
yadif_mode |
int
|
Yadif mode (0–3). Defaults to 0. |
yadif_parity |
int
|
Field parity (-1=auto, 0=tff, 1=bff). Defaults to -1. |
yadif_deint |
int
|
Deinterlace all frames (0) or flagged-only (1). Defaults to 0. |
crf |
int
|
Quantizer passed to the video encoder (CRF analogue). Defaults to 23. |
preset |
str
|
x264 speed preset name. Defaults to |
video_encoder |
str | None
|
Inherited from |
validate_preset(value)
classmethod
Reject blank preset strings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
Candidate preset name. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Validated preset name. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
yadif_filter()
Return an FFmpeg-style yadif filter label for metadata parity.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Filter description aligned with the FFmpeg backend. |
GstDeinterlaceProcessConfig
Bases: DeinterlaceProcessConfig
Per-invocation GStreamer overrides (inherits shared DeinterlaceProcessConfig).
GstDeinterlaceProcessor(config=None)
Bases: BaseGstreamerProcessor[GstDeinterlaceConfig], DeinterlaceProcessor
Deinterlace video with GStreamer yadif + x264enc.
Pipeline topology::
filesrc → decodebin → queue → videoconvert → yadif → videoconvert
→ x264enc → mp4mux → filesink
Audio pads are dropped when strip_audio is True.
Initialize the GStreamer deinterlace processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | DeinterlaceConfig | GstDeinterlaceConfig | None
|
Shared or GStreamer-specific config. Defaults to None. |
None
|
config_model()
classmethod
Return the stable configuration model.
process(attachment, **kwargs)
async
Deinterlace one video attachment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Input interlaced (or progressive) video. |
required |
**kwargs
|
Any
|
May include |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Progressive video attachment. |
process_config_model()
classmethod
Return the per-invocation configuration model.
GstFrameDecodeConfig
GstFrameDecodeProcessConfig
Bases: FrameDecodeProcessConfig
Per-invocation GStreamer decode overrides (inherits shared process config).
GstFrameDecodeProcessor(config=None)
Bases: BaseGstreamerProcessor[GstFrameDecodeConfig], FrameDecodeProcessor
Decode every video frame to PNG attachments with system GStreamer.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
GstFrameDecodeConfig
|
Runtime configuration. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If GStreamer is not available. |
Initialise without encoder selection (decode needs no encoder).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | FrameDecodeConfig | GstFrameDecodeConfig | None
|
Shared or GStreamer-specific config. Defaults to None. |
None
|
config_model()
classmethod
Return the stable configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstFrameDecodeConfig]
|
type[GstFrameDecodeConfig]: The stable configuration model. |
process(attachment, **kwargs)
async
Decode every frame of one video attachment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Input video attachment to decode. |
required |
**kwargs
|
Any
|
Runtime options, typically |
{}
|
Returns:
| Type | Description |
|---|---|
list[Attachment]
|
list[Attachment]: PNG image attachments in decode order. |
process_config_model()
classmethod
Return the per-invocation configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstFrameDecodeProcessConfig]
|
type[GstFrameDecodeProcessConfig]: The configuration class
accepted by
|
GstFrameExtractionConfig
Bases: FrameExtractionConfig, GstBaseConfig
GStreamer-specific stable config (inherits shared FrameExtractionConfig).
Shared yadif knobs are forwarded to
DeinterlaceProcessor
when deinterlace runs. Encode settings remain on
GstDeinterlaceConfig.
Attributes:
| Name | Type | Description |
|---|---|---|
output_format |
str
|
Image encode format (JPEG/PNG). Defaults to JPEG. |
deinterlace |
DeinterlaceMode
|
Whether to run the nested deinterlace
family before extracting frames. Defaults to |
yadif_mode |
int
|
Shared frame setting used by the FFmpeg backend; GStreamer uses its native deinterlace mode. Defaults to 0. |
yadif_parity |
int
|
Shared field parity used by the FFmpeg backend. GStreamer reads field order from caps. Defaults to -1. |
yadif_deint |
int
|
Shared flagged-only setting used by the FFmpeg backend. GStreamer reads interlace flags from caps. Defaults to 0. |
default_timestamps |
list[float] | None
|
Used when |
validate_supported_image_format(value)
classmethod
Reject still-image formats this backend cannot encode.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str
|
Candidate format (already upper-cased by the parent). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Validated format. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the format is not JPEG/JPG/PNG. |
GstFrameExtractionProcessConfig
Bases: FrameExtractionProcessConfig
Per-invocation GStreamer overrides (inherits shared process config).
validate_optional_supported_image_format(value)
classmethod
Reject optional still-image format overrides this backend cannot encode.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
str | None
|
Candidate format or None. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: Validated format or None. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the format is not JPEG/JPG/PNG. |
GstFrameExtractionProcessor(config=None)
Bases: BaseGstreamerProcessor[GstFrameExtractionConfig], FrameExtractionProcessor
Extract image frames with GStreamer.
When deinterlace is force or auto (and the source is interlaced),
frames are taken from the progressive output of
DeinterlaceProcessor
rather than inlining yadif in the still-frame pipeline.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
GstFrameExtractionConfig
|
Stable constructor configuration. |
Initialize the GStreamer frame extraction processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | FrameExtractionConfig | GstFrameExtractionConfig | None
|
Shared or GStreamer-specific config. Defaults to None. |
None
|
config_model()
classmethod
Return the stable configuration model.
extract_frames(video_path, timestamps, *, params=None)
Extract encoded image bytes at each timestamp from a video path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
video_path
|
str
|
Source video path. |
required |
timestamps
|
list[float]
|
Non-empty non-negative timestamps. |
required |
params
|
GstFrameExtractionConfig | None
|
Effective
options. Defaults to |
None
|
Returns:
| Type | Description |
|---|---|
list[bytes]
|
list[bytes]: Encoded frames in the same order as |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If any extraction fails. |
ValueError
|
If timestamps are invalid. |
process(attachment, **kwargs)
async
Extract frames from one video attachment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
**kwargs
|
Any
|
May include |
{}
|
Returns:
| Type | Description |
|---|---|
list[Attachment]
|
list[Attachment]: One image attachment per timestamp. |
process_config_model()
classmethod
Return the per-invocation configuration model.
set_timestamps(timestamps)
Update constructor-level default timestamps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamps
|
list[float]
|
Non-empty non-negative timestamps. |
required |
stream_frames(attachment, sample_fps, deinterlace=None)
async
Stream sampled BGR frames from a single GStreamer decode pass.
Frames are resampled by videorate, auto-rotated by videoflip and pulled from a
bounded appsink, so decoding pauses while the consumer is busy. Closing the
generator early stops the pipeline.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source video. |
required |
sample_fps
|
float
|
Positive sampling rate in frames per second. |
required |
deinterlace
|
DeinterlaceMode | None
|
Deinterlace override. Defaults to None. |
None
|
Yields:
| Type | Description |
|---|---|
AsyncIterator[tuple[float, ndarray]]
|
tuple[float, np.ndarray]: Timestamp and BGR |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
RuntimeError
|
If decoding fails, the source has no video stream, or the pipeline stalls. |
GstFrameSamplingConfig
Bases: GstBaseConfig
Stable configuration for GstFrameSamplingProcessor.
Attributes:
| Name | Type | Description |
|---|---|---|
min_size_mb |
int
|
Minimum video size in MB to process. |
default_target_fps |
int
|
Default target FPS used when
|
validate_default_target_fps(default_target_fps)
classmethod
Validate that default_target_fps is positive.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
default_target_fps
|
int
|
The candidate default FPS to validate. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
int |
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
GstFrameSamplingProcessConfig
Bases: ProcessorProcessConfig
Per-invocation configuration for [process][gllm_multimodal.media_toolkit.media_toolkit.MediaToolkit.process].
Attributes:
| Name | Type | Description |
|---|---|---|
target_fps |
int
|
Target frames-per-second for the output video. |
validate_target_fps(target_fps)
classmethod
Validate that target_fps is positive.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
target_fps
|
int
|
The candidate target FPS to validate. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
int |
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
GstFrameSamplingProcessor(config=None)
Bases: BaseGstreamerProcessor[GstFrameSamplingConfig], FrameSamplingProcessor
Resamples a video attachment to a target frame-rate using GStreamer.
The processor is format-agnostic: it handles MP4, MKV, MOV, and similar containers and preserves any audio track (passthrough by default).
Attributes:
| Name | Type | Description |
|---|---|---|
config |
GstFrameSamplingConfig
|
Runtime configuration. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If GStreamer is not available or no suitable encoder is found. |
Initialise and select encoders.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | GstFrameSamplingConfig | None
|
Optional stable configuration. |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
config_model()
classmethod
Return the stable configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstFrameSamplingConfig]
|
type[GstFrameSamplingConfig]: The stable configuration model for this processor. |
process(attachment, **kwargs)
async
Resample one video attachment to the configured or requested FPS.
The method preserves the regular media-toolkit process contract while
exposing this concrete processor's runtime semantics in API docs:
constructor defaults establish baseline FPS behavior, and per-call
overrides can be supplied through process_config.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Input video attachment to resample. |
required |
**kwargs
|
Any
|
Runtime options, typically |
{}
|
Notes
- Delegates shared mimetype validation and dispatch to
MediaToolkit.process.
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Resampled video attachment, or the original attachment |
Attachment
|
when the input is skipped by validation rules. |
Example
processor = GstFrameSamplingProcessor(
config={"default_target_fps": 2},
)
sampled_attachment = await processor.process(
attachment=video_attachment,
process_config={"target_fps": 1},
)
process_config_model()
classmethod
Return the per-invocation configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstFrameSamplingProcessConfig]
|
type[GstFrameSamplingProcessConfig]: The configuration class |
type[GstFrameSamplingProcessConfig]
|
accepted by |
type[GstFrameSamplingProcessConfig]
|
|
type[GstFrameSamplingProcessConfig]
|
for per-call overrides. |
GstVideoClipConfig
Bases: GstBaseConfig
Stable configuration for GstVideoClipProcessor.
Set once at construction. Controls encoder selection, timeout, and other processor behaviour that does not change per attachment.
Attributes:
| Name | Type | Description |
|---|---|---|
default_windows |
list[tuple[float, float]] | None
|
Default clipping
windows used when |
validate_default_windows(default_windows)
classmethod
Validate the constructor-level default windows, if any.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
default_windows
|
list[tuple[float, float]] | None
|
The candidate default windows to validate. |
required |
Returns:
| Type | Description |
|---|---|
list[tuple[float, float]] | None
|
list[tuple[float, float]] | None: |
list[tuple[float, float]] | None
|
when |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
GstVideoClipProcessConfig
Bases: ProcessorProcessConfig
Per-invocation configuration for [process][gllm_multimodal.media_toolkit.media_toolkit.MediaToolkit.process].
Attributes:
| Name | Type | Description |
|---|---|---|
windows |
list[tuple[float, float]]
|
Ordered |
validate_windows(windows)
classmethod
Validate that every clipping window is non-empty and well ordered.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
windows
|
list[tuple[float, float]]
|
The windows to validate. |
required |
Returns:
| Type | Description |
|---|---|
list[tuple[float, float]]
|
list[tuple[float, float]]: The validated windows. |
GstVideoClipProcessor(config=None)
Bases: BaseGstreamerProcessor[GstVideoClipConfig], VideoClipProcessor
Clips a video attachment to one or more [start_time, end_time] windows.
Stable settings belong in GstVideoClipConfig (constructor).
Per-call clip windows belong in GstVideoClipProcessConfig
(process(..., process_config=...)).
Initialise the clip processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
dict[str, Any] | GstVideoClipConfig | None
|
Optional stable
configuration. |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
config_model()
classmethod
Return the stable configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstVideoClipConfig]
|
type[GstVideoClipConfig]: The stable configuration model for this processor. |
process(attachment, **kwargs)
async
Clip one video attachment using resolved time windows.
This method is the canonical caller-facing API for clipping. It accepts
stable default windows set at construction time and also supports
per-call window overrides through process_config.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Input video attachment to clip. |
required |
**kwargs
|
Any
|
Runtime options, typically |
{}
|
Notes
- Delegates shared mimetype validation and dispatch to
MediaToolkit.process.
Returns:
| Type | Description |
|---|---|
Attachment | list[Attachment]
|
Attachment | list[Attachment]: Single clipped attachment when exactly |
Attachment | list[Attachment]
|
one window is used, otherwise a list of clipped attachments in window order. |
Example
processor = GstVideoClipProcessor(
config={"default_windows": [(0.0, 10.0)]},
)
clips = await processor.process(
attachment=video_attachment,
process_config={
"windows": [(5.0, 12.5), (30.0, 45.0)],
},
)
process_config_model()
classmethod
Return the per-invocation configuration model for this processor.
Returns:
| Type | Description |
|---|---|
type[GstVideoClipProcessConfig]
|
type[GstVideoClipProcessConfig]: The per-invocation configuration model for this processor. |
set_windows(windows)
Update the constructor-level default windows.
Mutates config so process falls back to windows when
no process_config is supplied. Prefer passing
GstVideoClipProcessConfig to process for per-call
overrides.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
windows
|
list[tuple[float, float]]
|
The windows to set. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
|
None
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
MediaToolkit()
Bases: ABC, Generic[T_in, T_out]
Base abstraction for all media toolkit processing components.
This class provides the shared lifecycle and registry behavior used by both: - concrete leaf processors (e.g. backend-specific audio/video processors), and - composite components (e.g. segmenters, keyframe extractors) that orchestrate nested processors.
Key responsibilities:
- auto-register subclasses by class name for class-name-based construction via
build;
- provide consistent input validation against supported_mimetypes;
- define async processing contracts through process and
process_batch.
Contributor guidance:
- inherit this class directly for concrete processors with custom behavior;
- inherit BackendSelectableProcessor when one logical processor family maps
to multiple backend implementations;
- inherit composite bases (e.g. BaseSegmenter) for orchestration-style
components.
Example
Building a processor by class name
from gllm_multimodal.media_toolkit.media_toolkit import MediaToolkit
processor = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
result = await processor.process(video_attachment)
Checking mimetype support
if processor.is_supported(attachment):
result = await processor.process(attachment)
Listing registered processors
print(list(MediaToolkit.registry.keys()))
# ['GstAudioExtractionProcessor', 'GstVideoClipProcessor', ...]
Initialize processor logging.
name
property
Return a stable component name for optional plan metadata.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Component class name. |
registry = {}
class-attribute
Global class-name registry used by build.
Maps each registered subclass name to its concrete class type, enabling
string-based construction such as MediaToolkit.build("AudioExtractionProcessor").
supported_mimetypes = ['*/*']
class-attribute
MIME types this processor accepts (supports wildcards, e.g. 'video/*').
Defaults to ['*/*'] (accept all). Override as a class attribute in subclasses.
__init_subclass__(**kwargs)
Register every concrete subclass into the global registry.
This hook is triggered automatically by Python whenever a class inherits
from MediaToolkit (directly or indirectly). Registration happens at
class definition/import time, so classes become immediately discoverable
by build
without manual setup.
Registration key
- The subclass'
__name__(e.g."AudioExtractionProcessor"). - The value stored is the subclass type itself.
Why uniqueness is enforced
build(class_name=...)uses this registry for class resolution.- Duplicate class names would silently shadow earlier classes and could route builds to unintended implementations.
- To prevent that ambiguity, duplicate keys raise
TypeError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
**kwargs
|
Any
|
Extra class declaration keyword arguments forwarded to
parent |
{}
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If a subclass with the same class name is already registered, preventing silent dispatch to the wrong implementation. |
available_backends_for(class_name)
classmethod
Return backend keys registered for a processor family class name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
class_name
|
str
|
Registered processor family class name. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Available backend keys. Empty when the class is unknown or not a backend-selectable family base. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the class name is unknown. |
Example
backends = MediaToolkit.available_backends_for("VideoClipProcessor")
print(backends)
["gstreamer", "ffmpeg"]
build(class_name, backend=None, **kwargs)
classmethod
Build a processor by class name.
Family abstract classes (e.g. AudioExtractionProcessor) resolve a concrete
backend implementation via backend. Composite components (segmenters,
keyframe extractors) store backend on the instance for nested resolution.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
class_name
|
str
|
Registered subclass name. |
required |
backend
|
str | MediaBackend | None
|
Backend key for family classes, or nested processor preference for composite instances. Defaults to None. |
None
|
**kwargs
|
Any
|
Constructor kwargs passed to the processor class. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
MediaToolkit |
MediaToolkit
|
Instantiated processor. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the class name is unknown. |
Example
proc = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
seg = MediaToolkit.build(
"FixedDurationSegmenter",
backend="gstreamer",
config={"segment_durations": [2.0]},
)
print(proc)
print(seg)
<GstAudioExtractionProcessor instance>
<FixedDurationSegmenter instance>
build_from_registry(backend=None, **kwargs)
classmethod
Instantiate this registered class.
Subclasses override this hook to customize registry-based construction
(e.g. backend-selectable families resolve a concrete backend; composites
store backend for nested processor resolution).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
backend
|
str | MediaBackend | None
|
Backend key forwarded to subclass overrides. Ignored by the base implementation. |
None
|
**kwargs
|
Any
|
Constructor kwargs passed to the processor class. |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
MediaToolkit |
MediaToolkit
|
Instantiated processor. |
Example
# Called indirectly by MediaToolkit.build(...)
processor = SomeRegisteredProcessor.build_from_registry(custom_flag=True)
print(processor)
<SomeRegisteredProcessor instance>
is_supported(attachment)
Return whether the attachment's mimetype is accepted by this processor.
Callers can use this to check compatibility before calling
process or process_batch, avoiding a ValueError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
The attachment to check. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
|
Example
ok = processor.is_supported(Attachment(mime_type="video/mp4"))
print(ok)
True
list_available_backends()
classmethod
Return backend keys when this class supports backend selection.
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Available backend keys. Empty for classes that are not backend-selectable family bases. |
Example
Base classes are not backend-selectable.
print(MediaToolkit.list_available_backends())
[]
Family classes expose registered backend keys.
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
backends = VideoClipProcessor.list_available_backends()
print(backends)
["gstreamer", "ffmpeg", "moviepy"] # depends on registered backends
process(attachment, **kwargs)
async
Process a single attachment (or perform an aggregation on a list) and return the result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
T_in
|
The attachment or list of attachments to process. |
required |
**kwargs
|
Any
|
Additional keyword arguments forwarded to |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
T_out |
T_out
|
The result of the processing. |
Example
result = await processor.process(attachment)
print(result)
<processed attachment or transformed output>
process_batch(attachments, **kwargs)
async
Process a batch of attachments.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachments
|
list[T_in]
|
The batch of attachments to process. |
required |
**kwargs
|
Any
|
Additional keyword arguments forwarded to |
{}
|
Returns:
| Type | Description |
|---|---|
list[T_out]
|
list[T_out]: The result of the batch processing. |
Example
results = await processor.process_batch([attachment_1, attachment_2])
print(results)
[<result_1>, <result_2>]
ProcessorProcessConfig
Bases: BaseModel
Marker base class for per-invocation processor configuration.
Concrete processors define subclasses with the parameters that may change
on every process() call. Stable processor settings belong in each
processor's *Config model passed at construction time.
VideoClipProcessor()
Bases: BackendSelectableProcessor[Attachment, Attachment], ABC
Family base for clipping video attachments to time windows.
This class serves as a unified entry point for video clipping operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.
Why use this base class?
- Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
- Simplicity: No need to handle fallback logic or conditional imports yourself.
- Future-proofing: New backends can be added to the library without requiring changes to your application code.
Usage Example
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment
# Instantiates the best available backend automatically
processor = VideoClipProcessor.build()
# Set the target clipping window (start_time, end_time) in seconds
processor.set_windows([(10.0, 20.5)])
attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment
# Explicitly force the ffmpeg backend
processor = VideoClipProcessor.build(backend="ffmpeg")
processor.set_windows([(10.0, 20.5)])
attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment
# Explicitly force the moviepy backend
processor = VideoClipProcessor.build(backend="moviepy")
processor.set_windows([(10.0, 20.5)])
attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
set_windows(windows)
abstractmethod
Configure one or more [start, end] clipping windows for the next call.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
windows
|
list[tuple[float, float]]
|
List of (start, end) tuples in seconds. |
required |