Skip to content

Overview

Media toolkit processor public API and loading entrypoint.

This package exposes processor families and shared backend base classes used by builder/factory flows. Importing from here ensures family classes and concrete backend implementations are loaded so subclass registration is populated in MediaToolkit.registry and each family's backend map.

Design overview
  • Family package layout: <family>/base.py for the abstract family class and <family>/<backend>_backend.py for concrete implementations.
  • Construction: MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer").
  • Composite components (segmenters/keyframe extractors): cache nested processors by class name + backend and can switch backend by changing component.backend before calling get_processor(...).
Contributor quick-start for a new processor family
  1. Create processor/<new_family>/base.py subclassing BackendSelectableProcessor.
  2. Add backend implementations declaring BACKEND (+ IS_DEFAULT if needed).
  3. Re-export in processor/<new_family>/__init__.py and this module.
  4. Add tests for family registration and factory creation.

AudioExtractionProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for extracting audio tracks from video attachments.

This class serves as a unified entry point for audio extraction operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.

Why use this base class?

  • Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
  • Simplicity: No need to handle fallback logic or conditional imports yourself.
  • Future-proofing: New backends can be added to the library without requiring changes to your application code.

Usage Example

from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment

# Instantiates the best available backend automatically
processor = AudioExtractionProcessor.build()

attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment

# Explicitly force the ffmpeg backend
processor = AudioExtractionProcessor.build(backend="ffmpeg")

attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.audio_extraction_processor import AudioExtractionProcessor
from gllm_inference.schema import Attachment

# Explicitly force the moviepy backend
processor = AudioExtractionProcessor.build(backend="moviepy")

attachment = Attachment(url="file:///path/to/video.mp4")
audio_attachment = await processor.process(attachment)

BackendSelectableProcessor()

Bases: MediaToolkit[T_in, T_out], ABC

Abstract base for a processor family with pluggable backends.

Each family owns its backend registry. Concrete backend classes are auto-registered via __init_subclass__ by declaring:

  • BACKEND: backend key, e.g. gstreamer or ffmpeg
  • IS_DEFAULT: whether this backend is the family default
Why this exists

It lets callers construct by stable family class name while deferring runtime selection of backend implementation.

Minimal contributor pattern
# base.py
class ImageTilingProcessor(BackendSelectableProcessor):
    pass

# pil_backend.py
class PilImageTilingProcessor(ImageTilingProcessor):
    BACKEND = "pil"

# cv2_backend.py
class Cv2ImageTilingProcessor(ImageTilingProcessor):
    BACKEND = "cv2"
Then callers can use

MediaToolkit.build("ImageTilingProcessor", backend="pil").

__init_subclass__(**kwargs)

Automatically register backend classes into their family registry.

This hook runs at class definition/import time and maintains per-family backend mappings used by backend-selectable builds.

Registration flow
  1. If cls is the family base itself (e.g. VideoClipProcessor), reset _backends and _default_backend for that family.
  2. If cls is abstract, skip registration.
  3. Otherwise treat cls as a concrete backend implementation:
  4. require class variable BACKEND (e.g. "gstreamer"),
  5. reject duplicate backend keys within the same family,
  6. add mapping family_base._backends[BACKEND] = cls.
Default backend resolution
  1. If IS_DEFAULT=True on a concrete class, that backend becomes the family's default, even if a first-registered fallback already exists.
  2. If no explicit default exists yet, the first registered backend is used as fallback default.
  3. Multiple IS_DEFAULT=True declarations in one family raise ValueError.

Parameters:

Name Type Description Default
**kwargs Any

Extra class declaration keyword arguments forwarded to parent __init_subclass__ implementations.

{}

Raises:

Type Description
TypeError

If a concrete backend class omits BACKEND.

ValueError

If duplicate backend keys are registered in one family.

ValueError

If more than one backend declares IS_DEFAULT=True.

build(backend=None, **kwargs) classmethod

Build a backend implementation for this processor family.

Parameters:

Name Type Description Default
backend str | None

Backend key. Uses the family default when omitted.

None
**kwargs Any

Constructor kwargs forwarded to the backend class.

{}

Returns:

Name Type Description
BackendSelectableProcessor BackendSelectableProcessor

Instantiated backend processor.

Raises:

Type Description
ValueError

If the backend is unknown or no default is configured.

Example

For a family class AudioExtractionProcessor, AudioExtractionProcessor.build(backend="gstreamer") returns the registered GStreamer implementation class instance.

<GstAudioExtractionProcessor instance>

build_from_registry(backend=None, **kwargs) classmethod

Build a family backend or instantiate a concrete backend class.

Parameters:

Name Type Description Default
backend str | MediaBackend | None

Backend key for family resolution. Defaults to None.

None
**kwargs Any

Constructor kwargs forwarded to the backend class.

{}

Returns:

Name Type Description
BackendSelectableProcessor BackendSelectableProcessor

Instantiated processor.

Example

processor = AudioExtractionProcessor.build_from_registry(backend="gstreamer")
<GstAudioExtractionProcessor instance>

get_install_hint(backend=None) classmethod

Return the installation hint for a backend.

Reads INSTALL_HINT from the registered backend class. If the backend is not registered or carries no hint, returns a generic fallback.

Parameters:

Name Type Description Default
backend str | None

Backend key to look up. Defaults to _default_backend when None.

None

Returns:

Name Type Description
str str

Human-readable install hint for the backend, or a generic fallback.

is_family_base(processor_cls) classmethod

Return whether processor_cls is a family abstract base.

Parameters:

Name Type Description Default
processor_cls type[BackendSelectableProcessor]

Candidate processor class.

required

Returns:

Name Type Description
bool bool

True when the class directly subclasses BackendSelectableProcessor.

Example

assert BackendSelectableProcessor.is_family_base(AudioExtractionProcessor) is True
True

list_available_backends() classmethod

Return registered backend keys for backend-selectable family bases.

Returns:

Type Description
list[str]

list[str]: Backend keys available for build. Empty for concrete backend implementations and non-family classes.

Example

AudioExtractionProcessor.list_available_backends()
# e.g. ["gstreamer", "ffmpeg", "moviepy"]
["gstreamer", "ffmpeg", "moviepy"]

list_backends() classmethod

Return registered backend keys for this processor family.

Returns:

Type Description
list[str]

list[str]: Backend keys available for build.

Example

backends = AudioExtractionProcessor.list_backends()
["gstreamer", "ffmpeg", "moviepy"]

BaseFFmpegProcessor(config=None)

Bases: MediaToolkit[Attachment, Attachment], Generic[TConfig]

Shared helpers for processors that shell out to the ffmpeg binary.

Unlike BaseGstreamerProcessor, this base does not require ffmpeg at construction time — availability is checked when a command is run. That keeps import / build paths light when ffmpeg is only needed for optional I/O.

Attributes:

Name Type Description
INSTALL_HINT str

Human-readable install guidance for missing ffmpeg.

config TConfig

Coerced stable configuration from config_model().

Initialize the FFmpeg helper base and coerce constructor config.

Parameters:

Name Type Description Default
config dict[str, Any] | BaseModel | None

Stable config dict, family/backend model, or None for defaults.

None

config_model() classmethod

Return the stable configuration model for this processor.

Concrete backends must override with their FFmpeg config type.

Returns:

Type Description
type[TConfig]

type[TConfig]: The configuration class used at construction time.

BaseGstreamerProcessor(config=None, *, enable_video=True, enable_audio=True)

Bases: MediaToolkit[Attachment, Attachment], Generic[TConfig]

Abstract base class for GStreamer-powered processors.

Subclasses must implement only _execute_pipeline. All common concerns — availability checks, Conda environment configuration, GStreamer initialisation, encoder/element selection, the standard EOS/error bus loop, and temporary-file cleanup — are handled here.

Parameterise with the concrete config model, e.g. BaseGstreamerProcessor[GstFrameDecodeConfig].

Typical subclass skeleton::

class MyGstProcessor(BaseGstreamerProcessor[GstBaseConfig]):
    def __init__(self, my_param, config=None):
        super().__init__(config=config)
        self.my_param = my_param

    async def _process(self, attachment: Attachment) -> Attachment:
        return await self._process_single(attachment)

    async def _execute_pipeline(self, input_path, output_path):
        # build & run YOUR GStreamer pipeline here
        ...

Attributes:

Name Type Description
logger

Logger bound to the concrete subclass name.

config TConfig

Runtime configuration for this processor family.

video_encoder_info EncoderFormatInfo | None

Selected video encoder format info, or None when video is disabled.

audio_encoder_info EncoderFormatInfo | None

Selected audio encoder format info, or None when audio is disabled.

Verify GStreamer availability, configure the environment, and select encoders.

Parameters:

Name Type Description Default
config dict[str, Any] | GstBaseConfig | BaseModel | None

Optional configuration. A plain dict or shared family BaseModel is coerced into config_model. None uses all defaults of that model.

None
enable_video bool

When True, select a video encoder. Defaults to True.

True
enable_audio bool

When True, select an audio encoder. Defaults to True.

True

Raises:

Type Description
RuntimeError

If GStreamer is not installed or cannot be initialised.

RuntimeError

If no suitable video encoder is found in the registry.

TypeError

If config is neither dict, BaseModel, nor None.

config_model() classmethod

Return the stable configuration model for this processor.

Returns:

Type Description
type[TConfig]

type[TConfig]: The configuration class used at construction time.

DeinterlaceConfig

Bases: BaseModel

Backend-agnostic stable configuration for deinterlace processors.

Only knobs that apply across FFmpeg / GStreamer (and future engines) belong here. Engine-specific fields (yadif mode, x264 CRF, GST element props, …) live on backend subclasses that inherit this model.

Attributes:

Name Type Description
strip_audio bool

Drop audio when producing progressive output. Defaults to True.

DeinterlaceMode

Bases: StrEnum

How frame extraction applies deinterlace (e.g. FFmpeg yadif).

Attributes:

Name Type Description
OFF

Never deinterlace.

FORCE

Always deinterlace (ignore field-order metadata).

AUTO

Deinterlace only when ffprobe reports an interlaced field order.

DeinterlaceProcessConfig

Bases: ProcessorProcessConfig

Backend-agnostic per-invocation overrides for deinterlace.

None fields fall back to the stable constructor config. Backend process-config models inherit this class and add engine-specific overrides.

DeinterlaceProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for deinterlacing video attachments.

Why use this base class?

  • Portability: Swap FFmpeg vs GStreamer without changing call sites.
  • I/O boundary: Deinterlace is media transform I/O, not a segmenter algorithm.
  • Tunable: Shared knobs on DeinterlaceConfig; backends extend for engine-specific settings via inheritance.

Usage

from gllm_multimodal.media_toolkit.processor.deinterlace_processor import (
    DeinterlaceConfig,
    DeinterlaceProcessor,
)

processor = DeinterlaceProcessor.build(
    backend="ffmpeg",
    config=DeinterlaceConfig(strip_audio=True),
)
progressive = await processor.process(video_attachment)

is_interlaced(attachment) staticmethod

Return whether ffprobe reports an interlaced field order for an attachment.

Writes attachment.data to a temporary file, probes field_order, then cleans up. Returns False when ffprobe is unavailable, the probe fails, or the field order is progressive or unknown (fail-open: unknown is treated as progressive so callers skip the deinterlace pass; a warning is logged whenever the probe cannot determine interlacing).

Callers that need fail-safe behavior (deinterlace when unknown) should check ffprobe availability separately instead of relying on this gate.

Parameters:

Name Type Description Default
attachment Attachment

Video attachment to probe.

required

Returns:

Name Type Description
bool bool

True when field_order is one of tt, bb, tb, or bt.

is_interlaced_path(video_path) staticmethod

Return whether ffprobe reports an interlaced field order for a local path.

Public alias of is_interlaced_path kept on the family base so frame-extraction and other families do not reach into a private member.

Parameters:

Name Type Description Default
video_path str

Path to a local video file.

required

Returns:

Name Type Description
bool bool

True when field_order is interlaced.

FFmpegBaseConfig

Bases: BaseModel

Minimal shared configuration for FFmpeg CLI processors.

Family backends extend this (or their own shared family config) with engine-specific fields. Kept intentionally thin — FFmpeg leaves share process helpers more than stable knobs.

FFmpegDeinterlaceConfig

Bases: DeinterlaceConfig

FFmpeg-specific stable config (inherits shared DeinterlaceConfig).

Attributes:

Name Type Description
yadif_mode int

Yadif mode (0=send_frame, 1=send_field, 2=send_frame_nospatial, 3=send_field_nospatial). Defaults to 0.

yadif_parity int

Field parity (-1=auto, 0=tff, 1=bff). Defaults to -1.

yadif_deint int

Deinterlace all frames (0) or only flagged (1). Defaults to 0.

video_codec str

Video encoder name. Defaults to libx264.

preset str

x264 preset. Defaults to ultrafast.

crf int

Constant rate factor (0–51). Defaults to 23.

filter_override str | None

Full -vf string; when set, yadif_* fields are ignored. Defaults to None.

validate_non_empty(value) classmethod

Reject blank codec/preset strings.

Parameters:

Name Type Description Default
value str

Candidate string.

required

Returns:

Name Type Description
str str

Validated string.

Raises:

Type Description
ValueError

If value is blank.

yadif_filter()

Build the -vf filter string.

Returns:

Name Type Description
str str

Filter graph for FFmpeg -vf.

FFmpegDeinterlaceProcessConfig

Bases: DeinterlaceProcessConfig

Per-invocation FFmpeg overrides (inherits shared DeinterlaceProcessConfig).

Any field left as None falls back to the stable constructor config.

FFmpegDeinterlaceProcessor(config=None)

Bases: BaseFFmpegProcessor[FFmpegDeinterlaceConfig], DeinterlaceProcessor

Deinterlace video with FFmpeg yadif.

Attributes:

Name Type Description
config FFmpegDeinterlaceConfig

Stable constructor configuration.

Initialize the FFmpeg deinterlace processor.

Parameters:

Name Type Description Default
config dict[str, Any] | DeinterlaceConfig | FFmpegDeinterlaceConfig | None

Shared DeinterlaceConfig, FFmpeg-specific config, or dict. Shared-only configs are promoted with FFmpeg defaults for engine fields. Defaults to None.

None

config_model() classmethod

Return the stable configuration model.

deinterlace_video(video_path, output_path=None, *, params=None)

Deinterlace a video path with FFmpeg yadif into a progressive file.

Parameters:

Name Type Description Default
video_path str

Source video path.

required
output_path str | None

Destination path. When None, a temporary .mp4 path is created. Defaults to None.

None
params FFmpegDeinterlaceConfig | None

Effective options for this call. Defaults to self.config.

None

Returns:

Name Type Description
str str

Path to the progressive output video.

Raises:

Type Description
FileNotFoundError

If ffmpeg is not on PATH.

RuntimeError

If FFmpeg deinterlace fails.

process(attachment, **kwargs) async

Deinterlace one video attachment.

Parameters:

Name Type Description Default
attachment Attachment

Input interlaced (or progressive) video.

required
**kwargs Any

May include process_config overrides.

{}

Returns:

Name Type Description
Attachment Attachment

Progressive video attachment.

process_config_model() classmethod

Return the per-invocation configuration model.

FFmpegFrameDecodeConfig

Bases: FrameDecodeFieldsMixin, FFmpegBaseConfig

FFmpeg-specific stable configuration for dense frame decoding.

FFmpegFrameDecodeProcessConfig

Bases: FrameDecodeProcessConfig

Per-invocation FFmpeg overrides (inherits shared process config).

FFmpegFrameDecodeProcessor(config=None)

Bases: BaseFFmpegProcessor[FFmpegFrameDecodeConfig], FrameDecodeProcessor

Decode every video frame to PNG attachments with the FFmpeg CLI.

Attributes:

Name Type Description
config FFmpegFrameDecodeConfig

Stable constructor configuration.

Initialize the FFmpeg frame decode processor.

Parameters:

Name Type Description Default
config dict[str, Any] | FrameDecodeConfig | FFmpegFrameDecodeConfig | None

Shared or FFmpeg-specific config. Defaults to None.

None

build_filter(*, sample_fps, target_width)

Build the -vf filter chain for the requested knobs.

Parameters:

Name Type Description Default
sample_fps int | None

Downsample rate, if any.

required
target_width int | None

Downscale width, if any.

required

Returns:

Type Description
str | None

str | None: Comma-joined filter chain, or None for native decode.

config_model() classmethod

Return the stable configuration model.

process(attachment, **kwargs) async

Decode every frame of one video attachment.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
**kwargs Any

May include process_config.

{}

Returns:

Type Description
list[Attachment]

list[Attachment]: PNG image attachments in decode order.

process_config_model() classmethod

Return the per-invocation configuration model.

FFmpegFrameExtractionConfig

Bases: FrameExtractionConfig

FFmpeg-specific stable config (inherits shared FrameExtractionConfig).

Yadif fields apply when deinterlace is force or auto (and the probe selects yadif for auto).

Attributes:

Name Type Description
yadif_mode int

Yadif mode. Defaults to 0.

yadif_parity int

Field parity (-1=auto). Defaults to -1.

yadif_deint int

Deinterlace all (0) or flagged-only (1). Defaults to 0.

filter_override str | None

Full -vf string when set. Defaults to None.

yadif_filter()

Build the -vf filter string for deinterlaced extract.

Returns:

Name Type Description
str str

Filter graph for FFmpeg -vf.

FFmpegFrameExtractionProcessConfig

Bases: FrameExtractionProcessConfig

Per-invocation FFmpeg overrides (inherits shared process config).

FFmpegFrameExtractionProcessor(config=None)

Bases: BaseFFmpegProcessor[FFmpegFrameExtractionConfig], FrameExtractionProcessor

Extract image frames with FFmpeg (optional yadif).

Attributes:

Name Type Description
config FFmpegFrameExtractionConfig

Stable constructor configuration.

Initialize the FFmpeg frame extraction processor.

Parameters:

Name Type Description Default
config dict[str, Any] | FrameExtractionConfig | FFmpegFrameExtractionConfig | None

Shared or FFmpeg-specific config. Defaults to None.

None

config_model() classmethod

Return the stable configuration model.

extract_frames(video_path, timestamps, *, params=None)

Extract encoded image bytes at each timestamp from a video path.

Parameters:

Name Type Description Default
video_path str

Source video path.

required
timestamps list[float]

Non-empty non-negative timestamps.

required
params FFmpegFrameExtractionConfig | None

Effective options. Defaults to self.config.

None

Returns:

Type Description
list[bytes]

list[bytes]: Encoded frames in the same order as timestamps.

Raises:

Type Description
FileNotFoundError

If ffmpeg is unavailable.

RuntimeError

If any extraction fails.

ValueError

If timestamps are invalid.

process(attachment, **kwargs) async

Extract frames from one video attachment.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
**kwargs Any

May include process_config.

{}

Returns:

Type Description
list[Attachment]

list[Attachment]: One image attachment per timestamp.

process_config_model() classmethod

Return the per-invocation configuration model.

set_timestamps(timestamps)

Update constructor-level default timestamps.

Parameters:

Name Type Description Default
timestamps list[float]

Non-empty non-negative timestamps.

required

stream_frames(attachment, sample_fps, deinterlace=None) async

Stream sampled BGR frames from a single FFmpeg process.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
sample_fps float

Positive sampling rate in frames per second.

required
deinterlace DeinterlaceMode | None

Deinterlace override. Defaults to None.

None

Yields:

Type Description
AsyncIterator[tuple[float, ndarray]]

tuple[float, np.ndarray]: Timestamp and BGR uint8 frame at source size.

Raises:

Type Description
ValueError

If sample_fps is not finite and positive.

FileNotFoundError

If ffmpeg or ffprobe is unavailable.

RuntimeError

If probing or decoding fails.

FrameDecodeConfig

Bases: FrameDecodeFieldsMixin

Backend-agnostic stable configuration for dense frame decoding.

FrameDecodeProcessConfig

Bases: ProcessorProcessConfig

Backend-agnostic per-invocation frame decode config.

Attributes:

Name Type Description
sample_fps int | None

Optional rate override.

target_width int | None

Optional width override.

FrameDecodeProcessor()

Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC

Family base for dense video-frame decoding.

Why use this base class?

  • Portability: Swap FFmpeg vs GStreamer without changing call sites.
  • FIPS choice: backend="gstreamer" decodes with system plugins (no bundled-FFmpeg wheels); backend="ffmpeg" shells out to the system ffmpeg binary.
  • Separation: Frame decoding stays here; frame scoring (shot detection, keyframes) lives in segmenters/extractors.

Usage

from gllm_multimodal.media_toolkit.processor.frame_decode_processor import (
    FrameDecodeProcessor,
)

processor = FrameDecodeProcessor.build(backend="gstreamer")
frames = await processor.process(video_attachment)

frame_metadata(frame_index, fps)

Build the metadata dict attached to every decoded frame.

Parameters:

Name Type Description Default
frame_index int

Zero-based position in decode order.

required
fps float | None

Effective sampling rate, if known.

required

Returns:

Type Description
dict[str, Any]

dict[str, Any]: frame_index / timestamp / fps / frame_decode_backend mapping.

iter_rgb_frames_sync(attachment, *, sample_fps=None, target_width=None, as_rgb=True)

Yield decoded frames after the backend writes the full PNG sequence.

The decoder subprocess/pipeline completes first, so peak temp-disk usage is every sampled PNG at once. After that, this generator opens each file in order, yields the payload, and deletes the PNG so Python RAM stays O(1) in frames. Prefer this over process when callers only need a scored stream and can tolerate the peak-disk cost.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
sample_fps int | None

Override downsample rate. Defaults to the constructor config.

None
target_width int | None

Override downscale width. Defaults to the constructor config.

None
as_rgb bool

When True (default), yield RGB arrays. When False, yield raw PNG bytes (used by process).

True

Yields:

Type Description
tuple[Any, dict[str, Any]]

tuple[Any, dict[str, Any]]: (rgb_or_png_bytes, frame_metadata).

Raises:

Type Description
RuntimeError

If decoding fails or yields no frames.

FileNotFoundError

If a required decoder binary is missing.

FrameExtractionConfig

Bases: BaseModel

Backend-agnostic stable configuration for frame extraction.

Yadif knobs are shared so FFmpeg can inline -vf yadif and GStreamer can forward the same values to DeinterlaceProcessor.

Attributes:

Name Type Description
output_format str

Image encode format (JPEG/PNG). Defaults to JPEG.

deinterlace DeinterlaceMode

Deinterlace policy (off / force / auto). Bool aliases True→force, False→off are accepted. Defaults to off.

yadif_mode int

Yadif mode (0=send_frame, 1=send_field, 2=send_frame_nospatial, 3=send_field_nospatial). Defaults to 0.

yadif_parity int

Field parity (-1=auto, 0=tff, 1=bff). Defaults to -1.

yadif_deint int

Deinterlace all frames (0) or flagged-only (1). Defaults to 0.

default_timestamps list[float] | None

Used when process is called without process_config. Defaults to None (must pass per-call).

validate_default_timestamps(value) classmethod

Validate optional constructor-level timestamps.

Parameters:

Name Type Description Default
value list[float] | None

Candidate timestamps.

required

Returns:

Type Description
list[float] | None

list[float] | None: Validated timestamps or None.

validate_deinterlace_mode(value) classmethod

Coerce bool / string deinterlace values.

Parameters:

Name Type Description Default
value Any

Candidate mode.

required

Returns:

Name Type Description
DeinterlaceMode DeinterlaceMode

Normalized mode.

validate_output_format(value) classmethod

Reject blank output format strings.

Parameters:

Name Type Description Default
value str

Candidate format.

required

Returns:

Name Type Description
str str

Normalized upper-case format.

Raises:

Type Description
ValueError

If blank.

yadif_filter()

Build the shared yadif=mode:parity:deint filter description.

Returns:

Name Type Description
str str

Yadif filter string consumed by FFmpeg -vf and nested GStreamer deinterlace metadata.

FrameExtractionProcessConfig

Bases: ProcessorProcessConfig

Backend-agnostic per-invocation frame extraction config.

Attributes:

Name Type Description
timestamps list[float]

Required non-empty timestamps in seconds.

output_format str | None

Optional format override.

deinterlace DeinterlaceMode | None

Optional deinterlace override.

yadif_mode int | None

Optional yadif mode override.

yadif_parity int | None

Optional field-parity override.

yadif_deint int | None

Optional flagged-only override.

validate_optional_deinterlace_mode(value) classmethod

Coerce optional bool / string deinterlace overrides.

Parameters:

Name Type Description Default
value Any

Candidate mode or None.

required

Returns:

Type Description
DeinterlaceMode | None

DeinterlaceMode | None: Normalized mode or None.

validate_optional_output_format(value) classmethod

Normalize optional format override.

Parameters:

Name Type Description Default
value str | None

Candidate format.

required

Returns:

Type Description
str | None

str | None: Upper-case format or None.

validate_process_timestamps(value) classmethod

Validate per-call timestamps.

Parameters:

Name Type Description Default
value list[float]

Candidate timestamps.

required

Returns:

Type Description
list[float]

list[float]: Validated timestamps.

FrameExtractionProcessor()

Bases: BackendSelectableProcessor[Attachment, list[Attachment]], ABC

Family base for extracting image frames at timestamps.

Why use this base class?

  • Portability: Swap FFmpeg vs GStreamer without changing call sites.
  • Batch I/O: Extract many keyframes in one process call.
  • Separation: Keyframe planning stays in extractors; decode is here.

Usage

from gllm_multimodal.media_toolkit.processor.frame_extraction_processor import (
    FrameExtractionProcessConfig,
    FrameExtractionProcessor,
)

processor = FrameExtractionProcessor.build()
frames = await processor.process(
    video_attachment,
    process_config=FrameExtractionProcessConfig(timestamps=[1.5, 4.0]),
)

set_timestamps(timestamps) abstractmethod

Configure default timestamps used when process_config is omitted.

Parameters:

Name Type Description Default
timestamps list[float]

Non-empty list of non-negative seconds.

required

stream_frames(attachment, sample_fps, deinterlace=None) abstractmethod

Stream frames sampled at a fixed rate from a single decode pass.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
sample_fps float

Positive sampling rate in frames per second.

required
deinterlace DeinterlaceMode | None

Deinterlace override. Defaults to None.

None

Returns:

Type Description
AsyncIterator[tuple[float, ndarray]]

AsyncIterator[tuple[float, np.ndarray]]: Timestamps and BGR uint8 frames at source size.

FrameSamplingProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for resampling video attachments to a target frame rate.

This class serves as a unified entry point for frame sampling operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.

Why use this base class?

  • Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
  • Simplicity: No need to handle fallback logic or conditional imports yourself.
  • Future-proofing: New backends can be added to the library without requiring changes to your application code.

Usage Example

from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import (
    FrameSamplingProcessor,
    GstFrameSamplingConfig,
)
from gllm_inference.schema import Attachment

# Instantiates the best available backend automatically
processor = FrameSamplingProcessor.build(
    config=GstFrameSamplingConfig(default_target_fps=2)
)

attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.frame_sampling_processor import FrameSamplingProcessor
from gllm_inference.schema import Attachment

# Explicitly force the ffmpeg backend
processor = FrameSamplingProcessor.build(backend="ffmpeg")

attachment = Attachment(url="file:///path/to/video.mp4")
sampled_video = await processor.process(attachment)

GstAudioExtractionConfig

Bases: GstBaseConfig

Configuration for GstAudioExtractionProcessor.

Attributes:

Name Type Description
sample_rate int | None

Force output sample rate in Hz (e.g. 16000). None preserves the source rate. Defaults to None.

channels int | None

Force output channel count (e.g. 1 for mono). None preserves the source count. Defaults to None.

output_format str | None

Preferred container/extension to extract to ("wav", "mp3", "m4a", or "ogg"). When set, only encoders for that format are considered. Defaults to None (auto-select from all candidates, preferring WAV when available).

audio_encoder str | None

Pin a specific GStreamer audio encoder element (e.g. "lamemp3enc"). When both audio_encoder and output_format are set, the candidate list is first narrowed to the pinned encoder, then further filtered to the requested format; an incompatible combination raises ValueError. Defaults to None.

GstAudioExtractionProcessConfig

Bases: ProcessorProcessConfig

Per-invocation configuration for [process][gllm_multimodal.media_toolkit.media_toolkit.MediaToolkit.process].

Attributes:

Name Type Description
sample_rate int | None

Force output sample rate in Hz. None is a legitimate per-call override meaning "do not force a sample rate" — it suppresses the constructor self.config.sample_rate so the source rate is preserved.

channels int | None

Force output channel count. None is a legitimate per-call override meaning "do not force a channel count" — it suppresses the constructor self.config.channels.

GstAudioExtractionProcessor(config=None)

Bases: BaseGstreamerProcessor[GstAudioExtractionConfig], AudioExtractionProcessor

Extracts the audio track from a video Attachment using GStreamer.

The processor demuxes the input video, re-encodes (or passes through) the audio into the best available format, and returns the result as an Attachment whose mime_type reflects the audio container.

If the video has no audio track the original Attachment is returned unchanged, so callers do not need to handle None.

Attributes:

Name Type Description
audio_encoder_info EncoderFormatInfo | None

Selected audio encoder format info.

config GstAudioExtractionConfig

Runtime configuration.

Raises:

Type Description
RuntimeError

If GStreamer is unavailable or no suitable audio encoder is found.

ValueError

If output_format or audio_encoder is unsupported.

Example
processor = GstAudioExtractionProcessor(config={"output_format": "mp3"})
audio_attachment = await processor.process(video_attachment)
# audio_attachment.mime_type == "audio/mpeg"
# audio_attachment.filename  == "audio_my_video.mp3"

Initialise GStreamer and select the audio encoder / output format.

Parameters:

Name Type Description Default
config dict[str, Any] | GstAudioExtractionConfig | None

Optional configuration. A plain dict is coerced into GstAudioExtractionConfig. None uses defaults. Use output_format (e.g. "mp3") or audio_encoder (e.g. "lamemp3enc") to override auto-selection.

None

Raises:

Type Description
RuntimeError

If GStreamer is unavailable or no audio encoder is found.

ValueError

If output_format or audio_encoder is unsupported.

config_model() classmethod

Return the stable configuration model for this processor.

Returns:

Type Description
type[GstAudioExtractionConfig]

type[GstAudioExtractionConfig]: The configuration class used at

type[GstAudioExtractionConfig]

construction time for stable (per-instance) settings.

process(attachment, **kwargs) async

Extract the audio track of one video attachment.

This public entrypoint keeps configuration ergonomics simple for callers: stable defaults can be supplied in the constructor, while per-call overrides (for sample rate/channels) can be passed through process_config in kwargs.

Parameters:

Name Type Description Default
attachment Attachment

Input video attachment whose audio should be extracted.

required
**kwargs Any

Additional runtime options, typically process_config as GstAudioExtractionProcessConfig or dict.

{}
Notes
  1. Delegates shared mimetype validation and dispatch to MediaToolkit.process.

Returns:

Name Type Description
Attachment Attachment

Extracted audio attachment, or the original attachment when

Attachment

no audio stream is detected.

Example
processor = GstAudioExtractionProcessor(
    config={"output_format": "mp3"},
)
audio_attachment = await processor.process(
    attachment=video_attachment,
    process_config={"sample_rate": 16000, "channels": 1},
)

process_config_model() classmethod

Return the per-invocation configuration model for this processor.

Returns:

Type Description
type[GstAudioExtractionProcessConfig]

type[GstAudioExtractionProcessConfig]: The configuration class

type[GstAudioExtractionProcessConfig]

accepted by

type[GstAudioExtractionProcessConfig]
type[GstAudioExtractionProcessConfig]

for per-call overrides.

GstBaseConfig

Bases: BaseModel

Minimal shared configuration for all GStreamer-based processors.

Attributes:

Name Type Description
timeout int

Maximum seconds a GStreamer pipeline may run before being forcibly terminated. Defaults to 300 seconds.

audio_passthrough bool

When True (the default), encoded audio streams are passed directly to the muxer without being decoded and re-encoded. Set to False to force re-encoding via the best available audio encoder.

video_encoder str | None

Pin a specific GStreamer video encoder element name (e.g. "x264enc"). When set, the candidate-list scan is skipped entirely and this element is used directly. The element must exist in the GStreamer registry. Defaults to None (auto-select from _VIDEO_FORMAT_MAP).

audio_encoder str | None

Pin a specific GStreamer audio encoder element name (e.g. "voaacenc"). When set, the candidate-list scan is skipped entirely and this element is used directly. The element must exist in the GStreamer registry. Defaults to None (auto-select from _AUDIO_FORMAT_MAP).

GstDeinterlaceConfig

Bases: DeinterlaceConfig, GstBaseConfig

GStreamer-specific stable config (inherits shared DeinterlaceConfig).

Yadif / encode defaults mirror FFmpegDeinterlaceConfig so both backends produce comparable progressive output.

Attributes:

Name Type Description
yadif_mode int

Yadif mode (0–3). Defaults to 0.

yadif_parity int

Field parity (-1=auto, 0=tff, 1=bff). Defaults to -1.

yadif_deint int

Deinterlace all frames (0) or flagged-only (1). Defaults to 0.

crf int

Quantizer passed to the video encoder (CRF analogue). Defaults to 23.

preset str

x264 speed preset name. Defaults to ultrafast.

video_encoder str | None

Inherited from GstBaseConfig; None auto-selects the first available encoder (x264enc preferred), matching sibling Gst backends.

validate_preset(value) classmethod

Reject blank preset strings.

Parameters:

Name Type Description Default
value str

Candidate preset name.

required

Returns:

Name Type Description
str str

Validated preset name.

Raises:

Type Description
ValueError

If value is blank.

yadif_filter()

Return an FFmpeg-style yadif filter label for metadata parity.

Returns:

Name Type Description
str str

Filter description aligned with the FFmpeg backend.

GstDeinterlaceProcessConfig

Bases: DeinterlaceProcessConfig

Per-invocation GStreamer overrides (inherits shared DeinterlaceProcessConfig).

GstDeinterlaceProcessor(config=None)

Bases: BaseGstreamerProcessor[GstDeinterlaceConfig], DeinterlaceProcessor

Deinterlace video with GStreamer yadif + x264enc.

Pipeline topology::

filesrc → decodebin → queue → videoconvert → yadif → videoconvert
    → x264enc → mp4mux → filesink

Audio pads are dropped when strip_audio is True.

Initialize the GStreamer deinterlace processor.

Parameters:

Name Type Description Default
config dict[str, Any] | DeinterlaceConfig | GstDeinterlaceConfig | None

Shared or GStreamer-specific config. Defaults to None.

None

config_model() classmethod

Return the stable configuration model.

process(attachment, **kwargs) async

Deinterlace one video attachment.

Parameters:

Name Type Description Default
attachment Attachment

Input interlaced (or progressive) video.

required
**kwargs Any

May include process_config overrides.

{}

Returns:

Name Type Description
Attachment Attachment

Progressive video attachment.

process_config_model() classmethod

Return the per-invocation configuration model.

GstFrameDecodeConfig

GstFrameDecodeProcessConfig

Bases: FrameDecodeProcessConfig

Per-invocation GStreamer decode overrides (inherits shared process config).

GstFrameDecodeProcessor(config=None)

Bases: BaseGstreamerProcessor[GstFrameDecodeConfig], FrameDecodeProcessor

Decode every video frame to PNG attachments with system GStreamer.

Attributes:

Name Type Description
config GstFrameDecodeConfig

Runtime configuration.

Raises:

Type Description
RuntimeError

If GStreamer is not available.

Initialise without encoder selection (decode needs no encoder).

Parameters:

Name Type Description Default
config dict[str, Any] | FrameDecodeConfig | GstFrameDecodeConfig | None

Shared or GStreamer-specific config. Defaults to None.

None

config_model() classmethod

Return the stable configuration model for this processor.

Returns:

Type Description
type[GstFrameDecodeConfig]

type[GstFrameDecodeConfig]: The stable configuration model.

process(attachment, **kwargs) async

Decode every frame of one video attachment.

Parameters:

Name Type Description Default
attachment Attachment

Input video attachment to decode.

required
**kwargs Any

Runtime options, typically process_config as GstFrameDecodeProcessConfig or dict.

{}

Returns:

Type Description
list[Attachment]

list[Attachment]: PNG image attachments in decode order.

process_config_model() classmethod

Return the per-invocation configuration model for this processor.

Returns:

Type Description
type[GstFrameDecodeProcessConfig]

type[GstFrameDecodeProcessConfig]: The configuration class accepted by process for per-call overrides.

GstFrameExtractionConfig

Bases: FrameExtractionConfig, GstBaseConfig

GStreamer-specific stable config (inherits shared FrameExtractionConfig).

Shared yadif knobs are forwarded to DeinterlaceProcessor when deinterlace runs. Encode settings remain on GstDeinterlaceConfig.

Attributes:

Name Type Description
output_format str

Image encode format (JPEG/PNG). Defaults to JPEG.

deinterlace DeinterlaceMode

Whether to run the nested deinterlace family before extracting frames. Defaults to off.

yadif_mode int

Shared frame setting used by the FFmpeg backend; GStreamer uses its native deinterlace mode. Defaults to 0.

yadif_parity int

Shared field parity used by the FFmpeg backend. GStreamer reads field order from caps. Defaults to -1.

yadif_deint int

Shared flagged-only setting used by the FFmpeg backend. GStreamer reads interlace flags from caps. Defaults to 0.

default_timestamps list[float] | None

Used when process is called without process_config. Defaults to None.

validate_supported_image_format(value) classmethod

Reject still-image formats this backend cannot encode.

Parameters:

Name Type Description Default
value str

Candidate format (already upper-cased by the parent).

required

Returns:

Name Type Description
str str

Validated format.

Raises:

Type Description
ValueError

If the format is not JPEG/JPG/PNG.

GstFrameExtractionProcessConfig

Bases: FrameExtractionProcessConfig

Per-invocation GStreamer overrides (inherits shared process config).

validate_optional_supported_image_format(value) classmethod

Reject optional still-image format overrides this backend cannot encode.

Parameters:

Name Type Description Default
value str | None

Candidate format or None.

required

Returns:

Type Description
str | None

str | None: Validated format or None.

Raises:

Type Description
ValueError

If the format is not JPEG/JPG/PNG.

GstFrameExtractionProcessor(config=None)

Bases: BaseGstreamerProcessor[GstFrameExtractionConfig], FrameExtractionProcessor

Extract image frames with GStreamer.

When deinterlace is force or auto (and the source is interlaced), frames are taken from the progressive output of DeinterlaceProcessor rather than inlining yadif in the still-frame pipeline.

Attributes:

Name Type Description
config GstFrameExtractionConfig

Stable constructor configuration.

Initialize the GStreamer frame extraction processor.

Parameters:

Name Type Description Default
config dict[str, Any] | FrameExtractionConfig | GstFrameExtractionConfig | None

Shared or GStreamer-specific config. Defaults to None.

None

config_model() classmethod

Return the stable configuration model.

extract_frames(video_path, timestamps, *, params=None)

Extract encoded image bytes at each timestamp from a video path.

Parameters:

Name Type Description Default
video_path str

Source video path.

required
timestamps list[float]

Non-empty non-negative timestamps.

required
params GstFrameExtractionConfig | None

Effective options. Defaults to self.config.

None

Returns:

Type Description
list[bytes]

list[bytes]: Encoded frames in the same order as timestamps.

Raises:

Type Description
RuntimeError

If any extraction fails.

ValueError

If timestamps are invalid.

process(attachment, **kwargs) async

Extract frames from one video attachment.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
**kwargs Any

May include process_config.

{}

Returns:

Type Description
list[Attachment]

list[Attachment]: One image attachment per timestamp.

process_config_model() classmethod

Return the per-invocation configuration model.

set_timestamps(timestamps)

Update constructor-level default timestamps.

Parameters:

Name Type Description Default
timestamps list[float]

Non-empty non-negative timestamps.

required

stream_frames(attachment, sample_fps, deinterlace=None) async

Stream sampled BGR frames from a single GStreamer decode pass.

Frames are resampled by videorate, auto-rotated by videoflip and pulled from a bounded appsink, so decoding pauses while the consumer is busy. Closing the generator early stops the pipeline.

Parameters:

Name Type Description Default
attachment Attachment

Source video.

required
sample_fps float

Positive sampling rate in frames per second.

required
deinterlace DeinterlaceMode | None

Deinterlace override. Defaults to None.

None

Yields:

Type Description
AsyncIterator[tuple[float, ndarray]]

tuple[float, np.ndarray]: Timestamp and BGR uint8 frame at source size.

Raises:

Type Description
ValueError

If sample_fps is not finite and positive.

RuntimeError

If decoding fails, the source has no video stream, or the pipeline stalls.

GstFrameSamplingConfig

Bases: GstBaseConfig

Stable configuration for GstFrameSamplingProcessor.

Attributes:

Name Type Description
min_size_mb int

Minimum video size in MB to process. 0 disables the check. Defaults to 0.

default_target_fps int

Default target FPS used when GstFrameSamplingProcessor.process is called without a process_config. Must be greater than 0. Defaults to 1.

validate_default_target_fps(default_target_fps) classmethod

Validate that default_target_fps is positive.

Parameters:

Name Type Description Default
default_target_fps int

The candidate default FPS to validate.

required

Returns:

Name Type Description
int int

default_target_fps unchanged when valid.

Raises:

Type Description
ValueError

If default_target_fps is not greater than 0.

GstFrameSamplingProcessConfig

Bases: ProcessorProcessConfig

Per-invocation configuration for [process][gllm_multimodal.media_toolkit.media_toolkit.MediaToolkit.process].

Attributes:

Name Type Description
target_fps int

Target frames-per-second for the output video.

validate_target_fps(target_fps) classmethod

Validate that target_fps is positive.

Parameters:

Name Type Description Default
target_fps int

The candidate target FPS to validate.

required

Returns:

Name Type Description
int int

target_fps unchanged when valid.

Raises:

Type Description
ValueError

If target_fps is not greater than 0.

GstFrameSamplingProcessor(config=None)

Bases: BaseGstreamerProcessor[GstFrameSamplingConfig], FrameSamplingProcessor

Resamples a video attachment to a target frame-rate using GStreamer.

The processor is format-agnostic: it handles MP4, MKV, MOV, and similar containers and preserves any audio track (passthrough by default).

Attributes:

Name Type Description
config GstFrameSamplingConfig

Runtime configuration.

Raises:

Type Description
RuntimeError

If GStreamer is not available or no suitable encoder is found.

Initialise and select encoders.

Parameters:

Name Type Description Default
config dict[str, Any] | GstFrameSamplingConfig | None

Optional stable configuration.

None

Raises:

Type Description
ValueError

If config.default_target_fps is non-positive.

config_model() classmethod

Return the stable configuration model for this processor.

Returns:

Type Description
type[GstFrameSamplingConfig]

type[GstFrameSamplingConfig]: The stable configuration model for this processor.

process(attachment, **kwargs) async

Resample one video attachment to the configured or requested FPS.

The method preserves the regular media-toolkit process contract while exposing this concrete processor's runtime semantics in API docs: constructor defaults establish baseline FPS behavior, and per-call overrides can be supplied through process_config.

Parameters:

Name Type Description Default
attachment Attachment

Input video attachment to resample.

required
**kwargs Any

Runtime options, typically process_config as GstFrameSamplingProcessConfig or dict.

{}
Notes
  1. Delegates shared mimetype validation and dispatch to MediaToolkit.process.

Returns:

Name Type Description
Attachment Attachment

Resampled video attachment, or the original attachment

Attachment

when the input is skipped by validation rules.

Example
processor = GstFrameSamplingProcessor(
    config={"default_target_fps": 2},
)
sampled_attachment = await processor.process(
    attachment=video_attachment,
    process_config={"target_fps": 1},
)

process_config_model() classmethod

Return the per-invocation configuration model for this processor.

Returns:

Type Description
type[GstFrameSamplingProcessConfig]

type[GstFrameSamplingProcessConfig]: The configuration class

type[GstFrameSamplingProcessConfig]

accepted by

type[GstFrameSamplingProcessConfig]
type[GstFrameSamplingProcessConfig]

for per-call overrides.

GstVideoClipConfig

Bases: GstBaseConfig

Stable configuration for GstVideoClipProcessor.

Set once at construction. Controls encoder selection, timeout, and other processor behaviour that does not change per attachment.

Attributes:

Name Type Description
default_windows list[tuple[float, float]] | None

Default clipping windows used when GstVideoClipProcessor.process is called without a process_config. None (the default) means windows must be supplied per-call via GstVideoClipProcessConfig. When set, the value is validated by GstVideoClipProcessConfig's field_validator.

validate_default_windows(default_windows) classmethod

Validate the constructor-level default windows, if any.

Parameters:

Name Type Description Default
default_windows list[tuple[float, float]] | None

The candidate default windows to validate.

required

Returns:

Type Description
list[tuple[float, float]] | None

list[tuple[float, float]] | None: default_windows unchanged

list[tuple[float, float]] | None

when None or when every window is well-ordered.

Raises:

Type Description
ValueError

If default_windows is empty or any window has end <= start. Validation mirrors GstVideoClipProcessConfig's field_validator.

GstVideoClipProcessConfig

Bases: ProcessorProcessConfig

Per-invocation configuration for [process][gllm_multimodal.media_toolkit.media_toolkit.MediaToolkit.process].

Attributes:

Name Type Description
windows list[tuple[float, float]]

Ordered (start_time, end_time) windows in seconds. Each window must satisfy end > start.

validate_windows(windows) classmethod

Validate that every clipping window is non-empty and well ordered.

Parameters:

Name Type Description Default
windows list[tuple[float, float]]

The windows to validate.

required

Returns:

Type Description
list[tuple[float, float]]

list[tuple[float, float]]: The validated windows.

GstVideoClipProcessor(config=None)

Bases: BaseGstreamerProcessor[GstVideoClipConfig], VideoClipProcessor

Clips a video attachment to one or more [start_time, end_time] windows.

Stable settings belong in GstVideoClipConfig (constructor). Per-call clip windows belong in GstVideoClipProcessConfig (process(..., process_config=...)).

Initialise the clip processor.

Parameters:

Name Type Description Default
config dict[str, Any] | GstVideoClipConfig | None

Optional stable configuration. None uses GstVideoClipConfig defaults.

None

Raises:

Type Description
ValueError

If config.default_windows is empty or any window has end <= start.

config_model() classmethod

Return the stable configuration model for this processor.

Returns:

Type Description
type[GstVideoClipConfig]

type[GstVideoClipConfig]: The stable configuration model for this processor.

process(attachment, **kwargs) async

Clip one video attachment using resolved time windows.

This method is the canonical caller-facing API for clipping. It accepts stable default windows set at construction time and also supports per-call window overrides through process_config.

Parameters:

Name Type Description Default
attachment Attachment

Input video attachment to clip.

required
**kwargs Any

Runtime options, typically process_config as GstVideoClipProcessConfig or dict.

{}
Notes
  1. Delegates shared mimetype validation and dispatch to MediaToolkit.process.

Returns:

Type Description
Attachment | list[Attachment]

Attachment | list[Attachment]: Single clipped attachment when exactly

Attachment | list[Attachment]

one window is used, otherwise a list of clipped attachments in window order.

Example
processor = GstVideoClipProcessor(
    config={"default_windows": [(0.0, 10.0)]},
)
clips = await processor.process(
    attachment=video_attachment,
    process_config={
        "windows": [(5.0, 12.5), (30.0, 45.0)],
    },
)

process_config_model() classmethod

Return the per-invocation configuration model for this processor.

Returns:

Type Description
type[GstVideoClipProcessConfig]

type[GstVideoClipProcessConfig]: The per-invocation configuration model for this processor.

set_windows(windows)

Update the constructor-level default windows.

Mutates config so process falls back to windows when no process_config is supplied. Prefer passing GstVideoClipProcessConfig to process for per-call overrides.

Parameters:

Name Type Description Default
windows list[tuple[float, float]]

The windows to set.

required

Returns:

Name Type Description
None None

self.config is replaced with a copy whose

None

default_windows is the validated windows.

Raises:

Type Description
ValueError

If windows is empty or any window has end <= start. Validation reuses GstVideoClipProcessConfig's field_validator.

MediaToolkit()

Bases: ABC, Generic[T_in, T_out]

Base abstraction for all media toolkit processing components.

This class provides the shared lifecycle and registry behavior used by both: - concrete leaf processors (e.g. backend-specific audio/video processors), and - composite components (e.g. segmenters, keyframe extractors) that orchestrate nested processors.

Key responsibilities: - auto-register subclasses by class name for class-name-based construction via build; - provide consistent input validation against supported_mimetypes; - define async processing contracts through process and process_batch.

Contributor guidance: - inherit this class directly for concrete processors with custom behavior; - inherit BackendSelectableProcessor when one logical processor family maps to multiple backend implementations; - inherit composite bases (e.g. BaseSegmenter) for orchestration-style components.

Example
Building a processor by class name
from gllm_multimodal.media_toolkit.media_toolkit import MediaToolkit

processor = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
result = await processor.process(video_attachment)
Checking mimetype support
if processor.is_supported(attachment):
    result = await processor.process(attachment)
Listing registered processors
print(list(MediaToolkit.registry.keys()))
# ['GstAudioExtractionProcessor', 'GstVideoClipProcessor', ...]

Initialize processor logging.

name property

Return a stable component name for optional plan metadata.

Returns:

Name Type Description
str str

Component class name.

registry = {} class-attribute

Global class-name registry used by build.

Maps each registered subclass name to its concrete class type, enabling string-based construction such as MediaToolkit.build("AudioExtractionProcessor").

supported_mimetypes = ['*/*'] class-attribute

MIME types this processor accepts (supports wildcards, e.g. 'video/*').

Defaults to ['*/*'] (accept all). Override as a class attribute in subclasses.

__init_subclass__(**kwargs)

Register every concrete subclass into the global registry.

This hook is triggered automatically by Python whenever a class inherits from MediaToolkit (directly or indirectly). Registration happens at class definition/import time, so classes become immediately discoverable by build without manual setup.

Registration key
  1. The subclass' __name__ (e.g. "AudioExtractionProcessor").
  2. The value stored is the subclass type itself.
Why uniqueness is enforced
  1. build(class_name=...) uses this registry for class resolution.
  2. Duplicate class names would silently shadow earlier classes and could route builds to unintended implementations.
  3. To prevent that ambiguity, duplicate keys raise TypeError.

Parameters:

Name Type Description Default
**kwargs Any

Extra class declaration keyword arguments forwarded to parent __init_subclass__ implementations.

{}

Raises:

Type Description
TypeError

If a subclass with the same class name is already registered, preventing silent dispatch to the wrong implementation.

available_backends_for(class_name) classmethod

Return backend keys registered for a processor family class name.

Parameters:

Name Type Description Default
class_name str

Registered processor family class name.

required

Returns:

Type Description
list[str]

list[str]: Available backend keys. Empty when the class is unknown or not a backend-selectable family base.

Raises:

Type Description
ValueError

If the class name is unknown.

Example

backends = MediaToolkit.available_backends_for("VideoClipProcessor")
print(backends)
Output:
["gstreamer", "ffmpeg"]

build(class_name, backend=None, **kwargs) classmethod

Build a processor by class name.

Family abstract classes (e.g. AudioExtractionProcessor) resolve a concrete backend implementation via backend. Composite components (segmenters, keyframe extractors) store backend on the instance for nested resolution.

Parameters:

Name Type Description Default
class_name str

Registered subclass name.

required
backend str | MediaBackend | None

Backend key for family classes, or nested processor preference for composite instances. Defaults to None.

None
**kwargs Any

Constructor kwargs passed to the processor class.

{}

Returns:

Name Type Description
MediaToolkit MediaToolkit

Instantiated processor.

Raises:

Type Description
ValueError

If the class name is unknown.

Example

proc = MediaToolkit.build("AudioExtractionProcessor", backend="gstreamer")
seg = MediaToolkit.build(
    "FixedDurationSegmenter",
    backend="gstreamer",
    config={"segment_durations": [2.0]},
)
print(proc)
print(seg)
Output:
<GstAudioExtractionProcessor instance>
<FixedDurationSegmenter instance>

build_from_registry(backend=None, **kwargs) classmethod

Instantiate this registered class.

Subclasses override this hook to customize registry-based construction (e.g. backend-selectable families resolve a concrete backend; composites store backend for nested processor resolution).

Parameters:

Name Type Description Default
backend str | MediaBackend | None

Backend key forwarded to subclass overrides. Ignored by the base implementation.

None
**kwargs Any

Constructor kwargs passed to the processor class.

{}

Returns:

Name Type Description
MediaToolkit MediaToolkit

Instantiated processor.

Example

# Called indirectly by MediaToolkit.build(...)
processor = SomeRegisteredProcessor.build_from_registry(custom_flag=True)
print(processor)
Output:
<SomeRegisteredProcessor instance>

is_supported(attachment)

Return whether the attachment's mimetype is accepted by this processor.

Callers can use this to check compatibility before calling process or process_batch, avoiding a ValueError.

Parameters:

Name Type Description Default
attachment Attachment

The attachment to check.

required

Returns:

Name Type Description
bool bool

True if the attachment's mimetype matches any entry in supported_mimetypes (including wildcards). True is also returned when the attachment has no mimetype set.

Example

ok = processor.is_supported(Attachment(mime_type="video/mp4"))
print(ok)
Output:
True

list_available_backends() classmethod

Return backend keys when this class supports backend selection.

Returns:

Type Description
list[str]

list[str]: Available backend keys. Empty for classes that are not backend-selectable family bases.

Example

Base classes are not backend-selectable.

print(MediaToolkit.list_available_backends())
Output:
[]

Family classes expose registered backend keys.

from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor

backends = VideoClipProcessor.list_available_backends()
print(backends)
Output:
["gstreamer", "ffmpeg", "moviepy"]  # depends on registered backends

process(attachment, **kwargs) async

Process a single attachment (or perform an aggregation on a list) and return the result.

Parameters:

Name Type Description Default
attachment T_in

The attachment or list of attachments to process.

required
**kwargs Any

Additional keyword arguments forwarded to _process.

{}

Returns:

Name Type Description
T_out T_out

The result of the processing.

Example

result = await processor.process(attachment)
print(result)
Output:
<processed attachment or transformed output>

process_batch(attachments, **kwargs) async

Process a batch of attachments.

Parameters:

Name Type Description Default
attachments list[T_in]

The batch of attachments to process.

required
**kwargs Any

Additional keyword arguments forwarded to process.

{}

Returns:

Type Description
list[T_out]

list[T_out]: The result of the batch processing.

Example

results = await processor.process_batch([attachment_1, attachment_2])
print(results)
Output:
[<result_1>, <result_2>]

ProcessorProcessConfig

Bases: BaseModel

Marker base class for per-invocation processor configuration.

Concrete processors define subclasses with the parameters that may change on every process() call. Stable processor settings belong in each processor's *Config model passed at construction time.

VideoClipProcessor()

Bases: BackendSelectableProcessor[Attachment, Attachment], ABC

Family base for clipping video attachments to time windows.

This class serves as a unified entry point for video clipping operations. It automatically routes requests to the most appropriate, available backend implementation based on your system environment.

Why use this base class?

  • Portability: Your code will run regardless of which underlying libraries are installed on the host machine.
  • Simplicity: No need to handle fallback logic or conditional imports yourself.
  • Future-proofing: New backends can be added to the library without requiring changes to your application code.

Usage Example

from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment

# Instantiates the best available backend automatically
processor = VideoClipProcessor.build()

# Set the target clipping window (start_time, end_time) in seconds
processor.set_windows([(10.0, 20.5)])

attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment

# Explicitly force the ffmpeg backend
processor = VideoClipProcessor.build(backend="ffmpeg")
processor.set_windows([(10.0, 20.5)])

attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)
from gllm_multimodal.media_toolkit.processor.video_clip_processor import VideoClipProcessor
from gllm_inference.schema import Attachment

# Explicitly force the moviepy backend
processor = VideoClipProcessor.build(backend="moviepy")
processor.set_windows([(10.0, 20.5)])

attachment = Attachment(url="file:///path/to/video.mp4")
clipped_video = await processor.process(attachment)

set_windows(windows) abstractmethod

Configure one or more [start, end] clipping windows for the next call.

Parameters:

Name Type Description Default
windows list[tuple[float, float]]

List of (start, end) tuples in seconds.

required