Skip to content

Overview

Segmenter components for temporal media splitting.

This module provides segmenters that split media files (audio, video) into fixed-duration or content-aware segments for parallel processing.

Exported Classes

BaseSegmenter(config=None)

Bases: CompositeMediaMixin, MediaToolkit[Attachment, list[Attachment]], ABC, Generic[TConfig]

Abstract base class for segmenting media attachments.

Segmenters are responsible for dividing a media attachment (usually video or audio) into semantically or temporally distinct segments. The resulting attachments should contain VideoSegment metadata.

Attributes:

Name Type Description
backend MediaBackend | str | None

Preferred backend for nested processor lookup. Set when built via the registry with an explicit or resolved backend.

Processing contract
  • await segment(attachment) computes boundaries only.
  • await process(attachment) computes + materializes media segments.
  • Materialization can call nested processors via get_processor(...) and uses this segmenter's backend preference.
Example

segmenter = build_media_toolkit( "FixedDurationSegmenter", backend=MediaBackend.GSTREAMER, ... ) then await segmenter.process(video_attachment) uses the GStreamer video clip processor during materialization.

Validate and store the concrete segmenter's configuration.

Parameters:

Name Type Description Default
config TConfig | BaseSegmenterConfig | dict[str, Any] | None

Config model, base config, mapping, or None for model defaults. Concrete config models decide whether defaults satisfy their required fields.

None

config_model() classmethod

Return the concrete Pydantic model used to validate this segmenter's config.

The base implementation provides BaseSegmenterConfig for simple subclasses that do not define additional configuration fields.

Returns:

Type Description
type[TConfig]

type[TConfig]: The config model associated with the concrete segmenter.

cuts_to_segments(cut_times, duration, *, min_shot_duration) staticmethod

Convert interior cut timestamps into contiguous shot windows.

Parameters:

Name Type Description Default
cut_times list[float]

Interior cut timestamps in seconds.

required
duration float

Total video duration in seconds.

required
min_shot_duration float

Windows shorter than this (after the first) are merged into the previous window.

required

Returns:

Type Description
list[VideoSegment]

list[VideoSegment]: Shot windows spanning [0, duration].

materialize(attachment, segment, segment_index=0) async

Materialize one segment plan into a media attachment.

Guarantees mimetype validation before calling _materialize.

Parameters:

Name Type Description Default
attachment Attachment

Source attachment to clip or reference.

required
segment VideoSegment

Segment plan with boundary metadata.

required
segment_index int

Zero-based index within the segment plan list. Defaults to 0.

0

Returns:

Name Type Description
Attachment Attachment

Materialized segment attachment with

Attachment

[VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata.

output_validator(attachment)

Validate that each output attachment has VideoSegment-compatible metadata dict.

Parameters:

Name Type Description Default
attachment Attachment | list[Attachment]

The attachment(s) to validate.

required

Returns:

Type Description
Attachment | list[Attachment]

Attachment | list[Attachment]: The validated attachment(s).

Raises:

Type Description
TypeError

If any attachment metadata is not a dictionary.

ValueError

If any metadata dictionary cannot be parsed as [VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment].

segment(attachment) async

Compute segment boundaries for one attachment without materializing media.

Guarantees mimetype validation before calling _segment.

Parameters:

Name Type Description Default
attachment Attachment

The attachment to segment.

required

Returns:

Type Description
list[VideoSegment]

list[VideoSegment]: Segment plans containing at least start_time and end_time.

BaseSegmenterConfig

Bases: BaseModel

Empty base config used to carry values into a concrete segmenter config.

Concrete segmenter configs declare and validate only the settings their implementation consumes. Extra fields are retained here so callers can up-cast an exact base config with from_dict.

from_dict(data=None) classmethod

Build a validated config from a dict, base config, or existing instance.

Sibling subclass configs are rejected: re-validating an unrelated model would silently drop (or carry over) fields the caller never set. Pass a dict or an exact BaseSegmenterConfig to up-cast into the concrete subclass.

Parameters:

Name Type Description Default
data Self | BaseSegmenterConfig | dict[str, Any] | None

Raw config payload. None uses field defaults. Defaults to None.

None

Returns:

Name Type Description
Self Self

Validated config of the concrete subclass.

Raises:

Type Description
TypeError

If data is a BaseModel that is neither an instance of cls nor an exact BaseSegmenterConfig, or is of any other unsupported type.

Detector

Bases: StrEnum

Supported shot-detector algorithm names.

FixedDurationSegmenter(config=None)

Bases: BaseSegmenter[FixedDurationSegmenterConfig]

Segment attachments using explicit per-segment durations.

segment returns cumulative time windows from config.segment_durations. materialize clips each window into a separate attachment using the cached VideoClipProcessor obtained from CompositeMediaMixin.

config_model() classmethod

Return this segmenter's concrete config model.

Returns:

Type Description
type[FixedDurationSegmenterConfig]

type[FixedDurationSegmenterConfig]: The model used to validate fixed-duration segmenter configuration.

materialize(attachment, segment, segment_index=0) async

Clip and return one attachment for a precomputed segment window.

This method is convenient when segment planning and clip extraction are performed in separate stages, and only selected windows should be materialized.

Parameters:

Name Type Description Default
attachment Attachment

Source media attachment to clip.

required
segment VideoSegment

Segment boundary plan to materialize.

required
segment_index int

Zero-based index used in generated output filenames. Defaults to 0.

0

Returns:

Name Type Description
Attachment Attachment

Clipped attachment with [VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata.

process(attachment, **kwargs) async

Materialize fixed-duration clips from one media attachment.

Unlike calling segment directly, this method returns real clipped attachment outputs with [VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata embedded on each result. It is the main runtime entrypoint when you need files/bytes for every configured duration window, not only boundary plans.

Parameters:

Name Type Description Default
attachment Attachment

Source media attachment to split.

required
**kwargs Any

Forwarded processing arguments accepted by the base media-toolkit contract.

{}
Notes
  1. Delegates shared validation and orchestration to MediaToolkit.process (inherited by BaseSegmenter).

Returns:

Type Description
list[Attachment]

list[Attachment]: One clipped attachment per configured segment window.

segment(attachment) async

Return computed fixed windows without creating clip attachments.

This is useful for previewing time boundaries (for inspection, logging, or downstream planning) before paying the cost of media clipping.

Parameters:

Name Type Description Default
attachment Attachment

Source media attachment. The payload itself is not read by this implementation when computing boundaries.

required
Notes
  1. Delegates shared validation to BaseSegmenter.segment.

Returns:

Type Description
list[VideoSegment]

list[VideoSegment]: Fixed cumulative windows derived from

list[VideoSegment]

segment_durations and start_time.

FixedDurationSegmenterConfig

Bases: BaseSegmenterConfig

Configuration for FixedDurationSegmenter.

Requires a non-empty segment_durations list.

Attributes:

Name Type Description
segment_durations list[float]

Ordered segment durations in seconds.

start_time float

Base start time for the first segment. Defaults to 0.0.

validate_segment_durations(value) classmethod

Reject non-positive segment durations.

Parameters:

Name Type Description Default
value list[float]

Candidate segment durations.

required

Returns:

Type Description
list[float]

list[float]: The durations unchanged when valid.

Raises:

Type Description
ValueError

If any duration is non-positive.

ShotBasedSegmenter(detector=Detector.CONTENT, config=None)

Bases: BaseSegmenter[ShotBasedSegmenterConfig]

Segment video into shots with swappable decode + numpy/skimage scoring.

Frame decoding is delegated to the nested FrameDecodeProcessor family (GStreamer by default; override via composite backend or set_processor_backend); shot scoring runs on numpy + scikit-image.

Attributes:

Name Type Description
detector Detector

content | adaptive | threshold.

Initialize the FIPS-friendly content segmenter.

Parameters:

Name Type Description Default
detector Detector | str

Detector algorithm. Defaults to content.

CONTENT
config ShotBasedSegmenterConfig | BaseSegmenterConfig | dict[str, Any] | None

Scoring / decode configuration. Defaults to None (use defaults).

None

Raises:

Type Description
ValueError

If detector is not a supported algorithm name.

ValidationError

If scoring / decode knobs fail ShotBasedSegmenterConfig.

ImportError

If numpy / scikit-image / Pillow are not installed.

config_model() classmethod

Return this segmenter's concrete config model.

Returns:

Type Description
type[ShotBasedSegmenterConfig]

type[ShotBasedSegmenterConfig]: The model used to validate shot-based segmenter configuration.

ShotBasedSegmenterConfig

Bases: BaseSegmenterConfig

Configuration for ShotBasedSegmenter.

Decode and detector settings specific to content-based shot detection.

Attributes:

Name Type Description
sample_fps int | None

Decode sampling rate. Defaults to 5.

target_width int | None

Decode downscale width. Defaults to 320.

min_shot_duration float

Minimum accepted shot length in seconds. Defaults to 1.0.

threshold float

Detector threshold. Defaults to 27.0.

min_content_val float

adaptive-only absolute score floor. Defaults to 15.0.

window_width int

adaptive-only neighbour half-window. Defaults to 2.