Overview
Segmenter components for temporal media splitting.
This module provides segmenters that split media files (audio, video) into fixed-duration or content-aware segments for parallel processing.
Exported Classes
BaseSegmenter-- Abstract base for media segmentation.BaseSegmenterConfig-- Empty base configuration used to up-cast values into concrete segmenter configs.FixedDurationSegmenter-- Fixed-duration segment splitting.FixedDurationSegmenterConfig-- Configuration for fixed-duration segmentation.Detector-- Shot-detector algorithm names.ShotBasedSegmenter-- FIPS-friendly content-based shot detection (GStreamer/FFmpeg decode + numpy / scikit-image scoring).ShotBasedSegmenterConfig-- Configuration for shot-based segmentation.
BaseSegmenter(config=None)
Bases: CompositeMediaMixin, MediaToolkit[Attachment, list[Attachment]], ABC, Generic[TConfig]
Abstract base class for segmenting media attachments.
Segmenters are responsible for dividing a media attachment (usually video or audio)
into semantically or temporally distinct segments. The resulting attachments
should contain VideoSegment metadata.
Attributes:
| Name | Type | Description |
|---|---|---|
backend |
MediaBackend | str | None
|
Preferred backend for nested processor lookup. Set when built via the registry with an explicit or resolved backend. |
Processing contract
await segment(attachment)computes boundaries only.await process(attachment)computes + materializes media segments.- Materialization can call nested processors via
get_processor(...)and uses this segmenter's backend preference.
Example
segmenter = build_media_toolkit(
"FixedDurationSegmenter", backend=MediaBackend.GSTREAMER, ...
)
then await segmenter.process(video_attachment) uses the GStreamer
video clip processor during materialization.
Validate and store the concrete segmenter's configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
TConfig | BaseSegmenterConfig | dict[str, Any] | None
|
Config model, base config, mapping, or |
None
|
config_model()
classmethod
Return the concrete Pydantic model used to validate this segmenter's config.
The base implementation provides BaseSegmenterConfig for simple
subclasses that do not define additional configuration fields.
Returns:
| Type | Description |
|---|---|
type[TConfig]
|
type[TConfig]: The config model associated with the concrete segmenter. |
cuts_to_segments(cut_times, duration, *, min_shot_duration)
staticmethod
Convert interior cut timestamps into contiguous shot windows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cut_times
|
list[float]
|
Interior cut timestamps in seconds. |
required |
duration
|
float
|
Total video duration in seconds. |
required |
min_shot_duration
|
float
|
Windows shorter than this (after the first) are merged into the previous window. |
required |
Returns:
| Type | Description |
|---|---|
list[VideoSegment]
|
list[VideoSegment]: Shot windows spanning |
materialize(attachment, segment, segment_index=0)
async
Materialize one segment plan into a media attachment.
Guarantees mimetype validation before calling _materialize.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source attachment to clip or reference. |
required |
segment
|
VideoSegment
|
Segment plan with boundary metadata. |
required |
segment_index
|
int
|
Zero-based index within the segment plan list. Defaults to 0. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Materialized segment attachment with |
Attachment
|
[ |
output_validator(attachment)
Validate that each output attachment has VideoSegment-compatible metadata dict.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment | list[Attachment]
|
The attachment(s) to validate. |
required |
Returns:
| Type | Description |
|---|---|
Attachment | list[Attachment]
|
Attachment | list[Attachment]: The validated attachment(s). |
Raises:
| Type | Description |
|---|---|
TypeError
|
If any attachment metadata is not a dictionary. |
ValueError
|
If any metadata dictionary cannot be parsed as
[ |
segment(attachment)
async
Compute segment boundaries for one attachment without materializing media.
Guarantees mimetype validation before calling _segment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
The attachment to segment. |
required |
Returns:
| Type | Description |
|---|---|
list[VideoSegment]
|
list[VideoSegment]: Segment plans containing at least |
BaseSegmenterConfig
Bases: BaseModel
Empty base config used to carry values into a concrete segmenter config.
Concrete segmenter configs declare and validate only the settings their
implementation consumes. Extra fields are retained here so callers can
up-cast an exact base config with from_dict.
from_dict(data=None)
classmethod
Build a validated config from a dict, base config, or existing instance.
Sibling subclass configs are rejected: re-validating an unrelated
model would silently drop (or carry over) fields the caller never
set. Pass a dict or an exact BaseSegmenterConfig to up-cast
into the concrete subclass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Self | BaseSegmenterConfig | dict[str, Any] | None
|
Raw config
payload. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
Self |
Self
|
Validated config of the concrete subclass. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
Detector
Bases: StrEnum
Supported shot-detector algorithm names.
FixedDurationSegmenter(config=None)
Bases: BaseSegmenter[FixedDurationSegmenterConfig]
Segment attachments using explicit per-segment durations.
segment
returns cumulative time windows from config.segment_durations.
materialize
clips each window into a separate attachment using the cached
VideoClipProcessor obtained from CompositeMediaMixin.
config_model()
classmethod
Return this segmenter's concrete config model.
Returns:
| Type | Description |
|---|---|
type[FixedDurationSegmenterConfig]
|
type[FixedDurationSegmenterConfig]: The model used to validate fixed-duration segmenter configuration. |
materialize(attachment, segment, segment_index=0)
async
Clip and return one attachment for a precomputed segment window.
This method is convenient when segment planning and clip extraction are performed in separate stages, and only selected windows should be materialized.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source media attachment to clip. |
required |
segment
|
VideoSegment
|
Segment boundary plan to materialize. |
required |
segment_index
|
int
|
Zero-based index used in generated output filenames. Defaults to 0. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
Attachment |
Attachment
|
Clipped attachment with
[ |
process(attachment, **kwargs)
async
Materialize fixed-duration clips from one media attachment.
Unlike calling
segment
directly, this method returns real clipped attachment outputs with
[VideoSegment][gllm_core.schema.multimodal.video_caption.VideoSegment] metadata embedded on
each result.
It is the main runtime entrypoint when you need files/bytes for every
configured duration window, not only boundary plans.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source media attachment to split. |
required |
**kwargs
|
Any
|
Forwarded processing arguments accepted by the base media-toolkit contract. |
{}
|
Notes
- Delegates shared validation and orchestration to
MediaToolkit.process(inherited byBaseSegmenter).
Returns:
| Type | Description |
|---|---|
list[Attachment]
|
list[Attachment]: One clipped attachment per configured segment window. |
segment(attachment)
async
Return computed fixed windows without creating clip attachments.
This is useful for previewing time boundaries (for inspection, logging, or downstream planning) before paying the cost of media clipping.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
attachment
|
Attachment
|
Source media attachment. The payload itself is not read by this implementation when computing boundaries. |
required |
Notes
- Delegates shared validation to
BaseSegmenter.segment.
Returns:
| Type | Description |
|---|---|
list[VideoSegment]
|
list[VideoSegment]: Fixed cumulative windows derived from |
list[VideoSegment]
|
|
FixedDurationSegmenterConfig
Bases: BaseSegmenterConfig
Configuration for FixedDurationSegmenter.
Requires a non-empty segment_durations list.
Attributes:
| Name | Type | Description |
|---|---|---|
segment_durations |
list[float]
|
Ordered segment durations in seconds. |
start_time |
float
|
Base start time for the first segment. Defaults to 0.0. |
validate_segment_durations(value)
classmethod
Reject non-positive segment durations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
list[float]
|
Candidate segment durations. |
required |
Returns:
| Type | Description |
|---|---|
list[float]
|
list[float]: The durations unchanged when valid. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any duration is non-positive. |
ShotBasedSegmenter(detector=Detector.CONTENT, config=None)
Bases: BaseSegmenter[ShotBasedSegmenterConfig]
Segment video into shots with swappable decode + numpy/skimage scoring.
Frame decoding is delegated to the nested FrameDecodeProcessor family
(GStreamer by default; override via composite backend or
set_processor_backend); shot scoring runs on numpy + scikit-image.
Attributes:
| Name | Type | Description |
|---|---|---|
detector |
Detector
|
|
Initialize the FIPS-friendly content segmenter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Detector | str
|
Detector algorithm. Defaults to |
CONTENT
|
config
|
ShotBasedSegmenterConfig | BaseSegmenterConfig | dict[str, Any] | None
|
Scoring / decode configuration. Defaults to None (use defaults). |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
ValidationError
|
If scoring / decode knobs fail |
ImportError
|
If numpy / scikit-image / Pillow are not installed. |
config_model()
classmethod
Return this segmenter's concrete config model.
Returns:
| Type | Description |
|---|---|
type[ShotBasedSegmenterConfig]
|
type[ShotBasedSegmenterConfig]: The model used to validate shot-based segmenter configuration. |
ShotBasedSegmenterConfig
Bases: BaseSegmenterConfig
Configuration for ShotBasedSegmenter.
Decode and detector settings specific to content-based shot detection.
Attributes:
| Name | Type | Description |
|---|---|---|
sample_fps |
int | None
|
Decode sampling rate. Defaults to 5. |
target_width |
int | None
|
Decode downscale width. Defaults to 320. |
min_shot_duration |
float
|
Minimum accepted shot length in seconds. Defaults to 1.0. |
threshold |
float
|
Detector threshold. Defaults to 27.0. |
min_content_val |
float
|
|
window_width |
int
|
|