Skip to content

Overview

Audio-to-transcript task implementations.

This package contains the various transcription strategies for converting audio into timestamped transcript segments.

Submodules

  • asr -- ASR-based transcription (Google Cloud, OpenAI Whisper, Prosa, Qwen, Snowflake).
  • lm_based -- LM-based transcription using language models (Gemini).
  • transcript_fetch -- Transcript fetching from external sources (YouTube).

Exported Classes

ASRBasedAudioToTranscript()

Bases: BaseAudioToTranscript

Abstract base class for dedicated speech-to-text API integrations.

ASRBasedGoogleCloudAudioToTranscript(credentials_json, bucket_name, language_code='id-ID', alternative_language_codes=None, model='latest_long', timeout=5 * 60, proxy=None)

Bases: ASRBasedAudioToTranscript

An audio to text converter using Google Cloud Speech-to-Text.

The ASRBasedGoogleCloudAudioToTranscript class is responsible for converting audio to text using the Google Cloud Speech-to-Text. It supports various audio input formats and can handle audio from local files, base64 encoded strings, or URLs pointing to audio files.

Attributes:

Name Type Description
speech_client SpeechClient

Google Cloud Speech-to-Text client.

storage_client Client

Google Cloud Storage client.

bucket_name str

Google Cloud Storage bucket name.

language_code str

Language code for transcription.

alternative_language_codes list[str] | None

Alternative language codes for transcription. Up to 3 alternatives.

model str

Transcription model name.

timeout int

Timeout for the transcription request.

proxy str | None

The proxy URL to use for the YouTube request.

Initialize the ASRBasedGoogleCloudAudioToTranscript instance.

Parameters:

Name Type Description Default
credentials_json str | dict

Google Cloud API credentials can either be a file path or a dictionary.

required
bucket_name str

Google Cloud Storage bucket name.

required
language_code str

Language code for transcription. Defaults to "id-ID".

'id-ID'
alternative_language_codes list[str] | None

Alternative language codes for transcription. Up to 3 alternatives. If the list has more than 3 elements, the first 3 will be used. Defaults to None, in which ["id-ID", "en-US", "en-GB"] is used.

None
model str

Transcription model name. Defaults to "latest_long".

'latest_long'
timeout int

Timeout for the transcription request. Defaults to 5 minutes.

5 * 60
proxy str | None

The proxy URL to use for the YouTube request. Defaults to None.

None

convert(source, **kwargs) async

Convert audio with Google Cloud Speech-to-Text and normalize into TextResult.

Behavior
  1. Validates source and resolves it from bytes, file path, base64, URL, or YouTube.
  2. Converts audio to mono FLAC for consistent recognition settings.
  3. Uploads audio to GCS when duration exceeds provider inline limits.
  4. Executes long-running recognition and parses response alternatives.
  5. Maps recognition chunks into AudioTranscript segments and returns TextResult.

Parameters:

Name Type Description Default
source str | bytes

Audio source to transcribe. Supported forms: 1. Raw bytes containing audio data. 2. Local audio file path. 3. Base64 encoded audio string. 4. Direct downloadable URL. 5. YouTube URL.

required
**kwargs Any

Additional transcription options accepted by the shared audio-to-text API. This implementation currently does not consume custom keyword arguments in _get_transcripts.

{}

Returns:

Name Type Description
TextResult TextResult

Transcript text with timestamp and language metadata.

ASRBasedOpenAIAudioToTranscript(api_key, model='whisper-1', language=None, prompt=None, temperature=0, timestamp_granularity='segment', proxy=None, skip_conversion_for_formats=None, base_url=None, merge_consecutive_duplicates=False)

Bases: PromptableASRBasedAudioToTranscript

An audio to text converter using OpenAI Whisper.

The ASRBasedOpenAIAudioToTranscript class is responsible for converting audio to text using OpenAI Whisper. It supports various audio sources such as file paths, base64 encoded strings, downloadable audio URLs, and YouTube URLs.

When base_url is omitted, the class operates using the official OpenAI API. When base_url points to a custom OpenAI-compatible endpoint, the API response is checked for an error field and a fallback AudioTranscript is produced when no segments or words are returned.

In both cases:

  • Audio is converted to mono FLAC before upload (unless the format appears in skip_conversion_for_formats).
  • Files exceeding 25 MB are progressively downsampled.

Both modes use the same synchronous openai.OpenAI client.

Attributes:

Name Type Description
client OpenAI

The OpenAI client instance used for API requests.

model str

The identifier of the OpenAI Whisper model to use for transcription.

language str | None

The language of the input audio content.

prompt str | None

The text prompt to guide the model's style or continue a previous audio segment.

temperature float

The sampling temperature to control output randomness.

timestamp_granularity str

The timestamp detail levels to include in the output.

proxy str | None

The proxy URL to use for the YouTube request.

skip_conversion_for_formats list[str]

List of audio formats that should skip mono FLAC conversion.

merge_consecutive_duplicates bool

Whether to merge consecutive duplicated transcripts.

Initialize the ASRBasedOpenAIAudioToTranscript instance.

Parameters:

Name Type Description Default
api_key str

The API key for authentication with OpenAI (or the custom endpoint).

required
model str

The model identifier to use for transcription. Defaults to "whisper-1".

'whisper-1'
language str | None

The language of the input audio content. Defaults to None.

None
prompt PromptBuilder | None

The prompt builder to guide transcription behavior. Defaults to None.

None
temperature float

The sampling temperature to control output randomness. Defaults to 0.

0
timestamp_granularity str

The timestamp detail levels to include in the output. The granularity can be "segment" or "word". When set to "segment", OpenAI will return transcripts divided into segments of speech. When set to "word", it will include word-level timestamps. Defaults to "segment".

'segment'
proxy str | None

The proxy URL to use for the YouTube request. Defaults to None.

None
skip_conversion_for_formats list[str] | None

List of audio formats (e.g., ['mp3', 'wav']) that should skip the mono FLAC conversion process. Formats are case-insensitive and can be specified with or without a leading dot. If None or empty, all audio will be converted to mono FLAC. Defaults to None.

None
base_url str | None

Base URL of a custom OpenAI-compatible Whisper endpoint (e.g. "https://my-whisper-server/v1"). When provided, the class checks for API-level errors, and produces a fallback transcript when the response contains no segments or words. Defaults to None (standard OpenAI API).

None
merge_consecutive_duplicates bool

Whether to merge consecutive transcripts that have the exact same text, combining their time ranges. Defaults to False.

False

convert(source, **kwargs) async

Convert audio to transcript text using OpenAI Whisper-compatible APIs.

Behavior
  1. Validates source and resolves bytes, file path, base64, downloadable URL, or YouTube URL.
  2. Detects audio format and optionally skips FLAC conversion for configured formats.
  3. Converts to mono FLAC and progressively downsamples when payload exceeds size limits.
  4. Sends audio to the transcription endpoint with rendered prompt context.
  5. Parses segments or words responses into AudioTranscript objects.
  6. Optionally merges consecutive duplicate transcript segments.

Parameters:

Name Type Description Default
source str | bytes

Audio source to transcribe. Supported forms: 1. Raw bytes containing audio content. 2. Local file path. 3. Base64 encoded audio string. 4. Direct downloadable URL. 5. YouTube URL.

required
**kwargs Any

Runtime options forwarded through the shared audio-to-text API. Prompt customization should be provided via set_prompt_context before invoking this method.

{}

Returns:

Name Type Description
TextResult TextResult

Transcript text output with structured metadata.

Example
converter = ASRBasedOpenAIAudioToText(api_key="...", model="whisper-1")
converter.set_prompt_context({"domain": "meeting"})
result = await converter.convert("https://example.com/audio.mp3")
print(result.result)

ASRBasedProsaAudioToTranscript(api_key, base_url=None, model='stt-general', polling_interval=MIN_POLLING_INTERVAL)

Bases: ASRBasedAudioToTranscript

An audio to text converter using Prosa STT.

The ASRBasedProsaAudioToTranscript class is responsible for converting audio to text using the Prosa STT API. It supports various audio input formats and can handle audio from local files, base64 encoded strings, or URLs pointing to audio files.

Attributes:

Name Type Description
url str

The URL of the Prosa STT API.

api_key str

The API key for authenticating with the Prosa STT API.

model str

The model to use for the transcription.

polling_interval int

The interval between polling requests to the Prosa STT API.

Initializes a new instance of the ASRBasedProsaAudioToTranscript class.

Parameters:

Name Type Description Default
api_key str

The API key for the Prosa STT API.

required
base_url str | None

The base URL of the Prosa STT API. Defaults to None and falls back to DEFAULT_PROSA_STT_BASE_URL.

None
model str

The model to use for the transcription. Defaults to "stt-general".

'stt-general'
polling_interval int

The interval between polling requests to the Prosa STT API. Defaults to MIN_POLLING_INTERVAL, which is set to 5 seconds.

MIN_POLLING_INTERVAL

polling_interval property writable

Get the polling interval in seconds.

Returns:

Name Type Description
int int

The current polling interval.

convert(source, **kwargs) async

Convert audio using Prosa STT job submission and polling flow.

Behavior
  1. Validates the source and resolves it from bytes, local file, base64, or URL.
  2. Enforces Prosa inline audio duration constraints for bytes/base64 payloads.
  3. Submits transcription requests as base64 payload or URL payload.
  4. Polls job status until completed output is available.
  5. Sorts and deduplicates channel transcripts, then returns unified TextResult.

Parameters:

Name Type Description Default
source str | bytes

Audio source to transcribe. Supported forms: 1. Raw bytes payload (must be <= 55 seconds duration). 2. Local audio file path (must be <= 55 seconds duration). 3. Base64 encoded audio string (must be <= 55 seconds duration). 4. HTTPS audio URL submitted as remote URI.

required
**kwargs Any

Additional transcription options accepted by the shared audio-to-text API. This implementation currently does not consume custom keyword arguments in _get_transcripts.

{}

Returns:

Name Type Description
TextResult TextResult

Transcript text with per-segment timing metadata.

ASRBasedQwenAudioToTranscript(api_key='EMPTY', model='Qwen/Qwen3-ASR-1.7B', base_url='http://localhost:8000/v1', prompt=None, align_text=None, temperature=0.01, max_tokens=1024, timestamp_granularity='segment', segment_gap_threshold=DEFAULT_SEGMENT_GAP_THRESHOLD, proxy=None, timeout=DEFAULT_REQUEST_TIMEOUT)

Bases: PromptableASRBasedAudioToTranscript

An audio to text converter using Qwen3-ASR served via a vLLM-compatible endpoint.

This class sends audio to a vLLM OpenAI-compatible chat completions endpoint as a base64-encoded audio_url data URI. The model returns a structured response containing a <|language|> tag, a <|text|> tag with the full transcription, and a <|timestamps|> tag with word-level timing information.

The timestamp_granularity parameter controls the output format:

  1. "word" — each word is returned as a separate AudioTranscript with its own start and end times.
  2. "segment" — consecutive words are merged into segments based on the time gap between them. A new segment is started whenever the gap between one word's end time and the next word's start time exceeds segment_gap_threshold seconds.

Attributes:

Name Type Description
client AsyncOpenAI

The OpenAI-compatible async client used for API requests.

model str

The model identifier (e.g. "Qwen/Qwen3-ASR-1.7B").

prompt PromptBuilder | None

Optional prompt builder to inject domain knowledge or context.

align_text str | None

The ground truth transcription to use for forced alignment.

temperature float

The sampling temperature for generation.

max_tokens int

The maximum number of tokens to generate.

timestamp_granularity str

The timestamp detail level — "word" or "segment".

segment_gap_threshold float

The maximum gap in seconds between consecutive words before starting a new segment (only used when timestamp_granularity is "segment").

proxy str | None

The proxy URL to use for YouTube requests.

Initialize the ASRBasedQwenAudioToTranscript instance.

Parameters:

Name Type Description Default
api_key str

The API key for authentication with the endpoint.

'EMPTY'
model str

The model identifier. Defaults to "Qwen/Qwen3-ASR-1.7B".

'Qwen/Qwen3-ASR-1.7B'
base_url str

The base URL of the vLLM-compatible endpoint. Defaults to "http://localhost:8000/v1".

'http://localhost:8000/v1'
prompt PromptBuilder | None

Optional prompt builder used to inject domain knowledge or context. Defaults to None.

None
align_text str | None

The ground truth transcription to use for forced alignment. If provided, the server will bypass generation and instead return accurate timestamps for the text. Defaults to None.

None
temperature float

The sampling temperature. Defaults to 0.01.

0.01
max_tokens int

The maximum number of tokens to generate. Defaults to 1024.

1024
timestamp_granularity str

The timestamp detail level. Use "word" for word-level transcripts or "segment" to merge words into segments split by time gaps. Must be one of "word" or "segment". Defaults to "segment".

'segment'
segment_gap_threshold float

The maximum gap in seconds between consecutive words before a new segment is started. Only applies when timestamp_granularity is "segment". Defaults to 2.0.

DEFAULT_SEGMENT_GAP_THRESHOLD
proxy str | None

The proxy URL for YouTube requests. Defaults to None.

None
timeout float

The timeout in seconds for API requests. Defaults to 120.0.

DEFAULT_REQUEST_TIMEOUT

Raises:

Type Description
ValueError

If timestamp_granularity is not "word" or "segment".

convert(source, **kwargs) async

Convert audio to transcript text using Qwen3-ASR via chat completions.

Behavior
  1. Validates input source and resolves audio bytes from supported formats.
  2. Encodes audio bytes into a base64 audio_url data URI payload.
  3. Sends the payload to the vLLM-compatible chat completion endpoint.
  4. Parses tagged response sections for language, full text, and word timestamps.
  5. Returns word-level or gap-grouped segment transcripts based on timestamp_granularity.

Parameters:

Name Type Description Default
source str | bytes

Audio source to transcribe. Supported forms: 1. Raw audio bytes. 2. Local file path. 3. Base64 encoded audio string. 4. Downloadable URL. 5. YouTube URL.

required
**kwargs Any

Runtime options forwarded through the shared audio-to-text API. Prompt customization should be provided via set_prompt_context before invoking this method.

{}

Returns:

Name Type Description
TextResult TextResult

Transcript text result with parsed timing metadata.

Example
converter = ASRBasedQwenAudioToText(
    base_url="http://localhost:8000/v1",
    timestamp_granularity="segment",
)
result = await converter.convert("/tmp/audio.wav")
print(result.result)

ASRBasedSnowflakeAudioToTranscript(session_config=None, timestamp_granularity=None, proxy=None, auto_cleanup_stage_file=True)

Bases: ASRBasedAudioToTranscript

An audio to text converter using Snowflake AI_TRANSCRIBE.

The ASRBasedSnowflakeAudioToTranscript class transcribes audio via Snowflake Cortex using the ai_transcribe Snowpark function. Audio must reside on a Snowflake stage with server-side encryption. When the input is not already a stage path, the converter uploads the resolved audio bytes to a temporary stage before transcription.

When session_config is provided, a Snowpark session is established at init time and reused across invocations. Call close when finished, or use the instance as a context manager (with ASRBasedSnowflakeAudioToTranscript(...) as transcriber:), to release that session deterministically. When omitted, get_active_session() is used at invoke time, suitable for code running inside Snowflake (notebooks, Streamlit); do not call close in that mode.

This class is not thread-safe. Do not invoke convert concurrently on the same instance.

Supported features
  1. Local file, bytes, URL, and YouTube audio sources
  2. Direct transcription from an existing Snowflake stage path
  3. Optional word-level or speaker-level timestamps (see module References [1])
  4. Automatic cleanup of uploaded stage files

Attributes:

Name Type Description
session_config SnowflakeSessionConfig | None

Validated Snowflake connection parameters.

timestamp_granularity str | None

Timestamp detail level forwarded to ai_transcribe. Use "word" for per-word segments or "speaker" for per-speaker turns with speaker_label. When omitted, Snowflake returns one segment without timestamps. See module References [1].

proxy str | None

Proxy URL for resolving YouTube audio sources.

auto_cleanup_stage_file bool

Whether to remove uploaded stage files after transcription.

stage_name str

Name of the temporary Snowflake stage used for uploads.

Examples:

Connect with explicit Snowflake credentials (SnowflakeSessionConfig) and transcribe a local audio file:

transcriber = ASRBasedSnowflakeAudioToTranscript(
    session_config=SnowflakeSessionConfig(
        account="my_account",
        user="my_user",
        password="secret",
        role="ANALYST",
        warehouse="COMPUTE_WH",
    ),
)
transcripts = await transcriber.convert("path/to/audio.ogg")
print(transcripts[0].text)

Prefer a named connection from the Snowflake connections file when you do not want to pass credentials inline:

transcriber = ASRBasedSnowflakeAudioToTranscript(
    session_config={"connection_name": "my_conn"},
)
transcripts = await transcriber.convert("path/to/audio.ogg")
print(transcripts[0].text)

Inside Snowflake (notebook / Streamlit), omit session_config so the active session is used, enable word-level timestamps, and transcribe a stage path:

transcriber = ASRBasedSnowflakeAudioToTranscript(timestamp_granularity="word")
transcripts = await transcriber.convert("@mystage/audio.ogg")
for segment in transcripts:
    print(segment.start_time, segment.text)

Call the async API from synchronous code with syncify:

from gllm_core.concurrency import syncify

transcriber = ASRBasedSnowflakeAudioToTranscript(session_config={"connection_name": "my_conn"})
transcripts = syncify(transcriber.convert)("path/to/audio.wav")

Use as a context manager so the owned Snowpark session is closed automatically:

with ASRBasedSnowflakeAudioToTranscript(session_config={"connection_name": "my_conn"}) as transcriber:
    transcripts = await transcriber.convert("path/to/audio.wav")

Build the same converter through the modality-converter builder:

from gllm_multimodal.builder.modality_converter_builder import build_modality_converter
from gllm_multimodal.constants import Modality, ModalityConverterApproach, ModalityConverterTask

converter = build_modality_converter(
    Modality.AUDIO,
    Modality.TEXT,
    task_type=ModalityConverterTask.TRANSCRIPT,
    approach_type=ModalityConverterApproach.ASR,
    model_id="snowflake/cortex-transcribe",
    session_config={"connection_name": "my_conn"},
)
transcripts = await converter.convert("path/to/audio.wav")

Initialize the ASRBasedSnowflakeAudioToTranscript instance.

Parameters:

Name Type Description Default
session_config SnowflakeSessionConfig | dict[str, Any] | None

Connection parameters for Session.builder.configs(). Accepts a SnowflakeSessionConfig instance or a dictionary of connector parameters, e.g. {"connection_name": "my_conn"} (References [3]) or {"account": "...", "user": "...", "password": "..."} (References [2]). Defaults to None, in which case get_active_session() is called at invoke time.

None
timestamp_granularity str | None

ai_transcribe timestamp granularity. Supported values are "word" (one segment per word with timestamps) and "speaker" (one segment per speaker turn with timestamps and speaker_label). When omitted, the entire file is transcribed as a single segment without timestamps. See module References [1]. Defaults to None.

None
proxy str | None

Proxy URL forwarded when resolving YouTube audio sources. Defaults to None.

None
auto_cleanup_stage_file bool

Whether to remove uploaded files from the stage after transcription completes. Defaults to True.

True

Raises:

Type Description
ValueError

If timestamp_granularity is not a supported value.

__enter__()

Enter a context manager that closes the session on exit.

Returns:

Name Type Description
Self Self

This converter instance.

__exit__(*_args)

Close the Snowpark session when leaving the context manager.

This method is a no-op if the session is already closed.

Parameters:

Name Type Description Default
*_args object

The arguments passed to the context manager.

()

close()

Close the Snowpark session created by this instance.

Only closes sessions opened via session_config. When using get_active_session(), this method is a no-op because Snowflake owns the session lifecycle.

convert(source, **kwargs) async

Convert audio through staged Snowflake AI_TRANSCRIBE orchestration.

Behavior
  1. Validates source input from bytes or string-based references.
  2. Accepts direct stage paths or resolves external sources into audio bytes.
  3. Detects and normalizes supported upload format and validates size limits.
  4. Uploads audio to a temporary Snowflake stage when needed.
  5. Calls ai_transcribe and parses segment or whole-audio response into transcripts.
  6. Removes temporary stage files when auto_cleanup_stage_file is enabled.

Parameters:

Name Type Description Default
source str | bytes

Audio source to transcribe. Supported forms: 1. Existing Snowflake stage file path (e.g. @mystage/audio.ogg). 2. Local file path. 3. Base64 encoded audio string. 4. Downloadable URL. 5. YouTube URL. 6. Raw bytes.

required
**kwargs Any

Additional transcription options accepted by the shared audio-to-text API. This implementation currently does not consume custom keyword arguments in _get_transcripts.

{}

Returns:

Name Type Description
TextResult TextResult

Transcript text with timing metadata from Snowflake output.

Example
converter = ASRBasedSnowflakeAudioToText(
    session_config={"connection_name": "my_conn"},
    timestamp_granularity="word",
)
result = await converter.convert("@mystage/audio.wav")
print(result.result)

AudioTranscriptPreset

Bases: StrEnum

Available presets for LM-based audio transcription.

Attributes:

Name Type Description
DEFAULT

Standard Gemini-based audio transcription.

GEMINI

Alias for the default Gemini transcription preset.

BaseAudioToTranscript()

Bases: BaseAudioToText

Abstract base class for transcript-based audio-to-text conversion.

FetchBasedYouTubeAudioToTranscript(preferred_lang_ids=None, allow_non_preferred_lang_ids=False, allow_auto_generated_transcripts=False, proxy=None)

Bases: TranscriptFetchAudioToTranscript

An audio to text converter using YouTube Transcript API.

The YoutubeTranscriptAudioToText class is responsible for converting audio from YouTube to text using YouTube Transcript API.

Attributes:

Name Type Description
preferred_lang_ids list[str]

The preferred language IDs for the transcript.

allow_non_preferred_lang_ids bool

Whether to allow non-preferred language IDs.

allow_auto_generated_transcripts bool

Whether to allow auto-generated transcripts.

proxy str | None

The proxy URL to use for the YouTube request.

Initialize the FetchBasedYouTubeAudioToTranscript instance.

Parameters:

Name Type Description Default
preferred_lang_ids list[str] | None

The preferred language IDs for the transcript. Defaults to None, in which case ["id", "en"] will be used.

None
allow_non_preferred_lang_ids bool

Whether to allow non-preferred language IDs. Defaults to False.

False
allow_auto_generated_transcripts bool

Whether to allow auto-generated transcripts. Defaults to False.

False
proxy str | None

The proxy URL to use for the YouTube request. Defaults to None.

None

convert(source, **kwargs) async

Convert a YouTube source by fetching transcript tracks and normalizing output.

Behavior
  1. Validates source type and emptiness checks from the shared audio API.
  2. Extracts the video ID from URL patterns or yt_dlp metadata lookup.
  3. Chooses the best transcript language based on preferred languages and fallback flags.
  4. Fetches transcript snippets from YouTube Transcript API.
  5. Normalizes snippets into timestamped AudioTranscript entries and returns TextResult.

Parameters:

Name Type Description Default
source str | bytes

YouTube media reference to transcribe. Expected input: 1. YouTube URL string in watch/embed/shorts/youtu.be format. 2. Bytes are rejected by this implementation and raise ValueError.

required
**kwargs Any

Additional conversion parameters accepted by the shared audio-to-text API. This implementation currently does not require extra keyword arguments.

{}

Returns:

Name Type Description
TextResult TextResult

Combined transcript text with per-snippet timing metadata.

LMBasedAudioToTranscript(lm_request_processor, postprocessors=None)

Bases: BaseAudioToTranscript, UsesLM

Audio transcription implementation using multimodal Language Models.

Prompt rendering is handled by LMRequestProcessor via process(**prompt_context). set_prompt_context / clear_prompt_context cache keyword arguments to be passed as context variables during the next LM invocation. This caching pattern is used internally by HybridVideoToCaption and similar orchestration components.

Initialize the LM-based audio transcription component.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

Language model request processor that supports multimodal inputs and structured transcript output.

required
postprocessors list[Postprocessor] | None

Post-processors applied to transcript segments. Defaults to [normalize_transcript_timestamps].

None

clear_prompt_context()

Remove cached prompt context kwargs.

from_preset(preset_name=AudioTranscriptPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs) classmethod

Initialize the LM-based audio transcription component using preset model configurations.

Parameters:

Name Type Description Default
preset_name AudioTranscriptPreset | str | None

Name of the preset to use.

DEFAULT
lm_invoker_kwargs dict | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
**kwargs Any

Additional kwargs for current-class __init__ parameters. from_lm_components forwards these kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_audio_to_transcript. 2. Final converter constructor parameters: LMBasedAudioToTranscript.

Returns:

Name Type Description
LMBasedAudioToTranscript LMBasedAudioToTranscript

Initialized audio transcription component using the preset model.

Example
from gllm_multimodal.modality_converter.audio_to_text.audio_to_transcript.lm_based import (
    LMBasedAudioToTranscript,
    AudioTranscriptPreset,
)

# Use default preset (Gemini flash-lite)
transcriber = LMBasedAudioToTranscript.from_preset()

# Use explicit Gemini preset with a custom model
transcriber = LMBasedAudioToTranscript.from_preset(
    preset_name=AudioTranscriptPreset.GEMINI,
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await transcriber.convert(source="audio.mp3")
print(result.result)

set_prompt_context(prompt_context=None)

Cache prompt kwargs for the next LMRequestProcessor.process call.

Parameters:

Name Type Description Default
prompt_context dict[str, Any] | None

Context variables for prompt rendering. Defaults to None.

None

LMBasedGeminiAudioToTranscript(lm_request_processor=None, *, api_key=None, model='gemini-3.1-flash-lite', prompt=None, max_retries=3, timeout=300, postprocessors=None, credentials_path=None, credentials_info=None, project_id=None, location='global')

Bases: LMBasedAudioToTranscript

Gemini audio transcription converter.

Prefer LMBasedAudioToTranscript with presets when constructing through the shared LM-based transcript API.

Initialize the Gemini audio transcription wrapper.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor | None

Pre-built request processor. When provided, all other kwargs are ignored and the processor is used directly. This supports the parent class from_lm_components path. Defaults to None.

None
api_key str | None

The API key for Google Gen AI authentication. Defaults to None.

None
model str

The model to use for transcription. Defaults to "gemini-3.1-flash-lite".

'gemini-3.1-flash-lite'
prompt PromptBuilder | None

Prompt source for Gemini transcription. Defaults to None, in which case the default system and user templates are used.

None
max_retries int

Maximum retry attempts for failed requests. Defaults to 3.

3
timeout float | None

Request timeout in seconds. Defaults to 300.

300
postprocessors list[Postprocessor] | None

Post-processors applied to transcript output.

None
credentials_path str | None

Path to a service account credentials JSON file for Google Vertex AI authentication. Defaults to None.

None
credentials_info dict | None

Service account credentials JSON contents for Google Vertex AI authentication. Defaults to None.

None
project_id str | None

The Google Cloud project ID for Vertex AI. Only used when authenticating with service account credentials. Defaults to None, in which case it will be loaded from the credentials.

None
location str

The location of the Google Cloud project for Vertex AI. Only used when authenticating with service account credentials. Defaults to "global".

'global'

Raises:

Type Description
ValueError

If timeout is not None and is less than or equal to 0.

convert(source, **kwargs) async

Convert audio into transcript text through the Gemini LM pipeline.

Behavior
  1. Validates source as str or bytes.
  2. Resolves bytes, local files, and YouTube URLs into an attachment.
  3. Calls LMRequestProcessor.process with prompt context and event emitter support.
  4. Extracts audio_transcripts from LM output and applies configured postprocessors.
  5. Converts transcripts into a unified TextResult payload.

Parameters:

Name Type Description Default
source str | bytes

Audio source to transcribe. Supported forms: 1. Raw bytes containing audio content. 2. Local file path string. 3. YouTube URL string.

required
**kwargs Any

Runtime options forwarded to LM invocation, including: 1. prompt_context (dict[str, Any], optional): Variables used to render prompt templates. 2. event_emitter (Any, optional): Event emitter passed to LMRequestProcessor.process. 3. Any additional implementation-specific options consumed by postprocessors.

{}

Returns:

Name Type Description
TextResult TextResult

Transcript text plus structured transcript metadata.

Example
converter = LMBasedGeminiAudioToText(api_key="...", model="gemini-3.1-flash-lite")
result = await converter.convert(
    source="meeting.mp3",
    prompt_context={
        "target_language": "en",
        "speaker_hints": ["host", "guest"],
        "glossary": {"RAG": "retrieval augmented generation"},
    },
    event_emitter=my_event_emitter,
)
print(result.result)

SnowflakeSessionConfig

Bases: BaseModel

Validated configuration for creating a Snowflake Snowpark session.

All fields are optional individually, but the configuration must provide a connection bootstrap: connection_name, both account and user, or a pre-existing connection object for an open Python connector connection. Additional connector parameters may be supplied via extra fields and are passed to Session.builder.configs().

See module References [2] for snowflake.connector.connect() parameter names (account, user, password, role, warehouse, etc.) and References [3] for connection_name via connections.toml.

Sensitive values such as passwords, tokens, and private keys are automatically redacted from repr(), str(), and to_loggable_dict via redact_sensitive_data. Use to_session_config only when passing credentials to Snowflake, never for logging.

Attributes:

Name Type Description
connection_name str | None

Named connection from a Snowflake connections.toml file. See module References [3].

account str | None

Snowflake account identifier.

user str | None

Snowflake username.

password str | None

Snowflake password.

role str | None

Default role for the session.

warehouse str | None

Default warehouse for the session.

database str | None

Default database for the session.

schema str | None

Default schema for the session.

authenticator str | None

Authentication method (e.g. "externalbrowser").

private_key str | None

PEM-encoded private key for key-pair authentication.

private_key_file str | None

Path to a private key file.

private_key_file_pwd str | None

Passphrase for the private key file.

token str | None

OAuth or programmatic access token.

host str | None

Custom host for the Snowflake account.

port int | None

Custom port for the Snowflake account.

protocol str | None

Connection protocol (e.g. "https").

connection Any | None

An existing snowflake.connector.Connection to reuse when creating the Snowpark session. See module References [2].

__repr__()

Return a redacted string representation safe for logging.

Returns:

Name Type Description
str str

A redacted string representation of the configuration.

__str__()

Return a redacted string representation safe for logging.

Returns:

Name Type Description
str str

A redacted string representation of the configuration.

to_loggable_dict()

Return a copy of the configuration with sensitive values redacted for logging.

Returns:

Type Description
dict[str, Any]

dict[str, Any]: Connection parameters safe to include in logs or error messages.

to_session_config()

Convert the configuration to a Snowpark Session.builder.configs() dictionary.

The returned mapping uses the same keys accepted by snowflake.connector.connect() and Snowpark session creation. See module References [2].

Returns:

Type Description
dict[str, Any]

dict[str, Any]: Connection parameters with None values excluded.

Note
Credential security

This method includes credential values and must not be logged directly. Use to_loggable_dict for logging or debugging output.

validate_connection_params()

Ensure the configuration contains enough information to connect.

Returns:

Name Type Description
Self Self

The validated configuration instance.

Raises:

Type Description
ValueError

If no connection bootstrap parameters are provided.

TranscriptFetchAudioToTranscript()

Bases: BaseAudioToTranscript

Abstract base class for integrations that fetch existing transcripts.