Audio To Text
Defines a base class for audio to text converter used in Gen AI applications.
Quick Start
# Assume `converter` is an instance of a BaseAudioToText subclass
# Returns a TextResult object with combined text and metadata
result = await converter.convert(source="path/to/audio.mp3")
BaseAudioToText()
Bases: BaseModalityConverter
A base class for audio to text used in Gen AI applications.
Extends BaseModalityConverter
so audio converters participate in the shared modality-converter hierarchy.
The primary convert API returns
TextResult, and subclasses implement _convert to produce TextResult.
convert(source, **kwargs)
async
Execute the audio to text conversion process.
This method validates the input parameters and delegates to _convert
to perform the actual conversion.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
The source of the audio to convert to text. |
required |
**kwargs
|
Any
|
Additional conversion parameters. Implementations may
accept keyword arguments such as |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
Combined transcript text with structured segment metadata. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
TypeError
|
If |