Overview
Image-to-caption modules for generating descriptive captions from images.
This package provides concrete implementations for captioning images using language models or preset configurations.
Submodules
output_formatter-- Caption output formatting (free-text and structured).
Exported Classes
LMBasedImageToCaption-- LM-based image captioning.ImageCaptionPreset-- Preset-based image captioning.
ImageCaptionPreset
Bases: StrEnum
Available presets for image captioning.
Attributes:
| Name | Type | Description |
|---|---|---|
DEFAULT |
Standard image captioning. |
|
STRUCTURED |
Structured data extraction from images. |
|
DJARUM_ENG_SVLM |
Specialized maintenance document analysis. |
|
KEYFRAME_CAPTIONING |
Video keyframe captioning with chronological context. |
LMBasedImageToCaption(lm_request_processor, formatter=DEFAULT_FORMATTER)
Bases: BaseImageToCaption, UsesLM
Image captioning implementation using Language Models.
This class implements the ImageToCaption interface using LMs for generating natural language captions.
Usage
Use from_preset to easily instantiate with predefined configuration:
from gllm_multimodal.modality_converter.image_to_text.image_to_caption.lm_based_image_to_caption import LMBasedImageToCaption
converter = LMBasedImageToCaption.from_preset()
Initialize the LM based image captioning component.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lm_request_processor
|
LMRequestProcessor
|
Language model request processor instance that supports multimodal inputs. |
required |
formatter
|
BaseCaptionOutputFormatter
|
The formatter to use for formatting the output. |
DEFAULT_FORMATTER
|
convert(source, **kwargs)
async
Convert image into caption text using the LM captioning pipeline.
Behavior
- Validates and loads media bytes from path, URL, base64, or bytes source.
- Optionally extracts image metadata and builds the
Captioncontext payload. - Builds prompt params from supported prompt keys and caption fields.
- Sends image plus multimodal context attachments to the LM request processor.
- Formats the generated caption using the configured output formatter.
- Returns formatted caption text and metadata as
TextResult.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | bytes
|
Image source to caption. Supported forms: 1. Raw image bytes. 2. Local file path. 3. URL string. 4. Base64 encoded image string. |
required |
**kwargs
|
Any
|
See Other Parameters below. |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
number_of_captions |
int
|
Number of captions to generate. |
text_one_liner |
str
|
Short one-line context/title. |
text_context |
str
|
Additional descriptive context. |
domain_knowledge |
str
|
Domain hints for caption quality. |
multimodal_context |
list[Attachment | str]
|
Additional context attachments. |
use_metadata |
bool
|
Whether image metadata is extracted and injected. |
formatter_kwargs |
dict[str, Any]
|
Extra kwargs for formatter output shaping. |
event_emitter |
Any
|
Event emitter passed to LM request processing. |
Returns:
| Name | Type | Description |
|---|---|---|
TextResult |
TextResult
|
Formatted caption result with structured metadata. |
Example
from gllm_inference.schema import Attachment
converter = LMBasedImageToCaption.from_preset()
result = await converter.convert(
source="diagram.png",
number_of_captions=3,
text_one_liner="System architecture diagram",
domain_knowledge="cloud networking",
multimodal_context=[
"focus on request flow",
Attachment.from_file("legend.png"),
],
formatter_kwargs={"include_bullets": True, "max_sentences": 6},
use_metadata=True,
)
print(result.result)
from_preset(preset_name=ImageCaptionPreset.DEFAULT, lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs)
classmethod
Initialize the LM based image captioning component using preset model configurations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preset_name
|
ImageCaptionPreset | str | None
|
Preset name forwarded to
image-caption preset registry. Defaults to |
DEFAULT
|
lm_invoker_kwargs
|
dict | None
|
Keyword arguments passed to LM invoker
creation in the preset factory. Defaults to None. Valid keys follow
[ |
None
|
prompt_builder_kwargs
|
dict | None
|
Keyword arguments passed to prompt
builder creation in the preset factory. Defaults to None. Valid keys follow
|
None
|
**kwargs
|
Any
|
Additional kwargs for current-class |
{}
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Any additional kwargs forwarded to |
Returns:
| Name | Type | Description |
|---|---|---|
LMBasedImageToCaption |
LMBasedImageToCaption
|
Initialized image captioning component using preset model. |
Example
from gllm_multimodal.modality_converter.image_to_text.image_to_caption import (
LMBasedImageToCaption,
ImageCaptionPreset,
)
# Use default preset (Gemini flash-lite)
captioner = LMBasedImageToCaption.from_preset()
# Use structured output preset with a custom model
captioner = LMBasedImageToCaption.from_preset(
preset_name=ImageCaptionPreset.STRUCTURED,
lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)
result = await captioner.convert(source="diagram.png")
print(result.result)