Skip to content

Overview

Image-to-mermaid modules for generating mermaid syntax from chart images.

LMBasedImageToMermaid(lm_request_processor)

Bases: BaseImageToMermaid, UsesLM

LM-based implementation for converting an image into Mermaid diagram syntax.

This class leverages a language model (LM) pipeline to generate structured Mermaid syntax from image inputs and optional metadata. It uses prompt builders, LM invokers, and output parsers defined via a preset system to streamline model usage.

Inherits

BaseImageToMermaid: Base class defining the image-to-mermaid interface. UsesLM: Mixin providing shared logic for components using language models.

Attributes:

Name Type Description
lm_request_processor LMRequestProcessor

Handles prompt creation, LM invocation, and output parsing. Core component that orchestrates the image-to-mermaid pipeline.

Initializes the LMBasedImageToMermaid instance with a language model request processor.

Parameters:

Name Type Description Default
lm_request_processor LMRequestProcessor

The processor handling prompt creation, LM invocation, and output parsing.

required

convert(source, **kwargs) async

Convert image input into Mermaid syntax using the LM conversion flow.

Behavior
  1. Validates and loads image bytes from supported source formats.
  2. Builds Mermaid metadata payload from method keyword arguments.
  3. Filters prompt params to keys recognized by the prompt template.
  4. Sends the image as attachment to the LM request processor.
  5. Returns Mermaid syntax text and metadata wrapped in TextResult.

Parameters:

Name Type Description Default
source str | bytes

Image source to convert. Supported forms: 1. Raw image bytes. 2. Local file path. 3. URL string. 4. Base64 encoded image string.

required
**kwargs Any

Mermaid generation options forwarded into metadata and LM calls, including: 1. title (str, optional): Diagram title context. 2. description (str, optional): Additional scene or flow context. 3. event_emitter (Any, optional): Event emitter passed to LM request processing. 4. Any Mermaid schema fields recognized by Mermaid model and prompt template.

{}

Returns:

Name Type Description
TextResult TextResult

Mermaid syntax output with associated metadata.

Example
converter = LMBasedImageToMermaid.from_preset()
result = await converter.convert(
    source="workflow.png",
    title="Ticket Escalation Flow",
    description="Extract decision points and handoff states",
    diagram_type="flowchart",
    direction="LR",
    event_emitter=my_event_emitter,
)
print(result.result)

from_preset(preset_name='default', lm_invoker_kwargs=None, prompt_builder_kwargs=None, **kwargs) classmethod

Constructs an LMBasedImageToMermaid instance using a named preset configuration.

Parameters:

Name Type Description Default
preset_name str | None

Name of the predefined preset configuration to use. Defaults to "default".

'default'
lm_invoker_kwargs dict | None

Keyword arguments passed to LM invoker creation in the preset factory. Defaults to None. Valid keys follow [build_lm_invoker][gllm_inference.lm_invoker.build_lm_invoker.build_lm_invoker].

None
prompt_builder_kwargs dict | None

Keyword arguments passed to prompt builder creation in the preset factory. Defaults to None. Valid keys follow PromptBuilder.

None
**kwargs Any

Additional kwargs for current-class __init__ parameters. from_lm_components forwards these kwargs when instantiating cls(...).

{}

Other Parameters:

Name Type Description
**kwargs Any

Any additional kwargs forwarded to from_lm_components. Valid kwargs references: 1. Preset composition and consumed kwargs: get_preset_image_to_mermaid. 2. Final converter constructor parameters: LMBasedImageToMermaid.

Returns:

Name Type Description
LMBasedImageToMermaid LMBasedImageToMermaid

An instance initialized with the preset's components.

Example
from gllm_multimodal.modality_converter.image_to_text.image_to_mermaid import (
    LMBasedImageToMermaid,
)

# Use default preset (Gemini flash-lite)
converter = LMBasedImageToMermaid.from_preset()

# Override model via lm_invoker_kwargs
converter = LMBasedImageToMermaid.from_preset(
    lm_invoker_kwargs={"model_id": "google/gemini-3.1-flash-lite"},
)

result = await converter.convert(source="architecture.png")
print(result.result)