Skip to content

Ocr Result

Schema for OCR output results in Gen AI applications.

This module defines the structured output of an OCR operation, carrying both the raw extracted text and optional enriched fields populated by specialized OCR engines.

OcrResult

Bases: BaseModel

Structured result of an OCR operation.

Attributes:

Name Type Description
text str

Full concatenated text extracted from the document. This value is mirrored in TextResult.result for API consistency.

lines list[str]

Engine-native line units when available (e.g. from specialized OCR backends). LM-based implementations populate this only when bounding boxes are requested, with one entry per text region; otherwise they leave it empty. Defaults to an empty list.

page_count int

Number of pages processed. Populated by engines that support multi-page documents (e.g., Azure Document Intelligence). Defaults to 1.

detected_objects list[DetectedObject]

Localized regions, such as text lines and tables, populated by engines that return layout information. Text lines come first in engine reading order, followed by tables. Filter by DetectedObject.class_name to select a region kind. Defaults to an empty list.

image_width int | None

Width in pixels of the source image that the detected object polygons are measured against. Use it with image_height to rescale the polygons for a resized image. None when no detected objects are returned. Defaults to None.

image_height int | None

Height in pixels of the source image that the detected object polygons are measured against. None when no detected objects are returned. Defaults to None.