Ocr Result
Schema for OCR output results in Gen AI applications.
This module defines the structured output of an OCR operation, carrying both the raw extracted text and optional enriched fields populated by specialized OCR engines.
OcrResult
Bases: BaseModel
Structured result of an OCR operation.
Attributes:
| Name | Type | Description |
|---|---|---|
text |
str
|
Full concatenated text extracted from the document. This value is mirrored in TextResult.result for API consistency. |
lines |
list[str]
|
Engine-native line units when available (e.g. from specialized OCR backends). LM-based implementations populate this only when bounding boxes are requested, with one entry per text region; otherwise they leave it empty. Defaults to an empty list. |
page_count |
int
|
Number of pages processed. Populated by engines that support multi-page documents (e.g., Azure Document Intelligence). Defaults to 1. |
detected_objects |
list[DetectedObject]
|
Localized regions, such as
text lines and tables, populated by engines that return layout
information. Text lines come first in engine reading order,
followed by tables. Filter by |
image_width |
int | None
|
Width in pixels of the source image that
the detected object polygons are measured against. Use it with
|
image_height |
int | None
|
Height in pixels of the source image that the detected object polygons are measured against. None when no detected objects are returned. Defaults to None. |