Converts raw document files to markdown text and optionally extracts structured entities.
extract_operatorThe operator follows hexagonal architecture principles, separating business logic from infrastructure concerns:
TextExtractionMode, EntityExtractionMode, extraction requests/results)TextExtractionPort, EntityExtractionPort)This architecture enables:
{
"type": "extract_operator",
"name": "extract_documents",
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {}
},
"entity_extraction": {
"provider": "none"
}
},
"depends_on": ["ingest_documents"]
}
Standard document extraction using the Docling library locally. Supports optional VLM (Vision-Language Model) pipeline for enhanced extraction and ASR (Automatic Speech Recognition) pipeline for audio/video processing.
Basic Configuration:
{
"max_workers": 4,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
}
},
"entity_extraction": {
"provider": "none"
}
}
VLM Pipeline Configuration:
Enable VLM pipeline for enhanced extraction with vision-language models:
{
"max_workers": 1,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"vlm_pipeline": {
"preset": "granite_docling",
"engine": "transformers",
"engine_options": {}
}
}
},
"entity_extraction": {
"provider": "none"
}
}
GPU Acceleration Configuration:
Enable GPU-accelerated standard pipeline processing for PDF and image documents. The adapter builds one DocumentConverter at initialization and reuses it across all documents — model weights are loaded onto the GPU once per adapter execution.
Note: Cannot be combined with
vlm_pipeline. Requiresmax_workers: 1anduse_processes: false.
When device is omitted, the best available GPU is auto-detected at runtime via torch (CUDA → MPS → XPU). Specify device explicitly to pin a particular GPU.
{
"max_workers": 1,
"use_processes": false,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"standard_pipeline": {
"accelerator": {
"num_threads": 6
}
}
}
},
"entity_extraction": {
"provider": "none"
}
}
Or with an explicit device:
{
"max_workers": 1,
"use_processes": false,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"standard_pipeline": {
"accelerator": {
"device": "cuda",
"num_threads": 6
}
}
}
},
"entity_extraction": {
"provider": "none"
}
}
Supported GPU Devices:
mps — Apple Metal Performance Shaders (Apple Silicon: M1/M2/M3/M4)cuda — NVIDIA CUDA (automatic device selection)cuda:<index> — NVIDIA CUDA on a specific device (e.g. cuda:0, cuda:1)xpu — Intel XPUSupported VLM Engines:
transformers: Local inference using Transformers librarymlx: Local inference optimized for macOS (Apple Silicon)api_ollama: Ollama APIapi_openai: OpenAI APIapi_watsonx: IBM watsonx.ai APIapi_lmstudio: LM Studio APIapi: Generic API endpointVLM Pipeline Configuration Examples:
Ollama:
{
"text_extraction": {
"provider": "docling_library",
"provider_config": {
"vlm_pipeline": {
"preset": "granite_docling",
"engine": "api_ollama",
"engine_options": {
"api_base": "http://localhost:11434",
"model_id": "ibm/granite-docling:258m"
}
}
}
}
}
OpenAI:
{
"text_extraction": {
"provider": "docling_library",
"provider_config": {
"vlm_pipeline": {
"preset": "qwen",
"engine": "api_openai",
"engine_options": {
"api_base": "https://api.openai.com",
"model_id": "gpt-4-vision-preview",
"api_key": "<your-api-key>"
}
}
}
}
}
ASR Pipeline Configuration:
Enable ASR pipeline for audio and video file transcription:
{
"max_workers": 2,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"asr_pipeline": {
"model_id": "whisper_turbo"
}
}
},
"entity_extraction": {
"provider": "none"
}
}
Use Cases:
Sample Flows:
sample_flows/quickstart/basic_ingest_extract.jsonsample_flows/quickstart/complete_pipeline_ollama.jsonsample_flows/use_cases/audio_video_extraction.jsonREST API-based extraction using the Docling-Serve service for scalable, production-ready document processing.
Configuration:
{
"text_extraction": {
"provider": "docling_serve",
"provider_config": {
"base_url": "http://localhost:5001",
"timeout": 300,
"poll_interval": 2,
"max_retries": 3,
"ocr": {
"enabled": true,
"engine": "easyocr",
"engine_options": {
"lang": ["en"]
}
},
"pdf_backend": "dlparse_v4",
"table_mode": "accurate",
"image_export_mode": "embedded"
}
},
"entity_extraction": {
"provider": "none"
}
}
File Handling:
.txt files are processed locally using basic text extraction (not sent to Docling Serve)MIME Type Mapping:
| Extension | MIME Type |
|---|---|
.html, .htm |
text/html |
.md |
text/markdown |
.txt |
text/plain |
.pdf |
application/pdf |
.docx |
application/vnd.openxmlformats-officedocument.wordprocessingml.document |
.doc |
application/msword |
.xlsx |
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet |
.xls |
application/vnd.ms-excel |
.pptx |
application/vnd.openxmlformats-officedocument.presentationml.presentation |
.ppt |
application/vnd.ms-powerpoint |
| Other | application/octet-stream |
Prerequisites:
# Start docling-serve locally
docker run -p 5001:5001 ds4sd/docling-serve:latest
# Or use docker-compose
docker-compose -f docker-compose.docling-serve.yml up -d
Features:
dlparse_v4: Latest Docling parser (recommended)dlparse_v3: Legacy Docling parserpypdfium2: PyPDFium2 backendaccurate: High accuracy (slower)fast: Fast processing (less accurate)placeholder: Replace images with a placeholder (default)embedded: Embed images in outputreferenced: Reference images by pathnone: Skip image exportUse Cases:
Sample Flow: sample_flows/quickstart/complete_pipeline_ollama.json
Entity extraction can be combined with any text extraction provider to extract structured data from the extracted text.
No entity extraction is performed. Only text extraction is executed.
Configuration:
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "none"
}
}
Vision-Language Model (VLM) based entity extraction using Docling’s VLM pipeline for structured data extraction from documents.
Basic Configuration:
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "docling",
"custom_schema": {
"invoice_number": "string",
"invoice_date": "string",
"total_amount": "number",
"line_items": [
{
"description": "string",
"quantity": "number",
"unit_price": "number"
}
]
}
}
}
Custom Model Configuration:
Users can configure custom inline VLM models for entity extraction using the vlm_pipeline parameter. Only inline models (HuggingFace) are supported as DocumentExtractor does not support remote API endpoints.
Inline Model (HuggingFace with Transformers):
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "docling",
"provider_config": {
"vlm_pipeline": {
"model_type": "inline",
"inline_model": {
"repo_id": "numind/NuExtract-2.0-2B",
"inference_framework": "transformers",
"scale": 2.0,
"temperature": 0.0,
"max_new_tokens": 4096,
"load_in_8bit": true,
"torch_dtype": "bfloat16"
}
}
},
"custom_schema": {
"invoice_number": "string",
"total_amount": "number"
}
}
}
Note: API model configuration is not supported. For API-based entity extraction, use entity_extraction.provider: "litellm" instead.
Supported Model Types:
Supported Backends (for inline models):
transformers: HuggingFace Transformers libraryvllm: vLLM inference enginemlx: Apple MLX framework (macOS only)Use Cases:
Sample Flow: sample_flows/quickstart/complete_pipeline_ollama.json
Multi-provider LLM extraction using LiteLLM for accessing 100+ LLM providers (OpenAI, Anthropic, Cohere, Ollama, etc.).
Configuration (OpenAI):
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/gpt-3.5-turbo",
"api_key": "your-api-key", # pragma: allowlist secret
"api_base": "https://api.openai.com/v1",
"temperature": 0.0,
"max_tokens": 2000
},
"custom_schema": {
"invoice_number": "string",
"total_amount": "number"
}
}
}
Configuration (Ollama via LiteLLM):
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.2",
"api_base": "http://localhost:11434/v1",
"api_key": "<ollama_key>",
"temperature": 0.0,
"max_tokens": 4096
},
"custom_schema": {
"invoice_number": "string",
"total_amount": "number"
}
}
}
Configuration (Remote vLLM with Streaming & Extended Timeout):
For high-concurrency scenarios with remote vLLM clusters processing large documents, use streaming and extended timeouts to prevent connection drops:
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/granite4:latest",
"api_base": "https://your-vllm-route/v1",
"api_key": "YOUR_API_KEY", # pragma: allowlist secret
"temperature": 0.0,
"max_tokens": 5000,
"stream": true,
"timeout": 1800
},
"custom_schema": {
"invoice_number": "string",
"total_amount": "number"
}
}
}
Advanced Provider Configuration Parameters:
stream (boolean, default: false): Enable HTTP chunked transfer encoding to keep connections alive during long-running requests. Recommended for remote vLLM clusters processing large documents.timeout (integer, default: 60): HTTP client read timeout in seconds. Set to 1800 (30 minutes) for large documents that require extended generation time.Why Streaming & Extended Timeout?
During high-concurrency scalability testing with remote vLLM clusters, connection issues were identified:
stream=false, vLLM waits until entire generation completes (150+ seconds for large documents) before sending response, causing connections to be dropped mid-generation.Solution: Combining stream=true with timeout=1800 ensures:
Supported Providers:
Prerequisites for Ollama:
http://localhost:11434ollama pull llama3.2Use Cases:
Sample Flows:
IBM WatsonX.ai LLM-based entity extraction for enterprise deployments.
Configuration:
{
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "watsonx",
"provider_config": {
"model_id": "ibm/granite-13b-chat-v2",
"api_key": "${WATSONX_API_KEY}",
"container_id": "${WATSONX_CONTAINER_ID}",
"api_base": "https://us-south.ml.cloud.ibm.com",
"container_kind": "project",
"temperature": 0.0,
"max_tokens": 2000
},
"custom_schema": {
"invoice_number": "string",
"total_amount": "number"
}
}
}
Environment Variables:
WATSONX_API_KEY: WatsonX API key (required)WATSONX_CONTAINER_ID: WatsonX project or space ID (required)WATSONX_API_BASE_URL: WatsonX API base URL (optional, defaults to us-south)WATSONX_CONTAINER_KIND: Container type - “project” or “space” (optional, defaults to “project”)Use Cases:
Sample Flow: sample_flows/operators/entity_extraction_watsonx.json
| Parameter | Type | Default | Description |
|---|---|---|---|
text_extraction.provider |
string | "docling_library" |
Text extraction strategy: "docling_library" or "docling_serve" |
text_extraction.doc_column |
string | "doc_content" |
Column name for storing extracted content |
text_extraction.provider_config.additional_formats |
array | [] |
Additional output formats (e.g., ["html", "markdown"]) |
max_workers |
integer | auto | Maximum number of parallel workers (auto-detected based on CPU) |
use_processes |
boolean | false |
Use ProcessPoolExecutor instead of ThreadPoolExecutor (top-level parameter) |
entity_extraction.provider |
string | "none" |
Entity extraction strategy: "litellm" (includes Ollama via openai/ prefix), "docling", "watsonx", or "none" |
entity_extraction.custom_schema |
object | {} |
Schema dictionary for structured extraction (top-level entity_extraction parameter) |
entity_extraction.expand_extracted_data |
boolean | false |
Expand entity data JSON into individual columns (entity extraction only) |
| Parameter | Type | Default | Description |
|---|---|---|---|
text_extraction.provider_config.vlm_pipeline |
object | null |
VLM (Vision-Language Model) pipeline configuration object. Provide empty dict {} to enable with defaults, or omit to disable. |
text_extraction.provider_config.vlm_pipeline.preset |
string | "granite_docling" |
VLM preset name. Valid presets: smoldocling, granite_docling, deepseek_ocr, granite_vision, pixtral, got_ocr, phi4, qwen, nanonets_ocr2, gemma_12b, gemma_27b, dolphin, glm_ocr, lightonocr, falcon_ocr |
text_extraction.provider_config.vlm_pipeline.engine |
string | "api_ollama" |
VLM engine type. Valid engines: api_ollama, api_openai, api_watsonx, api_lmstudio, api (generic), transformers (local), mlx (macOS) |
text_extraction.provider_config.vlm_pipeline.engine_options |
object | {} |
Engine-specific options (api_base, model_id, etc.) |
text_extraction.provider_config.asr_pipeline |
object | null |
ASR (Automatic Speech Recognition) pipeline configuration object. Provide empty dict {} to enable with defaults, or omit to disable. |
text_extraction.provider_config.asr_pipeline.model_id |
string | "whisper_turbo" |
ASR model name. Valid values: whisper_tiny, whisper_small, whisper_medium, whisper_base, whisper_large, whisper_turbo, and their _mlx/_native variants (e.g., whisper_tiny_mlx, whisper_tiny_native) |
text_extraction.provider_config.standard_pipeline |
object | null |
Standard pipeline acceleration block. Omit entirely to use default Docling behaviour. |
text_extraction.provider_config.standard_pipeline.accelerator |
object | null |
GPU accelerator options. When present, one DocumentConverter is built at init and reused. Requires max_workers: 1 and use_processes: false. Cannot be combined with vlm_pipeline. |
text_extraction.provider_config.standard_pipeline.accelerator.device |
string | auto-detected | GPU device. Accepted: mps, cuda, cuda:<index> (e.g. cuda:0), xpu. When omitted, best available device is auto-detected via torch (CUDA → MPS → XPU). Validated at runtime via torch backends. |
text_extraction.provider_config.standard_pipeline.accelerator.num_threads |
int | 4 |
CPU-side pipeline thread count. Must be a positive integer (booleans rejected). |
| Parameter | Type | Default | Description |
|---|---|---|---|
text_extraction.provider_config.base_url |
string | "http://localhost:5001" |
Docling Serve API endpoint URL |
text_extraction.provider_config.api_key |
string | null |
Optional API key for authentication |
text_extraction.provider_config.timeout |
integer | 300 |
Request timeout in seconds |
text_extraction.provider_config.poll_interval |
integer | 2 |
Polling interval in seconds |
text_extraction.provider_config.max_retries |
integer | 3 |
Maximum retry attempts |
text_extraction.provider_config.additional_formats |
array | [] |
Additional output formats beyond markdown (e.g., ["html", "json", "text", "doctags", "doclang"]) |
text_extraction.provider_config.ocr |
object | null |
Canonical OCR config block (see OCR Configuration section below) |
text_extraction.provider_config.pdf_backend |
string | "dlparse_v2" |
PDF backend: "dlparse_v4", "dlparse_v3", or "pypdfium2" |
text_extraction.provider_config.table_mode |
string | null |
Table extraction mode: "accurate" or "fast". Not set by default; the Docling Serve instance uses its own default. |
text_extraction.provider_config.image_export_mode |
string | "placeholder" |
Image export mode: "placeholder", "embedded", "referenced", or "none" |
Both docling_library and docling_serve providers accept an ocr block inside provider_config. Omitting the block entirely uses docling-pipelines defaults (OCR enabled, RapidOCR engine, default mode).
"provider_config": {
"ocr": {
"enabled": true,
"engine": "easyocr",
"mode": "pdf_aware_layout_regions",
"engine_options": {
"lang": ["en", "fr"],
"use_gpu": false,
"confidence_threshold": 0.5
}
}
}
| Field | Type | Default | Description |
|---|---|---|---|
ocr.enabled |
bool | true |
Enable OCR processing. Set false to skip OCR entirely. |
ocr.engine |
string | "rapidocr" |
OCR engine. "rapidocr" is the default. |
ocr.mode |
string | "default" |
OCR scanning mode. "pdf_aware_layout_regions" is most efficient for mixed PDFs. |
ocr.engine_options |
object | null |
Engine-specific parameters (see table below). |
| Engine | Value | Install extra | Notes |
|---|---|---|---|
| Auto-select | "auto" |
none | Optional runtime selection when you explicitly want Docling to choose an installed backend. |
| EasyOCR | "easyocr" |
easyocr |
Cross-platform alternative; 80+ languages. |
| Tesseract (Python bindings) | "tesserocr" |
tesserocr |
Linux-focused local OCR option; 3-letter ISO 639-2 lang codes; PSM control. |
| Tesseract (CLI) | "tesseract" |
tesseract binary |
Same as above via CLI; portable. |
| RapidOCR | "rapidocr" |
included in base install | Default OCR engine for new users; PaddlePaddle-based; multiple backends. |
| macOS Vision | "ocrmac" |
ocrmac |
Optional macOS-specific alternative; Apple-only. |
| KServe V2 | "kserve_v2_ocr" |
custom | Remote KServe/Triton inference server. |
| Nemotron OCR | "nemotron-ocr" |
custom | NVIDIA Nemotron v2. |
RapidOCR is included in the default PyPI install, and ocr.engine: "rapidocr" is the default runtime behaviour for first-run OCR.
| OS / environment | Default install behaviour | Notes |
|---|---|---|
| macOS | pip install docling-pipelines |
RapidOCR works out of the box. Install docling-pipelines[ocrmac] only if you want Apple’s Vision OCR explicitly. |
| Linux | pip install docling-pipelines |
RapidOCR works out of the box. Install docling-pipelines[tesserocr] only if you want Tesseract bindings explicitly. |
| Cross-platform / unsure | pip install docling-pipelines |
Recommended for new users. First OCR run should work without extra setup. |
| Mode | Value | Behaviour |
|---|---|---|
| Default | "default" |
Docling picks automatically |
| Full page | "full_page" |
Scan entire page as one region |
| Layout regions | "layout_regions" |
Scan only layout-detected text regions |
| PDF-aware layout regions | "pdf_aware_layout_regions" |
Skip regions with an existing PDF text layer — most efficient for mixed PDFs |
engine_options Reference| Engine | Key | Type | Notes |
|---|---|---|---|
easyocr |
lang |
list[str] | ISO 639-1 codes, e.g. ["en", "fr"] |
easyocr |
use_gpu |
bool/null | null = auto-detect |
easyocr |
confidence_threshold |
float | 0.0–1.0 |
tesserocr / tesseract |
lang |
list[str] | 3-letter ISO 639-2, e.g. ["eng", "fra"] |
tesserocr / tesseract |
psm |
int | Page segmentation mode 0–13 |
tesserocr / tesseract |
path |
str/null | Tessdata directory |
rapidocr |
lang |
list[str] | Language list |
rapidocr |
backend |
string | "onnxruntime", "openvino", "paddle", "torch" |
rapidocr |
text_score |
float | Detection confidence threshold |
ocrmac |
lang |
list[str] | Locale format, e.g. ["en-US"] |
ocrmac |
recognition |
string | "accurate" or "fast" |
| Parameter | Type | Default | Description |
|---|---|---|---|
entity_extraction.provider_config.vlm_pipeline |
object | null |
Custom VLM model configuration (see Custom Model Configuration section above for full details) |
vlm_pipeline Structure:
Only "inline" models (HuggingFace) are supported. API model types are not supported by Docling’s DocumentExtractor; use entity_extraction.provider: "litellm" for API-based extraction instead.
Example (all fields shown with their defaults; only repo_id is required):
{
"model_type": "inline",
"inline_model": {
"repo_id": "numind/NuExtract-2.0-2B",
"inference_framework": "transformers",
"scale": 2.0,
"temperature": 0.0,
"max_new_tokens": 4096,
"load_in_8bit": true,
"torch_dtype": "bfloat16",
"prompt": "",
"response_format": "markdown"
}
}
inline_model field |
Type | Default | Description |
|---|---|---|---|
repo_id |
string | required | HuggingFace repository ID (e.g., "numind/NuExtract-2.0-2B") |
inference_framework |
string | "transformers" |
Inference backend: "transformers", "vllm", or "mlx" |
scale |
float | 2.0 |
Image scale factor for rendering |
temperature |
float | 0.0 |
Sampling temperature |
max_new_tokens |
integer | 4096 |
Maximum tokens to generate |
load_in_8bit |
boolean | false |
Load model in 8-bit quantization |
torch_dtype |
string | "bfloat16" |
Torch data type: "bfloat16", "float16", "float32" |
prompt |
string | "" |
Custom prompt override (leave empty to use model default) |
response_format |
string | "markdown" |
Expected response format from the model |
All parameters are nested under entity_extraction.provider_config:
| Parameter | Type | Default | Description |
|---|---|---|---|
model_id |
string | "gpt-3.5-turbo" |
LLM model identifier. Must include provider prefix when using LiteLLM (e.g., openai/gpt-4, openai/llama3.2 for Ollama, anthropic/claude-3-opus) |
temperature |
float | 0.0 |
Sampling temperature |
max_tokens |
integer | 2000 |
Maximum response tokens |
api_key |
string | null |
API key for the provider. For Ollama, can be any value |
api_base |
string | null |
API base URL. For Ollama, set to http://localhost:11434/v1 |
stream |
boolean | false |
Enable HTTP chunked transfer encoding. For remote vLLM with large documents, use stream: true and timeout: 1800 |
timeout |
integer | 60 |
HTTP client read timeout in seconds. For remote vLLM with large documents, use stream: true and timeout: 1800 |
All parameters are nested under entity_extraction.provider_config:
| Parameter | Type | Default | Description |
|---|---|---|---|
model_id |
string | "ibm/granite-13b-chat-v2" |
WatsonX model identifier |
temperature |
float | 0.0 |
Sampling temperature |
max_tokens |
integer | 2000 |
Maximum response tokens |
api_key |
string | required | WatsonX API key |
container_id |
string | required | WatsonX project or space ID |
api_base |
string | "https://us-south.ml.cloud.ibm.com" |
WatsonX API base URL (optional) |
container_kind |
string | "project" |
Container type: “project” or “space” (optional) |
| Column | Type | Required | Description |
|---|---|---|---|
id |
string | Yes | Document identifier |
name |
string | Yes | Document name/filename |
path |
string | Yes | Document file path |
document_type |
string | No | Document type for template-based entity extraction |
All input columns are preserved. The operator appends:
| Column | PyArrow Type | Description |
|---|---|---|
doc_content |
string |
Extracted markdown text |
doc_id_hash |
string |
Hash ID generated from document content |
pages_processed |
int32 |
Estimated page count (3000 chars = 1 page) |
entities |
string |
JSON string of extracted entities (when entity extraction is enabled) |
extracted_data |
string |
Structured data from template extraction (when applicable) |
When expand_extracted_data: true, entity fields are expanded into individual columns.
{
"operator_params": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
}
},
"entity_extraction": {
"provider": "none"
}
}
}
{
"operator_params": {
"max_workers": 2,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.2",
"api_base": "http://localhost:11434/v1",
"api_key": "<ollama_key>",
"temperature": 0.0,
"max_tokens": 4096
},
"custom_schema": {
"invoice_number": "string",
"vendor_name": "string",
"total_amount": "number"
}
}
}
}
{
"operator_params": {
"max_workers": 1,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"vlm_pipeline": {
"preset": "granite_docling",
"engine": "transformers",
"engine_options": {}
}
}
},
"entity_extraction": {
"provider": "none"
}
}
}
{
"operator_params": {
"max_workers": 4,
"text_extraction": {
"provider": "docling_serve",
"doc_column": "content",
"provider_config": {
"base_url": "http://localhost:5001",
"ocr": {
"enabled": true,
"engine": "tesseract",
"mode": "layout_regions",
"engine_options": {
"lang": ["eng", "spa"]
}
},
"pdf_backend": "dlparse_v4",
"table_mode": "accurate"
}
},
"entity_extraction": {
"provider": "none"
}
}
}
{
"operator_params": {
"max_workers": 2,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "docling",
"custom_schema": {
"invoice_number": "string",
"invoice_date": "string",
"vendor_name": "string",
"vendor_address": "string",
"total_amount": "number",
"currency": "string",
"line_items": [
{
"description": "string",
"quantity": "number",
"unit_price": "number",
"total": "number"
}
]
}
}
}
}
{
"operator_params": {
"max_workers": 1,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"vlm_pipeline": {
"preset": "granite_docling",
"engine": "transformers",
"engine_options": {}
}
}
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.2",
"api_base": "http://localhost:11434/v1",
"api_key": "<ollama_key>",
"temperature": 0.0
}
}
}
}
{
"operator_params": {
"max_workers": 2,
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"asr_pipeline": {
"model_id": "whisper_turbo"
}
}
},
"entity_extraction": {
"provider": "none"
}
}
}
Requirement: ASR (Automatic Speech Recognition) dependencies must be installed to process audio and video files
Installation:
# Install ASR dependencies
uv pip install -e '.[asr]'
Supported Audio/Video Formats (when ASR is installed):
Note: If ASR dependencies are not installed, the operator will only support standard document formats (PDF, DOCX, PPTX, etc.) and will log a warning if use_asr_pipeline=true is configured.
Used By: Text extraction with ASR when use_asr_pipeline=true
Requirement: ffmpeg must be installed and available on your PATH for processing certain audio and video formats
Required For:
Not Required For:
Installation:
macOS (using Homebrew):
brew install ffmpeg
Linux (Ubuntu/Debian):
sudo apt update
sudo apt install ffmpeg
Linux (RHEL/CentOS/Fedora):
sudo dnf install ffmpeg
Verify Installation:
ffmpeg -version
Used By: Text extraction with ASR (Automatic Speech Recognition) when processing audio/video files
Requirement: Ollama server must be running on http://localhost:11434
Setup:
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Pull required model
ollama pull llama3.2
# Verify Ollama is running
curl http://localhost:11434/api/tags
Used By: Entity extraction when entity_extraction.provider="litellm" with entity_extraction.provider_config.model_id="openai/llama3.2" and entity_extraction.provider_config.api_base="http://localhost:11434/v1"
Requirement: Docling Serve must be running (default: http://localhost:5001)
Setup:
# Using Docker
docker run -p 5001:5001 ds4sd/docling-serve:latest
# Or using docker-compose
docker-compose -f docker-compose.docling-serve.yml up -d
# Verify service is running
curl http://localhost:5001/health
Used By: Text extraction when text_extraction.provider="docling_serve"
| Feature | Docling Library (Basic) | Docling Library (VLM) | Docling Serve |
|---|---|---|---|
| Processing Location | Local | Local/API | Remote API |
| OCR Support | No | Limited | Yes (EasyOCR, Tesseract) |
| Multi-language OCR | No | No | Yes |
| Scalability | Low | Medium | High |
| Setup Complexity | Low | Medium | Medium |
| Processing Speed | Fast | Slow | Medium |
| Accuracy | Good | Excellent | Excellent |
| External Dependencies | None | Model files | Docker container |
| Feature | LiteLLM (Ollama) | LiteLLM (Cloud) | Docling | WatsonX |
|---|---|---|---|---|
| Processing Location | Local | Remote API | Local | Remote API |
| Schema Support | Yes | Yes | Yes | Yes |
| Schema-Free Mode | Yes | Yes | No | Yes |
| Setup Complexity | Medium | Low | Low | Medium |
| Processing Speed | Medium | Fast | Fast | Medium |
| Accuracy | High | High | Good | High |
| External Dependencies | Ollama server | API keys | None | WatsonX credentials |
Use Docling Library Provider (Basic) When:
Use Docling Library Provider (VLM Pipeline) When:
Use Docling Serve Provider When:
Use None Provider When:
Use LiteLLM Provider When:
Use Docling Provider When:
Use WatsonX Provider When:
entity_max_doc_chars to control LLM input sizeuse_processes: true for CPU-intensive tasksuse_processes: false (default) for I/O-bound tasksThe operator provides the following metadata after execution:
page_type_stats (dict): Aggregate estimated pages grouped by source document format (e.g., {"pdf": 120, "docx": 45})total_pages_converted (int): Total estimated pages across all successfully processed documentsThese metrics are available through the operator’s metadata and can be used for tracking document processing volume and performance analysis.
Complete sample flows are available in sample_flows/:
sample_flows/quickstart/basic_ingest_extract.json - Basic text extraction without entity extractionsample_flows/quickstart/complete_pipeline_ollama.json - Complete extraction with VLM and entity extraction (Ollama)sample_flows/quickstart/complete_pipeline_watsonx.json - Complete extraction with WatsonXsample_flows/operators/entity_extraction_litellm.json - LiteLLM entity extractionsample_flows/operators/entity_extraction_watsonx.json - WatsonX entity extractionsample_flows/use_cases/audio_video_extraction.json - Audio and video extractionIssue: “Failed to initialize text extraction adapter”
text_extraction.provider value is valid: "docling_library" or "docling_serve"vlm_pipeline is configured), ensure required model files are availableIssue: “Failed to initialize entity extraction adapter”
entity_extraction.provider value is valid: "litellm", "docling", "watsonx", or "none"openai/ model prefixWATSONX_API_KEY and WATSONX_CONTAINER_ID are setIssue: “Ollama connection refused”
curl http://localhost:11434/api/tagsollama listIssue: “Docling Serve timeout”
timeout valuecurl http://localhost:5001/healthIssue: “Entity extraction returns empty results”
entity_max_doc_chars is not too restrictiveIssue: “Audio/video processing fails with codec errors”
ffmpeg -versionwhich ffmpeg (macOS/Linux) or where ffmpeg (Windows)brew install ffmpegsudo apt install ffmpeg or sudo dnf install ffmpegIssue: “ffmpeg not found” error during audio/video extraction
ffmpeg -versionexport PATH="/opt/homebrew/bin:$PATH"The ExtractOperator follows hexagonal architecture (ports and adapters pattern) with clear separation of concerns:
ExtractOperator (Orchestrator)
↓
┌─────────────────────────────────────────────────────────────┐
│ Domain Layer │
│ - EntityExtractionService (business logic) │
│ - Domain Models (TextExtractionMode, EntityExtractionMode) │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ Port Layer (Interfaces) │
│ - TextExtractionPort │
│ - EntityExtractionPort │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ Adapter Layer (Implementations) │
│ Text Extraction: │
│ - DoclingAdapter (docling_library provider, optional VLM/ASR) │
│ - DoclingServeAdapter (docling_serve provider) │
│ Entity Extraction: │
│ - LLMEntityAdapter (litellm and watsonx providers - unified) │
│ - DoclingEntityAdapter (docling provider) │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ Factory Layer │
│ - TextExtractionAdapterFactory │
│ - EntityExtractionAdapterFactory │
└─────────────────────────────────────────────────────────────┘
Architecture Components:
EntityExtractionService handles business logic (prompt building, schema validation, response parsing)Key Benefits:
litellm and watsonx use the same LLMEntityAdapterTextExtractionAdapterFactoryEntityExtractionAdapterFactory (if enabled)EntityExtractionService with the entity adapterEntityExtractionService builds prompts with schema