This guide explains how to create and configure Docling Pipelines pipeline flows using JSON configuration files.
Docling Pipelines pipelines are defined using JSON configuration files. Here’s the basic structure:
{
"flow_name": "complete-document-pipeline",
"description": "Complete document processing pipeline",
"global_config": {
"doc_column": "content",
"disable_validation": false,
"force_ingest": true,
"storage": "in-memory",
"execute_type": "local"
},
"flow": [
// Operator definitions go here
]
}
| Field | Type | Description | Example |
|---|---|---|---|
flow_name |
string | Human-readable flow name | "complete-document-pipeline" |
description |
string | Flow purpose description | "Complete document processing pipeline" |
global_config |
object | Global configuration | See below |
flow |
array | Operator definitions | See operator sections |
The global_config object supports the following options:
| Field | Type | Description | Default | Example |
|---|---|---|---|---|
doc_column |
string | Column name for document content | "content" |
"content" |
force_ingest |
boolean | Force re-ingestion of documents | false |
true |
disable_validation |
boolean | Disable flow validation | false |
true |
Example with global configuration
{
"global_config": {
"doc_column": "content",
"force_ingest": true,
"disable_validation": true
}
}
The flow_name field in your flow definition serves different purposes depending on how you execute the flow:
docling-pipelines)When using the docling-pipelines CLI, the flow_name is automatically used to generate a unique job_id for tracking flow executions and incremental processing.
Automatic job_id Generation:
a1b2c3d4-e5f6-5789-a012-b3c4d5e6f7a8){sanitized}_{hash}Important Considerations:
⚠️ flow_name Uniqueness: Use a unique flow_name for each logically different pipeline. Reusing the same flow_name across different pipelines will generate the same job_id, causing incremental metadata conflicts where documents processed in one pipeline may be incorrectly marked as processed in another.
⚠️ Incremental Processing Impact: Since incremental ingestion metadata is associated with the job_id (derived from flow_name), changing a flow’s flow_name will generate a new job_id, causing previously processed files to be reprocessed.
DocpipeFlowManager)When using the Python API, you must provide a unique job_id parameter for each flow execution:
from docpipe.lib.docpipe_flow_manager import DocpipeFlowManager
# You must provide a unique job_id
manager = DocpipeFlowManager(
flow_file="my_flow.json",
job_id="my-unique-job-id-12345" # Required for proper tracking
)
result = manager.execute()
If no job_id is provided, a random UUID will be generated, which means incremental processing will not work correctly across executions.
When creating flows via the REST API, a flow_id is automatically generated and used as the job_id for execution tracking.
Reads files from a local directory:
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./tests/fixtures/invoices"]},
"include_filter": ".pdf"
}
}
Note: Extension names should include the dot prefix and be comma-separated (e.g., ".pdf,.txt,.docx" not "*.pdf,*.txt" or "pdf,txt").
The extract_operator handles both text extraction and entity extraction.
Supported text extraction providers:
docling_librarydocling_serveSupported entity extraction providers:
doclinglitellmwatsonxnone{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "none"
}
}
}
For structured data extraction with predefined schemas, use entity_extraction.provider: "docling"
Important: When using any entity extraction provider (not none), you must provide either:
custom_schema in the operator configuration (as shown below), ORdocument_type column from an upstream classification operator (e.g., DocumentClassifierOperator)If neither is provided, the operator will throw a ConfigurationError.
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "docling",
"expand_extracted_data": true,
"custom_schema": {
"invoice_number": "string",
"invoice_date": "string",
"vendor_name": "string",
"total": "float"
}
}
}
}
For LLM-powered entity extraction using Ollama models, use entity_extraction.provider: "litellm" with the openai/ prefix:
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.1:latest",
"api_base": "http://localhost:11434/v1"
},
"expand_extracted_data": true,
"custom_schema": {
"invoice_number": "string",
"invoice_date": "string",
"vendor_name": "string",
"total": "float"
}
}
}
}
This approach uses LiteLLM to access Ollama models for flexible, LLM-powered entity extraction. This is useful when you need to extract specific fields from structured documents like invoices, forms, or receipts.
For hardware-accelerated PDF and image extraction using a local GPU. The adapter builds one DocumentConverter at initialization and reuses it for every document — GPU model weights are loaded once per flow execution.
Requires
max_workers: 1anduse_processes: false. Cannot be combined withvlm_pipeline.
When device is omitted, the best available GPU is auto-detected via torch at runtime (CUDA → MPS → XPU). Specify device explicitly to pin a particular GPU.
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content",
"provider_config": {
"standard_pipeline": {
"accelerator": {
"num_threads": 6
}
}
}
},
"entity_extraction": {
"provider": "none"
},
"max_workers": 1,
"use_processes": false
}
}
Supported devices: mps (Apple Silicon), cuda (NVIDIA), cuda:<index> (e.g. cuda:0), xpu (Intel). Auto-detected when device is not specified.
Splits documents into chunks:
{
"name": "chunk",
"type": "chunker",
"depends_on": ["extract"],
"config": {
"chunk_type": "hybrid",
"doc_column": "content",
"chunk_size": 512,
"chunk_overlap": 128,
"retain_original_content": false
}
}
Generates vector embeddings:
{
"name": "embeddings",
"type": "embeddings",
"depends_on": ["chunk"],
"config": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/granite4:latest",
"api_base": "http://localhost:11434"
},
"embeddings_column": "embeddings",
"overlap_ratio": 0.2,
"doc_column": "content"
}
}
Note: Use ollama list to see available models on your system. Common embedding models include granite4:latest, nomic-embed-text, and mxbai-embed-large.
Stores documents and embeddings in OpenSearch for vector similarity search.
{
"name": "vectordb",
"type": "vectordb",
"depends_on": ["embeddings"],
"config": {
"provider": "opensearch",
"doc_id_column": "doc_id_hash",
"embeddings_column": "embeddings",
"create_index": true,
"vector_dimension": 384,
"provider_config": {
"index_name": "documents",
"host": "localhost",
"port": 9200,
"username": "admin",
"password": "<your-opensearch-password>",
"use_ssl": false,
"verify_certs": false,
"engine": "faiss",
"algorithm": "hnsw",
"space_type": "l2",
"batch_size": 100
},
"feature_mappings": {
"content": "content",
"doc_name": "doc_name",
"file_path": "file_path",
"doc_id_hash": "doc_id_hash",
"chunk_id": "chunk_id",
"chunk_index": "chunk_index"
},
"available_features": {
"embeddings": {
"type": "vector",
"available_for_vector_db": true
},
"content": {
"type": "string",
"available_for_vector_db": true
},
"doc_name": {
"type": "string",
"available_for_vector_db": true
},
"file_path": {
"type": "string",
"available_for_vector_db": true
},
"doc_id_hash": {
"type": "string",
"available_for_vector_db": true
},
"chunk_id": {
"type": "string",
"available_for_vector_db": true
},
"chunk_index": {
"type": "integer",
"available_for_vector_db": true
}
}
}
}
embeddings with "type": "vector" and "available_for_vector_db": true{"pyarrow_column": "vector_db_field"}nomic-embed-text: 768, llama3.2: 4096, granite-embedding: 384src/docpipe/core/operators/vectordb/)
schemas/default_schema.v1.json, schemas/template_with_content_analyzer.v1.jsonavailable_features⚠️ Important: The embeddings column is mandatory. The operator validates embeddings exist in the input table and will fail if missing. You must explicitly configure embeddings in
available_featuresfor them to be stored in OpenSearch.
Schema templates provide reusable index configurations with consistent settings across pipelines. Instead of defining available_features and feature_mappings manually, you can use a pre-configured template.
Benefits of Schema Templates:
Built-in Templates:
Example with Schema Template:
{
"name": "vectordb",
"type": "vectordb",
"depends_on": ["embeddings"],
"config": {
"provider": "opensearch",
"doc_id_column": "doc_id_hash",
"embeddings_column": "embeddings",
"create_index": true,
"vector_dimension": 384,
"provider_config": {
"index_name": "document_chunks",
"schema_template_path": "schemas/template_with_content_analyzer.v1.json",
"host": "localhost",
"port": 9200,
"username": "admin",
"password": "<your-opensearch-password>",
"use_ssl": false,
"verify_certs": false,
"engine": "faiss",
"algorithm": "hnsw",
"space_type": "l2"
},
"feature_mappings": {
"content": "content",
"doc_id_hash": "id",
"embeddings": "embeddings"
},
"available_features": {
"embeddings": {
"type": "vector",
"available_for_vector_db": true
},
"content": {
"type": "content_text",
"available_for_vector_db": true
},
"doc_id_hash": {
"type": "string",
"available_for_vector_db": true
}
}
}
}
Note: When using schema templates, the template defines the index settings and field type mappings. You still need to specify available_features and feature_mappings to control which columns from your data are stored.
The VectorDBOperator automatically collects common metadata fields into a metadata object:
name, size, created_time, modified_time, source, mimetype, extension, page_countpath → source, pages_processed → page_count)extension and mimetype are derived when possibleOperators are connected using the depends_on field, which specifies which operators must complete before the current operator runs. Simply reference the name of the upstream operator(s).
Example:
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {...}
}
This creates a dependency where the extract operator will only run after the ingest operator completes successfully.
Multiple Dependencies:
{
"name": "merge",
"type": "merge_operator",
"depends_on": ["branch1", "branch2"],
"config": {...}
}
{
"flow_name": "basic-document-pipeline",
"description": "Basic document ingestion and extraction",
"global_config": {
"doc_column": "content",
"force_ingest": false
},
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./documents"]},
"include_filter": ".pdf,.txt"
}
},
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "none"
}
}
}
]
}
{
"flow_name": "complete-rag-pipeline",
"description": "Complete pipeline for RAG system",
"global_config": {
"doc_column": "content",
"force_ingest": true
},
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./documents"]},
"include_filter": ".pdf"
}
},
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library",
"doc_column": "content"
},
"entity_extraction": {
"provider": "none"
}
}
},
{
"name": "chunk",
"type": "chunker",
"depends_on": ["extract"],
"config": {
"chunk_type": "hybrid",
"chunk_size": 512,
"chunk_overlap": 128
}
},
{
"name": "embeddings",
"type": "embeddings",
"depends_on": ["chunk"],
"config": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/nomic-embed-text:latest",
"api_base": "http://localhost:11434"
},
"embeddings_column": "embeddings"
}
},
{
"name": "vectordb",
"type": "vectordb",
"depends_on": ["embeddings"],
"config": {
"provider": "opensearch",
"doc_id_column": "doc_id_hash",
"embeddings_column": "embeddings",
"vector_dimension": 768,
"provider_config": {
"index_name": "documents",
"host": "localhost",
"port": 9200,
"username": "admin",
"password": "<your-opensearch-password>"
},
"feature_mappings": {
"content": "content",
"doc_id_hash": "doc_id_hash",
"embeddings": "embeddings"
},
"available_features": {
"embeddings": {
"type": "vector",
"available_for_vector_db": true
},
"content": {
"type": "string",
"available_for_vector_db": true
},
"doc_id_hash": {
"type": "string",
"available_for_vector_db": true
}
}
}
}
]
}