The Flow Authoring Format is a simplified, user-friendly way to define Docling Pipelines pipelines without managing UUIDs, edges, and low-level DAG details.
The authoring format is a simplified JSON structure for defining Docling Pipelines flows that:
DocpipeFlowManager accepts authoring formatPOST /api/v1/flows (default format)Authoring Format (Recommended):
{
"flow_name": "My Pipeline",
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {"provider": "filesystem", "connection_params": {"paths": ["./data"]}}
},
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {}
}
]
}
Elyra Format (UI/Legacy):
{
"doc_type": "pipeline",
"pipelines": [{
"nodes": [
{
"id": "30953cfb-a3a2-4688-9aea-ff9fff10f7bd",
"app_data": {
"component_parameters": {
"operator": "ingest_source",
"config": {"provider": "filesystem", "connection_params": {"paths": ["./data"]}}
}
},
"outputs": [{
"id": "7cfd7577-b061-4fc9-92d5-120ae0fbde89"
}]
}
]
}]
}
Note: Both formats are supported. Elyra format is used by the UI, while authoring format is recommended for CLI, Python API, and programmatic usage.
{
"flow_name": "string", // Flow name (required)
"flow": [ // Array of operators (required)
{
"name": "string", // Unique operator name (required)
"type": "string", // Operator type (required)
"depends_on": ["string"], // Dependencies (optional, default: [])
"config": {} // Operator config (optional, default: {})
}
]
}
{
"description": "string", // Flow description
"global_config": {}, // Global configuration
"tags": ["string"], // Flow tags (HTTP API only)
"flow_source": "cli|api|programmatic|ui" // Auto-set based on usage
}
{
"flow_name": "Document Processing",
"description": "Ingest and extract documents",
"flow": [
{
"name": "ingest_docs",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./documents"]},
"include_filter": "pdf,txt"
}
},
{
"name": "extract_text",
"type": "extract_operator",
"depends_on": ["ingest_docs"],
"config": {
"text_extraction": {"provider": "docling_library"}
}
}
],
"global_config": {
"doc_column": "content",
"disable_validation": false
}
}
Each operator in the flow array has:
{
"name": "unique_operator_name", // Required: Unique identifier
"type": "operator_type", // Required: Operator class name
"depends_on": ["parent1", "parent2"], // Optional: Dependencies
"config": { // Optional: Operator-specific config
"param1": "value1",
"param2": "value2"
}
}
.) - reserved for branch referencesingest_pdfs, extract_entities)Use the operator’s class name or short name:
{
"type": "ingest_source" // ✅ Short name
}
{
"type": "IngestSourceOperator" // ✅ Class name (also works)
}
List available operators:
docling-pipelines --list-operators
Reference operators by name:
{
"name": "chunk_docs",
"type": "chunker",
"depends_on": ["extract_text"], // Depends on one operator
"config": {}
}
{
"name": "merge_results",
"type": "merge",
"depends_on": ["path_a", "path_b", "path_c"], // Merge multiple inputs
"config": {}
}
For branching operators, reference specific branches:
{
"name": "classifier",
"type": "branching",
"config": {
"branches": {
"invoices": {"condition": "type == 'invoice'"},
"receipts": {"condition": "type == 'receipt'"}
}
}
},
{
"name": "process_invoices",
"type": "extract_operator",
"depends_on": ["classifier.invoices"], // Reference specific branch
"config": {}
}
At least one operator must have no dependencies (entry point):
{
"name": "start_here",
"type": "ingest_source",
"depends_on": [], // Entry point - no dependencies
"config": {}
}
{
"global_config": {
"doc_column": "content", // Document content column
"disable_validation": false, // Enable/disable validation
"force_ingest": true, // Force re-ingestion
"enable_micro_batching": true, // Enable batching
"micro_batch_size": 10, // Batch size
"data_storage_type": "local" // Storage backend
}
}
Operators can override global settings in their config:
{
"global_config": {
"doc_column": "content"
},
"flow": [
{
"name": "special_extract",
"type": "extract_operator",
"config": {
"doc_column": "raw_text" // Overrides global setting
}
}
]
}
# Execute authoring format flow
docling-pipelines --flow-file my_flow.json
# Validate without executing
docling-pipelines validate-flow my_flow.json
from docpipe.lib.docpipe_flow_manager import DocpipeFlowManager
# From file
manager = DocpipeFlowManager(flow_file="my_flow.json")
result = manager.execute()
# From dictionary
flow_def = {
"flow_name": "My Pipeline",
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {"provider": "filesystem", "connection_params": {"paths": ["./data"]}}
}
]
}
manager = DocpipeFlowManager(flow_def=flow_def)
result = manager.execute()
# Create flow with authoring format (default)
curl -X POST http://localhost:8000/api/v1/flows \
-H "Content-Type: application/json" \
-d @my_flow.json
# Explicitly specify authoring format
curl -X POST "http://localhost:8000/api/v1/flows?is_elyra=false" \
-H "Content-Type: application/json" \
-d @my_flow.json
# Use Elyra format (for UI compatibility)
curl -X POST "http://localhost:8000/api/v1/flows?is_elyra=true" \
-H "Content-Type: application/json" \
-d @elyra_flow.json
{
"flow_name": "PDF Extraction Pipeline",
"description": "Extract text and entities from PDF documents",
"flow": [
{
"name": "ingest_pdfs",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./pdfs"]},
"include_filter": "pdf"
}
},
{
"name": "extract_content",
"type": "extract_operator",
"depends_on": ["ingest_pdfs"],
"config": {
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.2",
"api_base": "http://localhost:11434/v1"
}
}
}
}
],
"global_config": {
"doc_column": "content",
"disable_validation": false
}
}
{
"flow_name": "RAG Pipeline",
"description": "Complete pipeline for RAG: ingest, extract, chunk, embed, store",
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./documents"]},
"include_filter": "pdf,txt,md"
}
},
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {"provider": "docling_library"}
}
},
{
"name": "chunk",
"type": "chunker",
"depends_on": ["extract"],
"config": {
"chunk_type": "semantic",
"chunk_size": 512,
"chunk_overlap": 50
}
},
{
"name": "embed",
"type": "embeddings",
"depends_on": ["chunk"],
"config": {
"provider": "litellm",
"model_id": "openai/nomic-embed-text",
"provider_config": {
"api_base": "http://localhost:11434"
}
}
},
{
"name": "store",
"type": "vectordb",
"depends_on": ["embed"],
"config": {
"provider": "opensearch",
"vector_dimension": 768,
"provider_config": {
"index_name": "documents"
}
}
}
],
"global_config": {
"doc_column": "content",
"enable_micro_batching": true,
"micro_batch_size": 10
}
}
{
"flow_name": "Document Classification Pipeline",
"description": "Classify and route documents by type",
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./mixed_docs"]}
}
},
{
"name": "classify",
"type": "branching",
"depends_on": ["ingest"],
"config": {
"branches": {
"invoices": {"condition": "doc_type == 'invoice'"},
"contracts": {"condition": "doc_type == 'contract'"},
"other": {"condition": "True"}
}
}
},
{
"name": "process_invoices",
"type": "extract_operator",
"depends_on": ["classify.invoices"],
"config": {
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.2",
"api_base": "http://localhost:11434/v1"
},
"custom_schema": "invoice_template"
}
}
},
{
"name": "process_contracts",
"type": "extract_operator",
"depends_on": ["classify.contracts"],
"config": {
"text_extraction": {
"provider": "docling_library"
},
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/llama3.2",
"api_base": "http://localhost:11434/v1"
},
"custom_schema": "contract_template"
}
}
},
{
"name": "merge_results",
"type": "merge",
"depends_on": ["process_invoices", "process_contracts"],
"config": {}
}
]
}
{
"flow_name": "Scalability Test - HuggingFace Local",
"description": "High-concurrency pipeline using native HuggingFace local inference",
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["/data/documents"]},
"include_filter": "txt,pdf",
"max_files": 10000
}
},
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": {
"provider": "docling_library"
}
}
},
{
"name": "chunk",
"type": "chunker",
"depends_on": ["extract"],
"config": {
"chunk_type": "simple",
"chunk_size": 4000,
"chunk_overlap": 300
}
},
{
"name": "embed",
"type": "embeddings",
"depends_on": ["chunk"],
"config": {
"provider": "huggingface",
"model_id": "sentence-transformers/all-MiniLM-L6-v2",
"provider_config": {
"use_local": true,
"device": "cpu",
"batch_size": 16
}
}
},
{
"name": "store",
"type": "vectordb",
"depends_on": ["embed"],
"config": {
"provider": "opensearch",
"vector_dimension": 384,
"provider_config": {
"index_name": "documents"
}
}
}
],
"global_config": {
"enable_micro_batching": true,
"micro_batch_size": 10,
"max_concurrent_batches": 600
}
}
ingest_customer_docs vs op1validate-flow before execution