This document provides a comprehensive reference for all global configuration parameters that can be set in docpipe flow definitions. Global configuration parameters control flow-level behavior and are specified in the global_config section of your flow JSON.
Global configuration parameters are set at the flow level and apply to all operators in the pipeline unless overridden at the operator level. These parameters control:
Global configuration is specified in the global_config section of your flow definition:
{
"flow_name": "My Pipeline",
"description": "Example pipeline with global configuration",
"global_config": {
"force_ingest": false,
"micro_batch_size": 50,
"prefect": {
"batch_execution": {
"strategy": "thread-pool"
}
}
},
"flow": [
{
"type": "ingest_source",
"name": "ingest",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./documents"]}
}
},
{
"type": "extract_operator",
"name": "extract",
"config": {
"text_extraction": {"provider": "docling_library"}
},
"depends_on": ["ingest"]
}
]
}
Parameters that control how the flow executes and processes data.
disable_validationType: boolean
Default: false
Description: Disables flow validation before execution. Not recommended for production use.
Valid Values:
true: Skip flow validation (faster startup, risky)false: Validate flow before execution (recommended)Example:
{
"global_config": {
"disable_validation": true
}
}
Warning: Disabling validation skips the main flow validation pass and can lead to runtime errors that validation would normally catch.
What Gets Skipped When disable_validation: true:
skip_custom_op_validationType: boolean
Default: false
Description: Skips validation for custom operators while still validating built-in operators.
Valid Values:
true: Skip custom operator validationfalse: Validate all operators including custom onesExample:
{
"global_config": {
"skip_custom_op_validation": true
}
}
output_folderType: string
Description: Directory for storing final output files. Can be either a relative path (relative to workspace directory) or an absolute path. If not specified, the system generates a unique path based on job execution IDs.
Valid Values: Valid directory path (relative or absolute)
Example:
{
"global_config": {
"data_local_config": {
"output_folder": "./data/output"
}
}
}
data_storage_typeType: string
Description: Controls intermediate execution data storage behavior during flow execution.
Valid Values:
"memory": Use in-memory intermediate data handling"local": Use local filesystem-backed intermediate data handlingExample:
{
"global_config": {
"data_storage_type": "local"
}
}
Note: Memory storage is fastest but limited by available RAM. Use "local" for large datasets.
memmap_thresholdType: integer
Default: 100
Description: Threshold in MB after which persistent storage is used for chunks and embeddings.
Valid Values: Must be greater than 1.
Example:
{
"global_config": {
"memmap_threshold": 100
}
}
Configuration for tracking and processing only changed documents.
force_ingestType: boolean
Default: false
Description: Forces re-ingestion of all documents, even if they were previously processed. Useful for reprocessing data after operator configuration changes.
Valid Values:
true: Re-ingest all documents regardless of previous processingfalse: Skip documents that were already processed (incremental mode)Example:
{
"global_config": {
"force_ingest": true
}
}
Use Cases:
retain_deleted_docsType: boolean
Default: true
Description: Controls whether documents deleted from the source should be retained in the output or removed.
Valid Values:
true: Keep documents in output even if deleted from sourcefalse: Remove documents from output when deleted from sourceExample:
{
"global_config": {
"retain_deleted_docs": false
}
}
Use Cases:
true)false)Important Limitations:
./sample_documents), the system will report an error: Path does not existforce_ingest: true to reset incremental metadataDocling Pipelines uses a centralized configuration system for incremental metadata storage. Instead of configuring incremental metadata in each flow JSON file, you configure it once in a repository-level docling-pipelines-config.yaml file.
The incremental metadata configuration is defined in docling-pipelines-config.yaml:
Option 1: Inherit from global_storage
global_storage:
type: "filesystem" # Options: filesystem, postgresql
config:
base_dir: "./data"
lock_timeout: 30.0
# PostgreSQL configuration (only needed if type is "postgres")
postgres:
host: "localhost"
port: 5432
database: "docpipe"
user: "docpipe_user"
password: "${POSTGRES_PASSWORD}" # pragma: allowlist secret
schema: "incremental_metadata"
incremental_metadata: {} # inherit from global_storage
Option 2: Override with service-specific configuration
global_storage:
type: "filesystem"
config:
base_dir: "./data"
lock_timeout: 30.0
# incremental_metadata overrides global_storage
incremental_metadata:
storage:
type: "filesystem"
config:
base_dir: "./incremental_data" # Different path
lock_timeout: 60.0
Docling Pipelines supports two storage types for incremental metadata:
config.base_dir and config.lock_timeoutYou can override configuration using environment variables:
DOCPIPE_INCREMENTAL_BASE_DIR: Override the base directory for incremental metadataDOCPIPE_INCREMENTAL_STORAGE_BACKEND: Override the storage backend (filesystem, postgresql)DOCPIPE_CONFIG_PATH: Specify a custom path to docling-pipelines-config.yamlExample:
export DOCPIPE_INCREMENTAL_BASE_DIR="/data/incremental"
export DOCPIPE_INCREMENTAL_STORAGE_BACKEND="filesystem"
docling-pipelines --flow-file my_flow.json
Configuration is resolved in the following order (highest to lowest priority):
DOCPIPE_INCREMENTAL_BASE_DIR, DOCPIPE_INCREMENTAL_STORAGE_BACKEND)incremental_metadata.storage in docling-pipelines-config.yaml)global_storage in docling-pipelines-config.yaml)./data)In your flow JSON files, you no longer need to specify incremental metadata configuration. The system automatically uses the centralized configuration:
{
"flow_name": "My Pipeline",
"description": "Pipeline with centralized incremental metadata",
"global_config": {
"force_ingest": false,
"retain_deleted_docs": true
},
"flow": [
{
"type": "ingest_source",
"name": "ingest",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./documents"]}
}
}
]
}
If you have existing flows with flow-level incremental metadata configuration, follow these steps:
incremental_metadata:
storage:
type: "filesystem"
config:
base_dir: "./incremental_data"
lock_timeout: 60.0
incremental_metadata section from global_configdocling-pipelines --flow-file your_flow.json --validate
Related Documentation :
Configuration for prefect-based workflow orchestration and batch execution strategies.
prefectType: object
Description: Prefect orchestration settings including batch execution strategy and work pool configuration.
Structure :
{
"prefect": {
"batch_execution": {
"strategy": "thread-pool",
"work_pool_name": "my-pool",
"deployment_name": "my-deployment",
"batch_storage": {
"type": "local",
"path": "./batch_data"
}
}
}
}
The batch_execution section controls how batches are executed.
strategyType: string
Default: "thread-pool"
Description: Execution strategy for batch processing.
Valid Values:
"thread-pool": Local execution using ThreadPoolTaskRunner (default, simplest)"work-pool-process": Distributed execution using Prefect process work pools"work-pool-docker": Distributed execution using Docker containersExample - Thread Pool (Local):
{
"prefect": {
"batch_execution": {
"strategy": "thread-pool"
}
}
}
Example - Docker Work Pool:
{
"prefect": {
"batch_execution": {
"strategy": "work-pool-docker",
"work_pool_name": "docpipe-docker-pool",
"deployment_name": "batch-processor",
"deployment_path": "/opt/docpipe",
"image": "docpipe:latest",
"batch_storage": {
"type": "local",
"path": "/data/batches"
}
}
}
}
Required when using work pool strategies (work-pool-*).
work_pool_nameType: string
Required: Yes (for work pool strategies)
Description: Name of the Prefect work pool to use for batch execution.
Example:
{
"work_pool_name": "docpipe-production-pool"
}
deployment_nameType: string
Required: Yes (for work pool strategies)
Description: Name for the Prefect deployment that will execute batches.
Example:
{
"deployment_name": "document-processing-v1"
}
deployment_pathType: string
Default: Current working directory
Description: Runtime path where flow code is available in the worker environment.
Example:
{
"deployment_path": "/opt/docpipe"
}
envType: object (dictionary of string key-value pairs)
Default: {}
Description: Environment variables injected into the worker job process or container. Used to pass configuration, credentials, or runtime settings to workers.
Example:
{
"env": {
"OLLAMA_HOST": "http://ollama-service:11434",
"LOG_LEVEL": "INFO",
"PREFECT_MODE": "cloud",
"PREFECT_URL": "https://api.prefect.cloud",
"DOCPIPE_STORAGE_BACKEND": "postgresql",
"PYTHONPATH": "/app/src",
"LOCAL_FLOWS_DIR": "/app/flows",
"DOCPIPE_DATA_PATH": "/data/docpipe"
}
}
Note: System-required environment variables (PREFECT_API_URL, PYTHONPATH, etc.) are automatically set by the adapter. User-provided env values supplement or override these defaults.
imageType: string
Required: Yes (for containerized work pools)
Description: Container image for Docker work pools.
Example:
{
"image": "myregistry/docpipe:v1.2.3"
}
image_pull_policyType: string
Default: "Never"
Description: Policy for pulling container images.
Valid Values:
"Always": Always pull the image"IfNotPresent": Pull only if not present locally"Never": Never pull, use local image onlyExample:
{
"image_pull_policy": "Never"
}
networksType: array of strings
Default: []
Description: Docker networks for spawned containers.
Example:
{
"networks": ["docpipe-network", "monitoring-network"]
}
The batch_storage section controls where batch data is stored during distributed execution.
typeType: string
Default: "inline"
Description: Storage type for batch data.
Valid Values:
"inline": Serialize batch data directly in deployment parameters (small batches only)"local": Store batch data on local filesystempathType: string
Required: Yes (when type is "local")
Description: Filesystem path for storing batch data when using local storage type.
Example - Local Storage:
{
"batch_storage": {
"type": "local",
"path": "/data/batches"
}
}
Related Documentation: Prefect Distributed Execution Guide
Parameters for controlling micro-batching behavior. Micro-batching must be explicitly enabled and splits large datasets into smaller batches for parallel processing.
micro_batch_sizeType: integer
Default: 100
Description: Note that the batch_size is not strictly enforced. The files are adjusted in the batches to make the batch size uniform across the batches, but limiting the number of documents in a batch to the given batch_size.
Valid Values: Positive integer
Example:
{
"global_config": {
"micro_batch_size": 50
}
}
Tuning Guidelines:
max_concurrent_batchesType: integer
Default: 10
Description: Maximum number of batches that can execute concurrently. Controls parallelism and resource usage.
Valid Values: Positive integer
Example:
{
"global_config": {
"max_concurrent_batches": 5
}
}
Tuning Guidelines:
Note: Higher concurrency requires more memory and CPU resources.
You can override configuration for specific operators using their name or ID.
<operator_name> or <operator_id>Type: object
Description: Override configuration for a specific operator by its name or ID.
Example:
{
"global_config": {
"doc_column": "content",
"extract_operator": {
"doc_column": "document",
"max_workers": 8
}
}
}
In this example, all operators use doc_column: "content" except the extract_operator which uses doc_column: "document".
Here’s a comprehensive example showing the separation between flow JSON and docling-pipelines-config.yaml:
{
"flow_name": "Production Document Processing Pipeline",
"description": "Process documents with micro-batching and Docker execution",
"global_config": {
"force_ingest": false,
"retain_deleted_docs": true,
"micro_batch_size": 100,
"max_concurrent_batches": 10,
"data_local_config": {
"output_folder": "./output"
},
"data_storage_type": "local",
"prefect": {
"batch_execution": {
"strategy": "work-pool-docker",
"work_pool_name": "docpipe-docker-pool",
"deployment_name": "doc-processor-v1",
"deployment_path": "/opt/docpipe",
"image": "myregistry/docpipe:1.0.0",
"env": {
"OLLAMA_HOST": "http://ollama-service:11434",
"LOG_LEVEL": "INFO"
},
"batch_storage": {
"type": "local",
"path": "/data/batches"
}
}
}
},
"flow": [
{
"type": "ingest_source",
"name": "ingest_source_filesystem",
"config": {
"provider": "filesystem",
"connection_params": {"paths": ["./sample_documents"]},
"include_filter": "pdf,txt,docx"
}
},
{
"type": "extract_operator",
"name": "extract_with_docling",
"config": {
"doc_column": "content"
},
"depends_on": ["ingest_source_filesystem"]
}
]
}
docling-pipelines-config.yaml)# Repository-level configuration for incremental metadata
incremental_metadata:
storage:
type: "filesystem"
config:
base_dir: "./incremental_data"
lock_timeout: 60.0
# PostgreSQL configuration (only needed if type is "postgres")
# postgres:
# host: "localhost"
# port: 5432
# database: "docpipe"
# user: "docpipe_user"
# password: "${POSTGRES_PASSWORD}" # pragma: allowlist secret
# schema: "incremental_metadata"
# The flow automatically uses the centralized incremental metadata configuration
docling-pipelines --flow-file production_pipeline.json
Key Points:
docling-pipelines-config.yaml