This guide covers advanced configuration topics for production deployments, distributed execution, and performance optimization.
Docling Pipelines supports pluggable job stats storage for job runs, node execution state, and micro-batch progress tracking.
Job-management components are wired through JobManagementFactory. Backend selection is controlled by environment variables and defaults.
The main user-facing configuration is controlled via environment variables:
job_management:
framework:
type: default
config: {}
store:
type: filesystem
config:
base_dir: ./data/job_stats_store_data
This allows users to configure:
DOCPIPE_FRAMEWORK_TYPE environment variableDOCPIPE_STORAGE_BACKEND environment variableLOCAL_FLOWS_DIR environment variableCommon overrides include:
DOCPIPE_CONFIG_PATHDOCPIPE_STORAGE_BACKENDDOCPIPE_FRAMEWORK_TYPEDOCPIPE_JOB_STATS_BASE_DIRDOCPIPE_POSTGRES_HOSTDOCPIPE_POSTGRES_PORTDOCPIPE_POSTGRES_DBDOCPIPE_POSTGRES_USERDOCPIPE_POSTGRES_PASSWORDEffective precedence for job-management runtime selection is:
JobManagementFactoryFilesystem storage is useful for single-host execution and simple local testing.
Important requirements:
base_dir to an absolute path before injecting it into the worker environmentDuckDB is an embedded analytical database that provides persistent storage without requiring a server. It’s ideal for local development and testing.
Use DuckDB when:
Configuration example:
job_management:
store:
type: duckdb
config:
database_path: ./data/duckdb/job_stats.duckdb
Important notes:
examples/duckdb_job_stats/ for usage examplesPostgreSQL is the recommended backend for multi-process and distributed execution because it provides durable shared storage and stronger concurrency behavior than file-backed JSON storage.
Use PostgreSQL when:
Node stats are aggregated on the read path, not in the storage adapter. When operators add new metadata fields, maintainers must review DEFAULT_STRATEGIES and update it if the field should not use the default LAST aggregation behavior.
See docs/internals/NODE_METADATA_AGGREGATION_STRATEGY.md for the maintainer workflow.
For distributed Prefect execution, work pool runtime configuration is modeled in work_pool_config.py and applied by WorkPoolAdapter.
Important behavior:
env values configured directly in the work pool take highest precedenceThis makes it possible to configure job management via environment variables per environment or per deployment.
For full distributed execution examples and work-pool-specific configuration, see DISTRIBUTED_EXECUTION_GUIDE.md.
Docling Pipelines supports incremental processing to avoid reprocessing unchanged input data. Incremental metadata stores processing state such as file identity and modification information so ingest operators can determine whether an item is new, changed, or already processed.
Use the incremental_metadata section in docling-pipelines-config.yaml:
incremental_metadata:
storage:
type: "filesystem" # Options: filesystem, postgresql
config:
base_dir: "./data"
lock_timeout: 30.0
This configuration is the single source of truth for incremental metadata backend selection and runtime settings.
JSON storage is the default backend. It is file-based, easy to inspect, and suitable for development or smaller single-host deployments.
incremental_metadata:
storage:
type: "json"
config:
base_dir: "./data/incremental_metadata"
lock_timeout: 30.0
Use JSON when:
Parquet storage uses a columnar file format and is better suited to larger datasets or analytics-oriented workflows.
incremental_metadata:
storage:
type: "parquet"
config:
base_dir: "./data/incremental_metadata"
lock_timeout: 30.0
Use Parquet when:
PostgreSQL is the recommended backend for production deployments that require stronger concurrency behavior and durable centralized storage.
incremental_metadata:
storage:
type: "postgresql"
config:
base_dir: "./data"
lock_timeout: 30.0
postgres:
host: "${POSTGRES_HOST:-localhost}"
port: 5432
database: "${POSTGRES_DB:-docpipe}"
user: "${POSTGRES_USER:-docpipe_user}"
password: "${POSTGRES_PASSWORD}"
schema: "incremental_metadata"
Use PostgreSQL when:
Use environment variable substitution in docling-pipelines-config.yaml for credentials and deployment-specific values.
incremental_metadata:
storage:
type: "postgresql"
config:
base_dir: "${DATA_DIR:-./data}"
lock_timeout: "${LOCK_TIMEOUT:-30.0}"
postgres:
host: "${INCR_META_DB_HOST:-localhost}"
port: "${INCR_META_DB_PORT:-5432}"
database: "${INCR_META_DB_NAME:-docpipe}"
user: "${INCR_META_DB_USER:-docpipe_user}"
password: "${INCR_META_DB_PASSWORD}"
schema: "${INCR_META_DB_SCHEMA:-incremental_metadata}"
Guidance:
${VAR_NAME} for required secrets${VAR_NAME:-default} for optional values with safe defaultsSee docling-pipelines-config.yaml.example for complete backend examples and environment variable patterns.
Docling Pipelines uses Prefect as its orchestration engine and supports two Prefect execution modes:
By default, Docling Pipelines runs Prefect in ephemeral mode with a temporary in-memory server. This mode is ideal for:
How it works:
Usage:
# Simply run your flow - Prefect ephemeral mode is automatic
docling-pipelines --flow-file sample_flows/quickstart/complete_pipeline_ollama.json
For production workloads and large-scale processing, Docling Pipelines supports Prefect’s distributed execution using work pools and workers.
When to use:
Deployment options:
For complete setup instructions, work pool configuration, and deployment guides, see:
Prefect Distributed Execution Guide
This guide covers: