docling-pipelines

Docling Pipelines Advanced Configuration Guide

This guide covers advanced configuration topics for production deployments, distributed execution, and performance optimization.

Table of Contents

  1. Job Stats Storage Configuration
  2. Incremental Metadata Configuration
  3. Execution Models

Job Stats Storage Configuration

Docling Pipelines supports pluggable job stats storage for job runs, node execution state, and micro-batch progress tracking.

Available Backends

Backend Selection

Job-management components are wired through JobManagementFactory. Backend selection is controlled by environment variables and defaults.

The main user-facing configuration is controlled via environment variables:

job_management:
  framework:
    type: default
    config: {}
  store:
    type: filesystem
    config:
      base_dir: ./data/job_stats_store_data

This allows users to configure:

Common overrides include:

Effective precedence for job-management runtime selection is:

  1. explicit environment variables
  2. built-in defaults in JobManagementFactory

Filesystem Storage Guidance

Filesystem storage is useful for single-host execution and simple local testing.

Important requirements:

DuckDB Guidance

DuckDB is an embedded analytical database that provides persistent storage without requiring a server. It’s ideal for local development and testing.

Use DuckDB when:

Configuration example:

job_management:
  store:
    type: duckdb
    config:
      database_path: ./data/duckdb/job_stats.duckdb

Important notes:

PostgreSQL Guidance

PostgreSQL is the recommended backend for multi-process and distributed execution because it provides durable shared storage and stronger concurrency behavior than file-backed JSON storage.

Use PostgreSQL when:

Metadata Aggregation Maintenance

Node stats are aggregated on the read path, not in the storage adapter. When operators add new metadata fields, maintainers must review DEFAULT_STRATEGIES and update it if the field should not use the default LAST aggregation behavior.

See docs/internals/NODE_METADATA_AGGREGATION_STRATEGY.md for the maintainer workflow.

Distributed Execution and Work Pool Environment Inheritance

For distributed Prefect execution, work pool runtime configuration is modeled in work_pool_config.py and applied by WorkPoolAdapter.

Important behavior:

This makes it possible to configure job management via environment variables per environment or per deployment.

For full distributed execution examples and work-pool-specific configuration, see DISTRIBUTED_EXECUTION_GUIDE.md.


Incremental Metadata Configuration

Docling Pipelines supports incremental processing to avoid reprocessing unchanged input data. Incremental metadata stores processing state such as file identity and modification information so ingest operators can determine whether an item is new, changed, or already processed.

Configuration

Use the incremental_metadata section in docling-pipelines-config.yaml:

incremental_metadata:
  storage:
    type: "filesystem"  # Options: filesystem, postgresql
    config:
      base_dir: "./data"
      lock_timeout: 30.0

This configuration is the single source of truth for incremental metadata backend selection and runtime settings.

Storage Backends

JSON

JSON storage is the default backend. It is file-based, easy to inspect, and suitable for development or smaller single-host deployments.

incremental_metadata:
  storage:
    type: "json"
    config:
      base_dir: "./data/incremental_metadata"
      lock_timeout: 30.0

Use JSON when:

Parquet

Parquet storage uses a columnar file format and is better suited to larger datasets or analytics-oriented workflows.

incremental_metadata:
  storage:
    type: "parquet"
    config:
      base_dir: "./data/incremental_metadata"
      lock_timeout: 30.0

Use Parquet when:

PostgreSQL

PostgreSQL is the recommended backend for production deployments that require stronger concurrency behavior and durable centralized storage.

incremental_metadata:
  storage:
    type: "postgresql"
    config:
      base_dir: "./data"
      lock_timeout: 30.0
  postgres:
    host: "${POSTGRES_HOST:-localhost}"
    port: 5432
    database: "${POSTGRES_DB:-docpipe}"
    user: "${POSTGRES_USER:-docpipe_user}"
    password: "${POSTGRES_PASSWORD}"
    schema: "incremental_metadata"

Use PostgreSQL when:

Environment Variables for Sensitive Data

Use environment variable substitution in docling-pipelines-config.yaml for credentials and deployment-specific values.

incremental_metadata:
  storage:
    type: "postgresql"
    config:
      base_dir: "${DATA_DIR:-./data}"
      lock_timeout: "${LOCK_TIMEOUT:-30.0}"
  postgres:
    host: "${INCR_META_DB_HOST:-localhost}"
    port: "${INCR_META_DB_PORT:-5432}"
    database: "${INCR_META_DB_NAME:-docpipe}"
    user: "${INCR_META_DB_USER:-docpipe_user}"
    password: "${INCR_META_DB_PASSWORD}"
    schema: "${INCR_META_DB_SCHEMA:-incremental_metadata}"

Guidance:

See docling-pipelines-config.yaml.example for complete backend examples and environment variable patterns.


Execution Models

Docling Pipelines uses Prefect as its orchestration engine and supports two Prefect execution modes:

Ephemeral Mode (Default)

By default, Docling Pipelines runs Prefect in ephemeral mode with a temporary in-memory server. This mode is ideal for:

How it works:

Usage:

# Simply run your flow - Prefect ephemeral mode is automatic
docling-pipelines --flow-file sample_flows/quickstart/complete_pipeline_ollama.json

Distributed Execution with Prefect Work Pools (Optional)

For production workloads and large-scale processing, Docling Pipelines supports Prefect’s distributed execution using work pools and workers.

When to use:

Deployment options:

  1. Local POC: Test distributed patterns with Prefect server and workers on a single machine
  2. Docker Compose: Multi-worker setup with containerized Prefect infrastructure

For complete setup instructions, work pool configuration, and deployment guides, see:

Prefect Distributed Execution Guide

This guide covers: