This document describes the current LLM adapter architecture used across all operators in docpipe.
All LLM integrations in Docling Pipelines share a common set of provider adapters, ports, and a single factory. The goals are:
Most operators follow the direct port pattern. Entity extraction is the one exception and uses an operator-specific adapter layer on top of the shared infrastructure — see Section 2: Entity Extraction for the rationale.
Operator
↓
Service Layer (business logic, optional for simple operators)
↓ calls
LLMAdapterFactory
↓ returns
Common Port Interface (LLMInferencePort / LLMEmbeddingPort / TextDetectionPort)
↑ implemented by
Consolidated Provider Adapters (WatsonXAdapter, LiteLLMAdapter)
provider_config from the flow JSON is passed directly and unchanged to LLMAdapterFactory.
It only ever contains connection-level fields (api_key, api_base, url, container_kind, etc.).
ExtractOperator
↓
EntityExtractionAdapterFactory (operator-specific factory)
↓ creates via registry
LiteLLMEntityAdapter / WatsonxEntityAdapter (thin registered subclasses)
↓ extends
LLMEntityAdapter (shared base: parallelism, schema building, response normalisation)
↓ calls
LLMAdapterFactory.create_inference_adapter()
↓ returns
LLMInferencePort (LiteLLMAdapter or WatsonXAdapter)
This extra layer exists because entity extraction’s transform() contract operates on a full
pa.Table with per-document parallelism — it cannot be expressed as a plain chat() call.
See Section 2 for full details.
Ollama is accessed through LiteLLM using the OpenAI-compatible API:
api_base: http://localhost:11434/v1
model_id: openai/<model_name>
Three shared interfaces cover all LLM capabilities:
| Port | Purpose | Used by |
|---|---|---|
LLMInferencePort |
Chat/generation APIs | Classification, Entity Extraction, Summarization, PII/HAP (LLM mode) |
LLMEmbeddingPort |
Embedding generation | Embeddings Operator |
TextDetectionPort |
Specialized detection APIs | PII/HAP (WatsonX mode) |
A single factory handles all adapter creation:
| Factory | Responsibility |
|---|---|
LLMAdapterFactory |
Inference, Embeddings, and Text Detection APIs |
File
src/docpipe/core/ports/llm_inference_port.py
Purpose: Common interface for chat() and generate() calls.
Used by: Classification, Entity Extraction, Summarization, PII/HAP (LLM mode)
File
src/docpipe/core/ports/llm_embedding_port.py
Purpose: Common interface for generate_embeddings(), generate_embeddings_batch(), get_embedding_dimension().
Used by: Embeddings Operator
File
src/docpipe/core/ports/text_detection_port.py
Purpose: Specialized interface for PII/HAP detection via WatsonX detection APIs.
Used by: PII/HAP Operator (WatsonX path only)
Each provider has a single consolidated adapter that implements one or more port interfaces.
File
src/docpipe/core/adapters/watsonx/watsonx_adapter.py
LLMInferencePort (chat, generate)
LLMEmbeddingPort (generate_embeddings, generate_embeddings_batch, get_embedding_dimension)
TextDetectionPort (detect, detect_entities, detect_entities_batch)
model_name parameterresponse_format passed as method parameter, not in __init__/ml/v1/text/detection API for PII/HAPfrom docpipe.core.adapters import WatsonXAdapter
adapter = WatsonXAdapter(
api_key="<your-watsonx-api-key>", # pragma: allowlist secret
container_id="<your-project-id>",
api_base="https://us-south.ml.cloud.ibm.com",
model_name="ibm/granite-13b-chat-v2"
)
# Inference with JSON response
response = adapter.chat(
messages=[{"role": "user", "content": "Classify this"}],
response_format={"type": "json_object"}
)
# Embeddings (model override)
embeddings = adapter.generate_embeddings(
model_name="ibm/slate-125m-english-rtrvr",
text="Sample text"
)
# Text detection
result = adapter.detect(text="Check for PII")
File
src/docpipe/core/adapters/litellm/litellm_adapter.py
LLMInferencePort (chat, generate)
LLMEmbeddingPort (generate_embeddings, generate_embeddings_batch, get_embedding_dimension)
gpt-4, text-embedding-3-smallclaude-3-opus-20240229openai/llama2, openai/nomic-embed-text (via OpenAI-compatible API)huggingface/sentence-transformers/all-MiniLM-L6-v2model_name parameterresponse_format passed as method parameter (not hardcoded)from docpipe.core.adapters import LiteLLMAdapter
adapter = LiteLLMAdapter(
model_name="openai/llama2",
api_key="ollama", # pragma: allowlist secret
api_base="http://localhost:11434/v1"
)
response = adapter.chat(
messages=[{"role": "user", "content": "Classify this"}],
response_format={"type": "json_object"}
)
embeddings = adapter.generate_embeddings(
model_name="openai/nomic-embed-text",
text="Sample text"
)
File
src/docpipe/core/adapters/llm_adapter_factory.py
Single unified factory for creating all types of LLM adapters across all providers.
# Create inference adapter (LLMInferencePort)
create_inference_adapter(provider, model_id, provider_config)
# Create embedding adapter (LLMEmbeddingPort)
create_embedding_adapter(provider, model_id, provider_config)
# Create text detection adapter (TextDetectionPort)
create_text_detection_adapter(provider, model_id, provider_config)
# Query supported providers for a capability
get_supported_providers(capability="inference")
| Provider | Inference | Embeddings | Text Detection | Notes |
|---|---|---|---|---|
watsonx |
✅ | ✅ | ✅ | IBM watsonx.ai — single consolidated adapter |
litellm |
✅ | ✅ | ❌ | 100+ providers (Ollama, OpenAI, Anthropic, etc.) |
huggingface |
❌ | ✅ | ❌ | Local or API-based HuggingFace models |
from docpipe.core.adapters import LLMAdapterFactory
# WatsonX inference
adapter = LLMAdapterFactory.create_inference_adapter(
provider="watsonx",
model_id="ibm/granite-13b-chat-v2",
provider_config={
"api_key": "<your-watsonx-api-key>", # pragma: allowlist secret
"url": "https://us-south.ml.cloud.ibm.com",
"container_id": "<your-project-id>",
"container_kind": "project"
}
)
# LiteLLM embedding (Ollama)
adapter = LLMAdapterFactory.create_embedding_adapter(
provider="litellm",
model_id="openai/nomic-embed-text",
provider_config={
"api_base": "http://localhost:11434/v1",
"api_key": "ollama" # pragma: allowlist secret
}
)
# HuggingFace embedding (local)
adapter = LLMAdapterFactory.create_embedding_adapter(
provider="huggingface",
model_id="sentence-transformers/all-MiniLM-L6-v2",
provider_config={"use_local": True, "device": "cpu"}
)
# WatsonX text detection
adapter = LLMAdapterFactory.create_text_detection_adapter(
provider="watsonx",
model_id="ibm/granite-13b-chat-v2",
provider_config={
"api_key": "<your-watsonx-api-key>", # pragma: allowlist secret
"url": "https://us-south.ml.cloud.ibm.com",
"container_id": "<your-project-id>",
"container_kind": "project"
}
)
LiteLLM provides unified access to multiple providers via model ID prefixes:
openai/<model> with api_base: http://localhost:11434/v1huggingface/<model> with HuggingFace API keygpt-4, text-embedding-3-small, etc.claude-3-opus-20240229, etc.provider_config Referenceprovider_config fields are scoped by operator, not global. Operators that pass provider_config
directly to LLMAdapterFactory (classification, PII/HAP, summarization, embeddings) only need
connection-level fields. Entity extraction is the exception — see the note below.
Used by: Classification, PII/HAP, Summarization, Embeddings
These operators pass provider_config unchanged to LLMAdapterFactory.
The Pydantic models live in src/docpipe/core/operators/shared/llm_provider_config.py.
| Provider | Model | Key fields |
|---|---|---|
litellm |
LLMProviderConfig |
model_id, api_base, api_key |
watsonx |
WatsonxProviderConfig |
model_id, url, api_key, container_kind, container_id |
Used by: ExtractOperator (entity_extraction block only)
Entity extraction puts temperature and max_tokens inside provider_config rather than as
separate operator-level fields. The Pydantic models live in
src/docpipe/core/operators/extract/adapters/outbound/entity_extraction/llm_entity_config.py.
| Provider | Model | Key fields |
|---|---|---|
litellm |
LLMEntityConfig |
model_id, api_base, api_key, temperature, max_tokens |
watsonx |
WatsonxEntityConfig |
above + url, container_kind, project_id |
File
src/docpipe/core/operators/quality/classification/document_classifier.py
DocumentClassifierOperator
↓
ClassificationService (src/docpipe/core/operators/quality/classification/classification_service.py)
↓ calls LLMAdapterFactory.create_inference_adapter()
↓ holds
LLMInferencePort
↑ implemented by
WatsonXAdapter / LiteLLMAdapter
File
src/docpipe/core/operators/quality/classification/classification_service.py
Responsibilities: prompt building, LLM call, response parsing, validation.
Actual constructor signature:
ClassificationService(
*,
model_id: str,
provider_name: str,
provider_config: dict[str, Any] | None = None,
temperature: float = 0.0,
max_tokens: int = 500,
)
provider_config is passed directly to LLMAdapterFactory. temperature and max_tokens
are separate parameters, not fields inside provider_config.
{
"type": "document_classifier",
"config": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/granite4:latest",
"api_base": "http://localhost:11434/v1",
"api_key": "${OLLAMA_API_KEY}"
},
"confidence_threshold": 7.0,
"doc_column": "content",
"output_column": "document_type"
}
}
File
src/docpipe/core/operators/extract/extract_operator.py
Entity extraction cannot follow Pattern A (direct port) for two reasons:
Unit of work: The EntityExtractionPort.transform() contract takes a full pa.Table and
returns tuple[list[pa.Table], dict]. This is the operator contract — it involves per-document
parallelism, expand_extracted_data flag handling, and output_column routing. None of that
can be expressed as a plain LLMInferencePort.chat() call.
Per-provider config schemas: The EntityExtractionAdapterFactory advertises a different
Pydantic config schema per provider for UI/metadata generation (get_metadata()). Each
registered adapter class owns its get_config_schema() method. A single adapter class cannot
own two different schemas simultaneously.
ExtractOperator
↓
EntityExtractionAdapterFactory
↓ routes through registry (@register_entity_extraction_adapter)
LiteLLMEntityAdapter / WatsonxEntityAdapter (thin subclasses — own ADAPTER_NAME + get_config_schema())
↓ extends
LLMEntityAdapter (shared base — parallelism, schema building, JSON normalisation, truncation)
↓ calls
LLMAdapterFactory.create_inference_adapter()
↓ returns
LLMInferencePort (LiteLLMAdapter or WatsonXAdapter)
For the docling mode, DoclingEntityAdapter is registered directly without going through
LLMAdapterFactory — it uses Docling’s own template extraction pipeline.
LLMEntityAdapter with ADAPTER_NAME, get_config_schema(), and the
@register_entity_extraction_adapter decorator.provider_config fields.entity_extraction/__init__.py so the decorator fires on package import.LLMAdapterFactory.INFERENCE_PROVIDERS — the
two registries are independent and both must know about the provider.provider_config convention noteEntity extraction puts temperature and max_tokens inside provider_config. This differs
from classification and summarization, where these are top-level operator config fields. This is
a historical inconsistency, not an intentional design difference.
{
"type": "extract_operator",
"config": {
"entity_extraction": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/granite4:latest",
"api_base": "http://localhost:11434/v1",
"api_key": "${OLLAMA_API_KEY}",
"temperature": 0.0,
"max_tokens": 2000
},
"max_doc_chars": 8000,
"expand_extracted_data": false
}
}
}
File
src/docpipe/core/operators/quality/pii_and_hap/pii_and_hap_annotator.py
Each provider adapter self-registers with PIIAndHAPDetectionFactory via
@register_pii_and_hap_detection_adapter and encapsulates its own detection path behind
the unified PIIAndHAPDetectionPort interface:
PIIAndHAPAnnotator._initialize_pii_hap_service()
↓
PIIAndHAPDetectionFactory.create(provider)
↓
WatsonxPIIAndHAPAdapter | LiteLLMPIIAndHAPAdapter
(each implements PIIAndHAPDetectionPort.detect() + validate())
↓ injected
PIIHAPService (thin wrapper — no if/elif, no factory calls)
The two adapters use different underlying ports because of provider API differences:
WatsonxPIIAndHAPAdapter: WatsonX exposes a specialized /ml/v1/text/detection API
(not a standard chat API), so it wraps TextDetectionPort.LiteLLMPIIAndHAPAdapter: LiteLLM uses standard prompt-based inference, so it wraps
LLMInferencePort.File
src/docpipe/core/operators/quality/pii_and_hap/services/pii_hap_service.py
Actual constructor signature:
PIIHAPService(*, adapter: PIIAndHAPDetectionPort)
Receives a pre-built PIIAndHAPDetectionPort instance via constructor injection. Calls
adapter.validate() at construction time (fail-fast) then delegates all detection to
adapter.detect(). Contains no provider branching.
{
"type": "pii_and_hap",
"config": {
"provider": "watsonx",
"provider_config": {
"model_id": "ibm/granite-13b-chat-v2",
"url": "https://us-south.ml.cloud.ibm.com",
"api_key": "${WATSONX_API_KEY}",
"container_kind": "project",
"container_id": "${WATSONX_CONTAINER_ID}"
}
}
}
File
src/docpipe/core/operators/functional/embeddings/embeddings_operator.py
EmbeddingsOperator
↓ calls LLMAdapterFactory.create_embedding_adapter()
↓ holds
LLMEmbeddingPort
↑ implemented by
WatsonXAdapter / LiteLLMAdapter / HuggingFaceAdapter
No service layer — the operator calls the port directly. This is appropriate because embeddings
generation has no business logic beyond calling generate_embeddings_batch().
{
"type": "embeddings",
"config": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/nomic-embed-text",
"api_base": "http://localhost:11434/v1",
"api_key": "${OLLAMA_API_KEY}"
}
}
}
File
src/docpipe/core/operators/functional/chunker.py
src/docpipe/core/operators/functional/summarization_service.py
ChunkerOperator
↓ (when summarization is enabled)
SummarizationService (src/docpipe/core/operators/functional/summarization_service.py)
↓ calls LLMAdapterFactory.create_inference_adapter()
↓ holds
LLMInferencePort
↑ implemented by
WatsonXAdapter / LiteLLMAdapter
Summarization is opt-in. When the summarization config block is absent, no LLM adapter is
created and the operator runs as a pure text chunker.
Actual constructor signature:
SummarizationService(
*,
llm_adapter: LLMInferencePort,
summary_sentences: int = 3,
summary_max_words: int = 100,
overlap_ratio: float = 0.1,
max_length: int = 8192,
)
Summarization config lives inside the summarization sub-key, not at the operator root:
{
"type": "chunker",
"config": {
"chunk_size": 512,
"summarization": {
"provider": "litellm",
"provider_config": {
"model_id": "openai/granite4:latest",
"api_base": "http://localhost:11434/v1",
"api_key": "${OLLAMA_API_KEY}"
}
}
}
}
One consolidated adapter per provider instead of provider-specific adapters in every operator.
| Layer | Responsibility |
|---|---|
| Operator | Workflow orchestration |
| Service | Business logic (prompts, parsing, etc.) |
| Adapter | Provider protocol translation |
| Port | Interface contract |
LLMInferencePort, LLMEmbeddingPort) to test services in isolationAdding a new provider for classification, PII/HAP, summarization, or embeddings requires:
LLMInferencePort or LLMEmbeddingPort)LLMAdapterFactoryNo operator changes needed for these operators.
Adding a new LLM-based provider for entity extraction additionally requires a thin subclass in the entity extraction adapter package. See Section 2 — Adding a new LLM-based entity extraction provider.