The Milvus adapter provides vector database capabilities for the docpipe project, supporting Milvus Lite (embedded local .db file), standalone Milvus deployments, and IBM watsonx.data (wx.data) managed Milvus instances. This integration follows the hexagonal architecture pattern, implementing the VectorStorePort interface for seamless integration with the VectorDBOperator.
.db fileMilvusAdapter (VectorStorePort)
├── MilvusClient - Connection management
├── MilvusIndexManager - Collection/schema operations
└── MilvusBatchProcessor - Bulk data operations
adapters/outbound/milvus/client.py)
uri, not host/port directly)token="user:password"adapters/outbound/milvus/index_manager.py)
adapters/outbound/milvus/batch_processor.py)
adapters/outbound/milvus/adapter.py)
Milvus adapter supports five authentication types via the auth_type parameter:
uri must be a local .db file path. Only FLAT index type is supported.ibmlhapikey_ prefix and API key as password)No Docker, Podman, or Kubernetes required. Data is persisted to a local .db file.
Installation: milvus-lite is included as a core dependency — no separate installation needed.
Constraints:
FLAT index type is supported (other index types are rejected at initialisation with an actionable error)enable_micro_batching: false in global_config — Milvus Lite holds a single-process file lock{
"type": "vectordb",
"name": "store_locally",
"config": {
"provider": "milvus",
"doc_id_column": "doc_id_hash",
"create_index": true,
"add_sparse_vector": false,
"provider_config": {
"collection_name": "documents",
"auth_type": "lite",
"uri": "./data/milvus/documents.db",
"index_type": "FLAT",
"metric_type": "COSINE",
"batch_size": 100
}
},
"depends_on": ["embed"]
}
Path resolution: relative paths (e.g. ./data/milvus/docs.db) are resolved from the working directory where docling-pipelines is invoked. The parent directory is created automatically if it does not exist. Use an absolute path to avoid ambiguity.
Persistence: data survives process restart — reopen the same .db file path to query or extend the collection.
Migration path: swap auth_type: "lite" for auth_type: "standalone" (and add host/port) to move from Milvus Lite to a standalone deployment without any other flow changes.
{
"id": "a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d",
"name": "milvus_vector_store",
"operator": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"vector_dimension": 384,
"add_sparse_vector": false,
"available_features": {
"doc_id_hash": {
"type": "keyword",
"available_for_vector_db": true
},
"content": {
"type": "text",
"available_for_vector_db": true
},
"embeddings": {
"type": "vector",
"available_for_vector_db": true
}
},
"feature_mappings": {
"doc_id_hash": "pk",
"content": "text",
"embeddings": "vector"
},
"provider_config": {
"collection_name": "my_collection",
"auth_type": "standalone",
"host": "localhost",
"port": 19530,
"username": "root",
"password": "<your-milvus-password>",
"database": "default",
"index_type": "HNSW",
"metric_type": "L2",
"secure": false,
"index_parameters": {
"M": 16,
"efConstruction": 200
}
}
}
}
{
"id": "b2c3d4e5-f6a7-4b8c-9d0e-1f2a3b4c5d6e",
"name": "wxdata_grpc_store",
"operator": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"vector_dimension": 384,
"add_sparse_vector": false,
"provider_config": {
"collection_name": "my_collection",
"auth_type": "grpc",
"host": "your-wxdata-instance.cloud.ibm.com",
"port": 19530,
"username": "ibmlhapikey_your_username",
"password": "<your-api-key>",
"database": "default",
"index_type": "HNSW",
"metric_type": "COSINE",
"secure": true
},
"available_features": {
"doc_id_hash": {
"type": "keyword",
"available_for_vector_db": true
},
"content": {
"type": "text",
"available_for_vector_db": true
},
"embeddings": {
"type": "vector",
"available_for_vector_db": true
}
},
"feature_mappings": {
"doc_id_hash": "pk",
"content": "text",
"embeddings": "vector"
}
}
}
{
"id": "c3d4e5f6-a7b8-4c9d-0e1f-2a3b4c5d6e7f",
"name": "wxdata_uri_store",
"operator": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"vector_dimension": 384,
"add_sparse_vector": false,
"available_features": {
"doc_id_hash": {
"type": "keyword",
"available_for_vector_db": true
},
"content": {
"type": "text",
"available_for_vector_db": true
},
"embeddings": {
"type": "vector",
"available_for_vector_db": true
}
},
"feature_mappings": {
"doc_id_hash": "pk",
"content": "text",
"embeddings": "vector"
},
"provider_config": {
"collection_name": "my_collection",
"auth_type": "uri",
"uri": "https://ibmlhapikey_your_username:<your-api-key>@your-wxdata-instance.cloud.ibm.com:19530",
"database": "default",
"index_type": "HNSW",
"metric_type": "COSINE"
}
}
}
{
"id": "d4e5f6a7-b8c9-4d0e-1f2a-3b4c5d6e7f8a",
"name": "wxdata_token_store",
"operator": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"vector_dimension": 384,
"add_sparse_vector": false,
"available_features": {
"doc_id_hash": {
"type": "keyword",
"available_for_vector_db": true
},
"content": {
"type": "text",
"available_for_vector_db": true
},
"embeddings": {
"type": "vector",
"available_for_vector_db": true
}
},
"feature_mappings": {
"doc_id_hash": "pk",
"content": "text",
"embeddings": "vector"
},
"provider_config": {
"collection_name": "my_collection",
"auth_type": "token",
"host": "your-wxdata-instance.cloud.ibm.com",
"port": 19530,
"username": "your_username",
"token": "your-iam-token",
"database": "default",
"index_type": "HNSW",
"metric_type": "COSINE"
}
}
}
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
auth_type |
string | Yes | - | Authentication type: lite, standalone, grpc, uri, or token |
uri |
string | Conditional* | - | Local .db file path (for lite) or full remote URI (for uri auth_type) |
host |
string | Conditional** | - | Milvus server host |
port |
integer | Conditional** | 19530 | Milvus server port |
token |
string | Conditional*** | - | IAM token (for token auth_type) |
username |
string | Conditional** | - | Username for authentication |
password |
string | Conditional***** | - | Password/API key for authentication |
database |
string | No | “default” | Database name (not used for lite) |
secure |
bool | No | false | Enable TLS/SSL (not used for lite) |
| auth_type | Required Parameters | Description |
|---|---|---|
lite |
uri (local .db path) |
Milvus Lite — embedded local database. No external service. Only FLAT index supported. |
standalone |
host, port |
Local Milvus. Optional: username, password |
grpc |
host, port, username, password |
IBM wx.data with gRPC. Username must have ibmlhapikey_ prefix, password is API key |
uri |
uri |
Pre-constructed URI with embedded API key (format: https://ibmlhapikey_<username>:<api-key>@<host>:<port>) |
token |
host, port, username, token |
IAM token-based. Constructs URI internally (format: https://ibmlhtoken_<username>:<token>@<host>:<port>) |
Required for lite and uri auth types
Required for standalone, grpc, and token auth types
***Required for token auth_type
**Required for grpc and token auth types
****Required for grpc auth_type (should be API key)
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
index_type |
string | No | “HNSW” | Index type (HNSW, IVF_FLAT, etc. Auto-set to SPARSE_INVERTED_INDEX in sparse mode) |
metric_type |
string | No | “L2” | Similarity metric (L2, IP, COSINE for dense; BM25 for sparse) |
index_parameters |
object | No | {} | Index-specific parameters |
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
provider |
string | Yes | - | Must be “milvus” |
provider_config.collection_name |
string | Yes | - | Name of the Milvus collection |
embeddings_column |
string | No | “embeddings” | Column containing vectors |
vector_dimension |
integer | No | 384 | Vector dimension (dense mode only) |
add_sparse_vector |
boolean | No | false | Use BM25 sparse vectors instead of dense |
batch_size |
integer | No | 100 | Batch size for operations |
primary_key_field |
string | No | “pk” | Primary key field name |
auto_id |
boolean | No | false | Auto-generate IDs |
M: Number of connections (default: 16)efConstruction: Build-time search depth (default: 200){
"index_type": "HNSW",
"index_parameters": {
"M": 16,
"efConstruction": 200
}
}
nlist: Number of clusters (default: 128){
"index_type": "IVF_FLAT",
"index_parameters": {
"nlist": 128
}
}
nlist: Number of clusters (default: 128)nlist: Number of clusters (default: 128)m: Number of subquantizers (default: 8)nbits: Bits per subquantizer (default: 8)| Metric | Description | Use Case | Vector Type |
|---|---|---|---|
L2 |
Euclidean distance | General purpose, normalized vectors | Dense only |
IP |
Inner product | Cosine similarity with normalized vectors | Dense only |
COSINE |
Cosine similarity | Text embeddings, semantic search | Dense only |
BM25 |
BM25 text scoring | Sparse vectors, keyword search | Sparse only (required) |
Docling Pipelines provides two example pipeline flows demonstrating different Milvus deployment scenarios:
Example Flow: sample_flows/vectordb/milvus_integration.json
Use Case: Local development, testing, or self-hosted Milvus
Key Configuration:
{
"provider": "milvus",
"doc_id_column": "doc_id_hash",
"embeddings_column": "embeddings",
"add_sparse_vector": true,
"create_index": true,
"provider_config": {
"collection_name": "sample_documents_sparse",
"auth_type": "standalone",
"host": "localhost",
"port": 19530,
"secure": false,
"username": "root",
"password": "<YOUR_PASSWORD>",
"metric_type": "BM25",
"index_type": "SPARSE_INVERTED_INDEX"
}
}
Features:
secure parameter not needed)Example Flow: sample_flows/vectordb/milvus_integration.json
Use Case: Enterprise deployment with IBM watsonx.data managed Milvus
Key Configuration:
{
"provider": "milvus",
"doc_id_column": "doc_id_hash",
"embeddings_column": "embeddings",
"add_sparse_vector": false,
"create_index": true,
"provider_config": {
"collection_name": "sample_documents_collection_dense",
"auth_type": "grpc",
"host": "<YOUR_WXDATA_HOST>",
"port": "<YOUR_PORT>",
"username": "<YOUR_USERNAME>",
"password": "<YOUR_API_KEY>",
"secure": true,
"metric_type": "L2",
"index_type": "HNSW"
}
}
CRITICAL REQUIREMENTS:
⚠️ secure: true is MANDATORY for watsonx.data connections
IBM watsonx.data is a cloud-managed service that requires encrypted communication. The secure: true parameter enables SSL/TLS encryption for gRPC connections, which is essential for:
Without secure: true, the connection attempt will fail with HTTPConnectionError because watsonx.data’s gRPC endpoint rejects unencrypted connections.
Additional Requirements:
GRPC_DNS_RESOLVER=native environment variable before running the pipeline. This resolves a known gRPC DNS resolution issue on macOS when connecting through VPNs or to cloud services.Dense vectors are traditional embeddings generated by models like Ollama, OpenAI, or HuggingFace. They require an upstream embeddings operator.
Configuration:
{
"embeddings_column": "embeddings",
"add_sparse_vector": false,
"vector_dimension": 768,
"feature_mappings": {
"embeddings": "dense_embeddings"
},
"provider_config": {
"index_type": "HNSW",
"metric_type": "COSINE"
}
}
Pipeline Requirements:
Sparse vectors are automatically generated from text content using Milvus’s built-in BM25 function during insertion. The pipeline still requires an embeddings operator for dense embeddings, but the BM25 function generates additional sparse vectors from the text content.
Configuration:
{
"embeddings_column": "embeddings",
"add_sparse_vector": true,
"feature_mappings": {
"doc_id_hash": "pk",
"content": "text",
"embeddings": "vector",
"sparse_embeddings": "sparse_vector"
},
"provider_config": {
"auth_type": "standalone",
"index_type": "SPARSE_INVERTED_INDEX",
"metric_type": "BM25"
}
}
Key Differences:
How It Works:
For a complete working example, see sample_flows/vectordb/milvus_integration.json.
Key Configuration:
{
"type": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"vector_dimension": 384,
"add_sparse_vector": false,
"feature_mappings": {
"doc_id_hash": "pk",
"embeddings": "vector",
"content": "text"
},
"provider_config": {
"collection_name": "documents",
"auth_type": "standalone",
"host": "localhost",
"port": 19530,
"username": "root",
"password": "<YOUR_PASSWORD>",
"database": "default",
"index_type": "HNSW",
"metric_type": "L2",
"index_parameters": {
"M": 16,
"efConstruction": 256
}
}
}
}
Pipeline Structure:
For a complete working example, see sample_flows/vectordb/milvus_integration.json.
Key Configuration:
{
"type": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"add_sparse_vector": true,
"feature_mappings": {
"doc_id_hash": "pk",
"embeddings": "vector",
"sparse_embeddings": "sparse_vector",
"content": "text"
},
"provider_config": {
"collection_name": "documents_sparse",
"auth_type": "standalone",
"host": "localhost",
"port": 19530,
"username": "root",
"password": "<YOUR_PASSWORD>",
"database": "default",
"index_type": "SPARSE_INVERTED_INDEX",
"metric_type": "BM25"
}
}
}
Pipeline Structure:
Note: Sparse vector pipeline INCLUDES embeddings operator - both dense embeddings and BM25 sparse vectors are stored together.
For a comprehensive Python example demonstrating Milvus operator usage, see examples/milvus_integration_example.py. This example includes:
Run the example:
# Set up environment variables in .env file
export MILVUS_HOST=localhost
export MILVUS_PORT=19530
export MILVUS_USERNAME=root
export MILVUS_PASSWORD=your-milvus-password
# Run the example
python examples/milvus_integration_example.py
{
"id": "b8c9d0e1-f2a3-4b4c-5d6e-7f8a9b0c1d2e",
"name": "wxdata_vector_store",
"operator": "vectordb",
"config": {
"provider": "milvus",
"embeddings_column": "embeddings",
"vector_dimension": 768,
"add_sparse_vector": false,
"available_features": {
"doc_id_hash": {
"type": "keyword",
"available_for_vector_db": true
},
"content": {
"type": "text",
"available_for_vector_db": true
},
"embeddings": {
"type": "vector",
"available_for_vector_db": true
}
},
"feature_mappings": {
"doc_id_hash": "pk",
"content": "text",
"embeddings": "vector"
},
"provider_config": {
"collection_name": "enterprise_docs",
"uri": "https://milvus.wxdata.ibm.com:19530",
"token": "${WXDATA_TOKEN}",
"db_name": "production",
"secure": true,
"index_type": "HNSW",
"metric_type": "COSINE",
"index_parameters": {
"M": 32,
"efConstruction": 400
}
}
}
}
Feature mappings define which PyArrow table columns are stored in Milvus:
{
"available_features": {
"embeddings": {
"type": "vector",
"available_for_vector_db": true
},
"content": {
"type": "text",
"available_for_vector_db": true
},
"doc_id_hash": {
"type": "keyword",
"available_for_vector_db": true
}
},
"feature_mappings": {
"embeddings": "vector_field",
"content": "text_content",
"doc_id_hash": "document_id"
}
}
Problem: Cannot connect to Milvus
MilvusDB Error: Failed to connect using standalone auth_type: Fail connecting to server on localhost:19530, illegal connection params or server unavailable
Solutions:
curl http://localhost:9091/healthz (should return OK)docker compose -f docker/docker-compose.milvus.yml up -dtelnet localhost 19530root / Milvus)Problem: Connection fails on macOS when accessing Milvus through VPN or watsonx.data
MilvusException: (code=2, message=Fail connecting to server on your-host-name:443,
illegal connection params or server unavailable)
Cause: macOS gRPC DNS resolver incompatibility with certain network configurations
Solution: Set the GRPC_DNS_RESOLVER environment variable before running your application:
export GRPC_DNS_RESOLVER=native
Or in Python code before importing Milvus libraries:
import os
os.environ["GRPC_DNS_RESOLVER"] = "native"
# Then import and use docpipe
from docpipe.core.operators.vectordb import VectorDBOperator
Verification Steps:
GRPC_DNS_RESOLVER=native environment variableNote: This issue primarily affects:
Problem: Collection creation fails
Failed to create collection 'my_collection'
Solutions:
Problem: Index parameters invalid
Invalid index type 'HNSW'. Supported: [...]
Solutions:
Problem: Out of memory during indexing
Memory allocation failed
Solutions:
cd /path/to/docling-pipelines
export PYTHONPATH="$(pwd)/src:${PYTHONPATH}"
uv run pytest tests/unit/operators/vectordb/test_milvus_client.py -v
The full end-to-end integration test (ingest → extract → chunk → embed → Milvus Lite) requires no running server:
source .venv/bin/activate
pytest tests/integration/test_ingest_extract_chunk_embed_milvus_lite.py -v
This test covers:
test_full_pipeline_with_feature_mappings — E2E flow; verifies row count, schema fields, and content stored correctlytest_incremental_reindex_no_duplicates — runs twice against the same .db; confirms row count is stable (no duplicates)test_unsupported_index_type_rejected — confirms non-FLAT index types are rejected at initialisation with a clear error# Start Milvus (Docker)
docker run -d --name milvus \
-p 19530:19530 \
-p 9091:9091 \
milvusdb/milvus:latest
# Run tests
uv run pytest tests/integration/milvus/ -v
To migrate from OpenSearch to Milvus:
{
"provider": "milvus", // Changed from "opensearch"
"provider_config": {
"collection_name": "my_collection",
"auth_type": "standalone", // Required: standalone, grpc, uri, or token
"host": "localhost",
"port": 19530, // Changed from 9200
"username": "root",
"password": "<YOUR_PASSWORD>",
"index_type": "HNSW", // Instead of "engine"
"metric_type": "L2" // Instead of "space_type"
}
}
faiss/hnsw → Milvus HNSWfaiss/ivf → Milvus IVF_FLATlucene/hnsw → Milvus HNSWl2 → Milvus L2cosine → Milvus COSINEinner_product → Milvus IPFor issues or questions: