Schema templates provide a flexible, reusable way to define OpenSearch index configurations for Docling Pipelines pipelines. Instead of manually configuring index schemas for each flow, you can use pre-built templates or create custom ones that automatically adapt to your pipeline’s requirements.
Schema templates are JSON files that define OpenSearch index configurations with placeholder values that are replaced at runtime. They provide:
Templates are stored in src/docpipe/core/operators/vectordb/schemas/ and referenced by path in the VectorDBOperator configuration.
Purpose: General-purpose schema for document storage with standard field types.
Use Cases:
Features:
Location: src/docpipe/core/operators/vectordb/schemas/default_schema.v1.json
Purpose: Template schema with custom content analyzer for text processing workflows.
Use Cases:
Features:
indexing_rules for declarative field-specific type mappingKey Feature: Shows how indexing_rules maps fields (e.g., content) to custom field types (content_text) with specialized analyzers, eliminating the need for manual feature_mappings configuration.
Location: src/docpipe/core/operators/vectordb/schemas/template_with_content_analyzer.v1.json
Add schema_template_path to your VectorDBOperator’s provider_config:
{
"operator_type": "docpipe.core.operators.vectordb.vectordb_operator.VectorDBOperator",
"operator_params": {
"provider": "opensearch",
"index_name": "my_documents",
"provider_config": {
"schema_template_path": "schemas/default_schema.v1.json",
"host": "localhost",
"port": 9200,
"engine": "faiss",
"algorithm": "hnsw",
"vector_dimension": 384
}
}
}
{
"provider_config": {
"schema_template_path": "schemas/template_with_content_analyzer.v1.json",
"host": "localhost",
"port": 9200,
"engine": "faiss",
"algorithm": "hnsw",
"vector_dimension": 768
}
}
If schema_template_path is not specified or the template cannot be loaded:
available_features configurationTemplates support the following placeholders that are replaced at runtime:
| Placeholder | Description | Example Values | Required |
|---|---|---|---|
__VECTOR_DIMENSION__ |
Vector embedding dimension | 384, 768, 1536 |
Yes |
__ENGINE__ |
KNN engine name | faiss, lucene, nmslib, jvector |
Yes |
__ALGORITHM__ |
KNN algorithm | hnsw, ivf |
Yes |
__SPACE_TYPE__ |
Similarity metric | l2, cosine, inner_product |
Yes |
__ENGINE_PARAMETERS__ |
Engine-specific parameters | {"ef_construction": 128, "m": 24} |
Yes |
{
"field_types": {
"vector": {
"type": "knn_vector",
"dimension": "__VECTOR_DIMENSION__",
"method": {
"name": "__ALGORITHM__",
"space_type": "__SPACE_TYPE__",
"engine": "__ENGINE__",
"parameters": "__ENGINE_PARAMETERS__"
}
}
}
}
When the template is loaded, placeholders are replaced with actual values from your configuration:
{
"field_types": {
"vector": {
"type": "knn_vector",
"dimension": 768,
"method": {
"name": "hnsw",
"space_type": "l2",
"engine": "faiss",
"parameters": {
"ef_construction": 128,
"m": 24
}
}
}
}
}
The indexing rules system enables flexible field-level customization without requiring full schema definitions. It allows you to override field types and properties for specific fields while using template-based field type definitions for the rest.
field_typesAllowlisted OpenSearch mapping properties that can be overridden:
analyzer, search_analyzer, normalizer: Text analysis configurationboost, copy_to: Relevance and field copyingindex, store, fields: Indexing behavior and sub-fieldssimilarity: Scoring algorithm selectionAdd an indexing_rules section to your schema template:
{
"schema_name": "my_schema",
"schema_version": 1,
"settings": { ... },
"field_types": {
"string": {
"type": "text",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
},
"content_text": {
"type": "text",
"analyzer": "content_analyzer",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
}
},
"indexing_rules": {
"content": {
"field_type": "content_text"
}
}
}
In this example, the content field uses the content_text field type instead of the default string type.
Override the field type for specific fields:
{
"indexing_rules": {
"title": {
"field_type": "string"
},
"content": {
"field_type": "content_text"
},
"summary": {
"field_type": "content_text"
}
}
}
Add or override specific properties:
{
"indexing_rules": {
"title": {
"field_type": "string",
"boost": 3.0,
"copy_to": ["all_text"]
},
"content": {
"field_type": "content_text",
"boost": 2.0,
"copy_to": ["all_text"]
}
}
}
Combine field type and property overrides:
{
"indexing_rules": {
"content": {
"field_type": "content_text",
"boost": 2.0,
"copy_to": ["all_text"],
"fields": {
"exact": {
"type": "keyword",
"normalizer": "lowercase"
}
}
}
}
}
When resolving field configurations, the system follows this priority:
indexing_rules for the feature name (e.g., content)indexing_rules for the mapped name (e.g., doc_content)system_type from available_features configurationSystem fields cannot be overridden via indexing rules:
_id_index_source_type_meta{
"schema_name": "document_chunks",
"schema_version": 1,
"settings": {
"index": {
"number_of_shards": 3,
"number_of_replicas": 1,
"knn": true
},
"analysis": {
"analyzer": {
"content_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "stop", "snowball"]
}
}
}
},
"field_types": {
"vector": {
"type": "knn_vector",
"dimension": "__VECTOR_DIMENSION__",
"method": {
"name": "__ALGORITHM__",
"space_type": "__SPACE_TYPE__",
"engine": "__ENGINE__",
"parameters": "__ENGINE_PARAMETERS__"
}
},
"string": {
"type": "text",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
},
"content_text": {
"type": "text",
"analyzer": "content_analyzer",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
}
},
"indexing_rules": {
"content": {
"field_type": "content_text",
"boost": 2.0
}
}
}
A schema template must include:
{
"schema_name": "my_custom_schema",
"schema_version": 1,
"settings": {
"index": {
"knn": true,
"number_of_shards": 2,
"number_of_replicas": 1
}
},
"field_types": {
"vector": {
"type": "knn_vector",
"dimension": "__VECTOR_DIMENSION__",
"method": {
"name": "__ALGORITHM__",
"space_type": "__SPACE_TYPE__",
"engine": "__ENGINE__",
"parameters": "__ENGINE_PARAMETERS__"
}
},
"string": {
"type": "text",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
}
}
}
{
"schema_name": "custom_analyzer_schema",
"schema_version": 1,
"settings": {
"index": {
"knn": true,
"number_of_shards": 3,
"number_of_replicas": 1
},
"analysis": {
"analyzer": {
"my_custom_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "stop", "snowball", "asciifolding"]
}
}
}
},
"field_types": {
"vector": {
"type": "knn_vector",
"dimension": "__VECTOR_DIMENSION__",
"method": {
"name": "__ALGORITHM__",
"space_type": "__SPACE_TYPE__",
"engine": "__ENGINE__",
"parameters": "__ENGINE_PARAMETERS__"
}
},
"analyzed_text": {
"type": "text",
"analyzer": "my_custom_analyzer",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
}
}
}
src/docpipe/core/operators/vectordb/schemas/my_schema.v1.json"schema_template_path": "schemas/my_schema.v1.json"Templates are validated before use. The following rules apply:
schema_name (string)schema_version (integer)settings (object)field_types (object)type: "knn_vector"dimension must be positive integermethod must include name, space_type, enginefaiss, lucene, nmslib, jvectorhnsw, ivfl2, cosine, inner_productHNSW Parameters:
ef_construction: 1-1000 (recommended: 100-512)m: 2-100 (recommended: 16-48)IVF Parameters:
nlist: 1-10000 (recommended: 100-1000)nprobe: 1-nlist (recommended: 8-128)The VectorDBOperator automatically normalizes and aggregates metadata columns.
Common column name variations are automatically mapped:
| Target Field | Source Aliases | Description |
|---|---|---|
source |
source, path |
Document source path |
page_count |
page_count, pages_processed |
Number of pages |
mimetype |
mimetype, mime_type, content_type |
MIME type |
Missing metadata fields are automatically derived:
name or source filenameextension using standard mappingsThese fields are automatically collected into a metadata object:
name: Document filenamesize: File size in bytescreated_time: Creation timestampmodified_time: Modification timestampsource: Document source pathmimetype: MIME typeextension: File extensionpage_count: Number of pagesInput:
{
"path": "/docs/report.pdf",
"name": "report.pdf",
"pages_processed": 10,
"content": "Document content..."
}
Output (automatically normalized):
{
"content": "Document content...",
"metadata": {
"source": "/docs/report.pdf",
"name": "report.pdf",
"page_count": 10,
"extension": "pdf",
"mimetype": "application/pdf"
}
}
document_chunks, product_catalog)my_schema.v1.json)vector_dimension matches your embedding modelfaiss: Best for large-scale similarity searchlucene: Good for smaller datasets, native to OpenSearchnmslib: Fast approximate searchef_construction and m based on accuracy/speed tradeoffl2: Euclidean distancecosine: Cosine similarity (normalized vectors)inner_product: Dot product similarity{
"operator_params": {
"provider": "opensearch",
"index_name": "documents",
"provider_config": {
"schema_template_path": "schemas/default_schema.v1.json",
"host": "localhost",
"port": 9200,
"engine": "faiss",
"algorithm": "hnsw",
"vector_dimension": 384
},
"feature_mappings": {
"doc_id_hash": "id",
"content": "content",
"embeddings": "embeddings"
},
"available_features": {
"embeddings": {
"type": "vector",
"available_for_vector_db": true
},
"content": {
"type": "string",
"available_for_vector_db": true
},
"doc_id_hash": {
"type": "string",
"available_for_vector_db": true
}
}
}
}
{
"operator_params": {
"provider": "opensearch",
"index_name": "document_chunks",
"provider_config": {
"schema_template_path": "schemas/template_with_content_analyzer.v1.json",
"host": "localhost",
"port": 9200,
"engine": "faiss",
"algorithm": "hnsw",
"vector_dimension": 768,
"engine_parameters": {
"ef_construction": 256,
"m": 32
}
},
"feature_mappings": {
"chunk_id": "id",
"content": "content",
"embeddings": "embeddings",
"doc_name": "doc_name",
"chunk_index": "chunk_index"
},
"available_features": {
"embeddings": {
"type": "vector",
"available_for_vector_db": true
},
"content": {
"type": "content_text",
"available_for_vector_db": true
},
"chunk_id": {
"type": "string",
"available_for_vector_db": true
},
"doc_name": {
"type": "string",
"available_for_vector_db": true
},
"chunk_index": {
"type": "integer",
"available_for_vector_db": true
}
}
}
}
{
"operator_params": {
"provider": "opensearch",
"index_name": "high_dim_vectors",
"provider_config": {
"schema_template_path": "schemas/default_schema.v1.json",
"host": "localhost",
"port": 9200,
"engine": "faiss",
"algorithm": "hnsw",
"vector_dimension": 1536,
"space_type": "cosine",
"engine_parameters": {
"ef_construction": 512,
"m": 48
}
}
}
}
Error: Schema template not found: schemas/my_schema.v1.json
Solutions:
src/docpipe/core/operators/vectordb/schemas/src/docpipe/core/operators/vectordb/Error: Schema validation failed: missing required field 'schema_name'
Solutions:
Error: Invalid vector dimension: __VECTOR_DIMENSION__
Solutions:
__PLACEHOLDER__Error: Vector dimension mismatch: expected 768, got 384
Solutions:
vector_dimension in config to embedding model outputnomic-embed-text: 768text-embedding-ada-002: 1536all-MiniLM-L6-v2: 384Error: Parameter 'ef_construction' out of valid range: 2000
Solutions:
ef_construction: 100-512m: 16-48Error: Unknown analyzer: my_custom_analyzer
Solutions:
settings.analysis section