The ACL Document Retrieval API provides secure, user-scoped access to documents stored in OpenSearch. All document access is automatically filtered based on the allowed_users field, ensuring users can only retrieve documents they are authorized to view.
allowed_users fieldsrc/docpipe/api/services/opensearch_service.py)
src/docpipe/api/services/acl_query_builder.py)
allowed_users filter into all queriessrc/docpipe/api/dto/document_dto.py)
DocumentResponse: Document retrieval response modelDocumentSearchRequest: Search request parametersDocumentSearchResponse: Search results with paginationsrc/docpipe/api/routes/documents.py)
GET /api/v1/documents/{document_id}: Retrieve single documentPOST /api/v1/documents/search: Search documentsEndpoint: GET /api/v1/documents/{document_id}
Authentication: Required (JWT Bearer token)
Description: Retrieve a single document by its ID. Access is granted only if the authenticated user’s username is present in the document’s allowed_users field.
Request:
curl -X GET "http://localhost:8080/api/v1/documents/doc-123" \
-H "Authorization: Bearer YOUR_JWT_TOKEN"
Response (200 OK):
{
"id": "doc-123",
"content": "Document content here...",
"title": "Sample Document",
"metadata": {
"category": "tech",
"author": "John Doe"
},
"created_at": "2026-05-01T10:00:00Z",
"updated_at": "2026-05-15T14:30:00Z"
}
Error Responses:
401 Unauthorized: Missing or invalid JWT token404 Not Found: Document not found OR user not authorized (security by obscurity)503 Service Unavailable: OpenSearch service unavailableEndpoint: POST /api/v1/documents/search
Authentication: Required (JWT Bearer token)
Description: Search documents with full-text search, filters, sorting, and pagination. Results are automatically filtered to only include documents where the authenticated user is in allowed_users.
Request:
curl -X POST "http://localhost:8080/api/v1/documents/search" \
-H "Authorization: Bearer YOUR_JWT_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"query": "machine learning",
"filters": {
"category": "tech",
"status": "published"
},
"sort": [
{"created_at": "desc"}
],
"limit": 20,
"offset": 0
}'
Request Body Parameters:
query (optional): Full-text search query across content, title, and metadatafilters (optional): Field filters for exact matching (supports single values or arrays)sort (optional): Sort specification with field and direction (asc/desc)limit (optional): Maximum results per page (1-100, default: 10)offset (optional): Number of results to skip (default: 0)Response (200 OK):
{
"documents": [
{
"id": "doc-1",
"content": "Machine learning content...",
"title": "ML Guide",
"metadata": {"category": "tech"},
"created_at": "2026-05-01T10:00:00Z",
"updated_at": "2026-05-15T14:30:00Z"
}
],
"total": 42,
"limit": 20,
"offset": 0,
"has_more": true
}
Error Responses:
401 Unauthorized: Missing or invalid JWT token422 Unprocessable Entity: Invalid request parameters503 Service Unavailable: OpenSearch service unavailableThe API implements a fail-closed security model:
allowed_users field: Document is NOT accessibleallowed_users array: Document is NOT accessibleallowed_users contains username: Document is accessibleThe API returns 404 Not Found for both:
This prevents attackers from enumerating document IDs by observing different error codes.
ACL filtering is applied at the OpenSearch query level, not in application code. This ensures:
Add to .env file:
# OpenSearch Connection
OPENSEARCH_HOST=localhost
OPENSEARCH_PORT=9200
OPENSEARCH_USE_SSL=false
OPENSEARCH_VERIFY_CERTS=false
OPENSEARCH_USERNAME=admin
OPENSEARCH_PASSWORD=admin
# Document Retrieval Configuration
OPENSEARCH_DEFAULT_INDEX=documents
OPENSEARCH_TIMEOUT=30
OPENSEARCH_MAX_RETRIES=3
Documents must have an allowed_users field:
{
"_id": "doc-123",
"_source": {
"content": "Document content",
"title": "Document title",
"metadata": {},
"allowed_users": ["john.doe", "jane.smith"],
"created_at": "2026-05-01T10:00:00Z",
"updated_at": "2026-05-15T14:30:00Z"
}
}
Index Mapping (recommended):
{
"mappings": {
"properties": {
"content": {"type": "text"},
"title": {"type": "text"},
"metadata": {"type": "object"},
"allowed_users": {"type": "keyword"},
"created_at": {"type": "date"},
"updated_at": {"type": "date"}
}
}
}
curl -X POST "http://localhost:8080/auth/login" \
-H "Content-Type: application/json" \
-d '{
"username": "john.doe",
"password": "${USER_PASSWORD}"
}'
Response:
{
"access_token": "ey...",
"token_type": "bearer"
}
curl -X GET "http://localhost:8080/auth/me" \
-H "Authorization: Bearer YOUR_JWT_TOKEN"
Response:
{
"username": "john.doe",
"email": "john.doe@example.com",
"full_name": "John Doe"
}
Use the JWT token in the Authorization header for all document requests.
import requests
# Authenticate
auth_response = requests.post(
"http://localhost:8080/auth/login",
json={"username": "john.doe", "password": "<YOUR_PASSWORD>"}
)
token = auth_response.json()["access_token"]
# Retrieve document
headers = {"Authorization": f"Bearer {token}"}
response = requests.get(
"http://localhost:8080/api/v1/documents/doc-123",
headers=headers
)
if response.status_code == 200:
document = response.json()
print(f"Title: {document['title']}")
print(f"Content: {document['content']}")
elif response.status_code == 404:
print("Document not found or not authorized")
import requests
# Authenticate (same as above)
token = "YOUR_JWT_TOKEN"
headers = {"Authorization": f"Bearer {token}"}
# Search documents
search_request = {
"query": "artificial intelligence",
"filters": {
"category": "tech",
"status": "published"
},
"sort": [{"created_at": "desc"}],
"limit": 20,
"offset": 0
}
response = requests.post(
"http://localhost:8080/api/v1/documents/search",
headers=headers,
json=search_request
)
results = response.json()
print(f"Found {results['total']} documents")
for doc in results['documents']:
print(f"- {doc['title']}")
import requests
token = "YOUR_JWT_TOKEN"
headers = {"Authorization": f"Bearer {token}"}
# Fetch all results with pagination
all_documents = []
offset = 0
limit = 100
while True:
response = requests.post(
"http://localhost:8080/api/v1/documents/search",
headers=headers,
json={"limit": limit, "offset": offset}
)
results = response.json()
all_documents.extend(results['documents'])
if not results['has_more']:
break
offset += limit
print(f"Retrieved {len(all_documents)} total documents")
Located in tests/unit/api/services/test_acl_query_builder.py:
# Run unit tests
pytest tests/unit/api/services/test_acl_query_builder.py -v
Located in tests/integration/api/test_documents_api.py:
# Run integration tests (requires OpenSearch)
pytest tests/integration/api/test_documents_api.py -v
# Run with coverage
pytest tests/unit/api/services/ tests/integration/api/test_documents_api.py \
--cov=src/docpipe/api/services \
--cov=src/docpipe/api/routes/documents \
--cov-report=html
allowed_users field: Create an index on the allowed_users field for fast filteringThe OpenSearchService maintains a connection pool for efficient resource usage:
Monitor these metrics:
Issue: 503 Service Unavailable
Issue: 401 Unauthorized
Issue: 404 Not Found for existing documents
allowed_users fieldIssue: Empty search results
Enable debug logging:
export DS_LOG_LEVEL=DEBUG
This will log:
Planned features for future releases:
allowed_groups in addition to allowed_usersallowed_users fieldFor issues or questions: