Metadata in the MCP Context Server provides a powerful way to enrich, organize, and query your context entries. This guide covers the complete lifecycle of metadata: adding it when storing context, updating it as your workflow evolves, and filtering by it when searching.
Key Capabilities:
- Flexible Structure: Store any JSON-serializable data (strings, numbers, booleans, arrays, nested objects)
- 16 Operators: From simple equality to advanced pattern matching and range queries
- Performance Optimized: Strategic indexing for common fields (status, agent_name, task_name, project, report_type)
- Configurable Indexing: Customize indexed fields via
METADATA_INDEXED_FIELDSenvironment variable - Case Control: Case-sensitive and case-insensitive string operations
- Partial Updates: RFC 7396 JSON Merge Patch for selective metadata modifications
- Query Statistics: Execution time and query plan analysis
Available in Multiple Tools:
store_context: Add metadata when creating context entriesupdate_context: Modify metadata on existing entries (full replacement or partial patch)search_context: Keyword and filter-based search with metadata filteringsemantic_search_context: Vector similarity search with metadata filtering (requires semantic search enabled)fts_search_context: Full-text search with metadata filtering (requires FTS enabled)hybrid_search_context: Combined FTS + semantic search with metadata filtering (requires hybrid search enabled)
All search tools support identical metadata filtering syntax and return consistent error responses.
The metadata field accepts any JSON-serializable structure with no predefined schema:
{
"status": "active",
"priority": 5,
"assignee": "alice@example.com",
"due_date": "2025-10-15",
"tags": ["urgent", "backend"],
"config": {
"retries": 3,
"timeout": 30
}
}Supported Types:
- Strings:
"active","completed" - Numbers:
42,3.14,0 - Booleans:
true,false - Null:
null - Lists:
["value1", "value2"] - Nested objects:
{"user": {"id": 123}}
The following metadata fields are indexed by default for faster filtering:
| Field | Type Hint | SQLite | PostgreSQL | Use Case |
|---|---|---|---|---|
status |
string | B-tree | B-tree | State tracking ("pending", "active", "done") |
agent_name |
string | B-tree | B-tree | Agent identification ("planner", "developer") |
task_name |
string | B-tree | B-tree | Task identification ("auth-impl", "data-export") |
project |
string | B-tree | B-tree | Project filtering ("my-app", "backend") |
report_type |
string | B-tree | B-tree | Report categorization ("research", "implementation") |
references |
object | Not indexed | GIN | Cross-references (uses containment queries) |
technologies |
array | Not indexed | GIN | Technology stack (uses containment queries) |
Backend Differences:
- SQLite: Only scalar fields (string, integer, boolean, float) can be indexed using expression indexes. Array and object fields cannot be efficiently indexed.
- PostgreSQL: Scalar fields use B-tree expression indexes. Array and object fields use GIN index for efficient containment queries (
@>operator).
Configurable Indexing:
You can customize which metadata fields are indexed via environment variables:
METADATA_INDEXED_FIELDS: Comma-separated list of fields with optional type hints (e.g.,status,priority:integer,tags:array)METADATA_INDEX_SYNC_MODE: How to handle index mismatches at startup (strict,auto,warn,additive)
Default: status,agent_name,task_name,project,report_type,references:object,technologies:array
See Environment Variables section for details.
Performance Note: Using indexed fields in filters significantly improves query performance, especially with large datasets.
When storing context entries, you can attach metadata to provide structured information about the entry. This metadata can later be used for filtering and organization.
# Store context with metadata
store_context(
thread_id="project-alpha",
source="agent",
text="Implement user authentication",
metadata={
"status": "active",
"priority": 8,
"assignee": "alice@example.com"
}
)For optimal query performance, use the indexed fields when possible:
# Store task with indexed metadata fields
store_context(
thread_id="project-alpha",
source="agent",
text="Implement user authentication",
metadata={
"status": "active", # Indexed - fast filtering
"agent_name": "developer", # Indexed - agent identification
"task_name": "auth-impl", # Indexed - task tracking
"project": "backend-api", # Indexed - project filtering
"report_type": "implementation", # Indexed - report categorization
"technologies": ["python", "fastapi"], # GIN indexed in PostgreSQL
"references": {"context_ids": ["0190abcdef1234567890abcdef100abc", "0190abcdef1234567890abcdef101abc"]}, # GIN indexed in PostgreSQL
"assignee": "alice@example.com", # Not indexed but still queryable
"due_date": "2025-10-20" # Not indexed but still queryable
}
)Metadata can include nested objects for more complex data:
# Store context with nested metadata
store_context(
thread_id="config-test",
source="agent",
text="Configuration loaded",
metadata={
"user": {
"id": 123,
"preferences": {
"theme": "dark",
"notifications": True
}
},
"settings": {
"timeout": 30,
"retries": 3
}
}
)Task Management:
metadata={
"status": "active",
"priority": 8,
"assignee": "alice@example.com",
"due_date": "2025-10-20",
"task_name": "auth-implementation",
"completed": False,
"tags": ["backend", "security"]
}Agent Coordination:
metadata={
"agent_name": "data-processor",
"task_name": "batch-processing",
"execution_time": 45.3,
"resource_usage": {
"cpu_percent": 65.2,
"memory_mb": 512
},
"records_processed": 1000,
"status": "completed"
}Knowledge Base:
metadata={
"category": "machine-learning",
"subcategory": "transformers",
"relevance_score": 8.5,
"source_url": "https://arxiv.org/...",
"author": "research-agent",
"year": 2025,
"peer_reviewed": True
}Debugging Context:
metadata={
"error_type": "TimeoutError",
"error_code": "DB_TIMEOUT",
"stack_trace": "File payment.py, line 42...",
"environment": "production",
"version": "v2.3.1",
"severity": "critical",
"timestamp": "2025-10-10T08:30:00Z"
}Analytics and Events:
metadata={
"user_id": "user_12345",
"session_id": "sess_abc123",
"event_type": "checkout_completed",
"timestamp": "2025-10-10T10:15:30Z",
"revenue": 149.99,
"items_count": 3,
"platform": "web"
}The update_context tool provides two ways to modify metadata on existing context entries:
- Full Replacement (
metadataparameter): Replace the entire metadata object - Partial Update (
metadata_patchparameter): Modify specific fields while preserving others
| Scenario | Use | Parameter |
|---|---|---|
| Replace entire metadata object | Full replacement | metadata |
| Update a single field | Partial update | metadata_patch |
| Add new fields while preserving existing | Partial update | metadata_patch |
| Delete specific fields | Partial update | metadata_patch |
| Clear all metadata | Full replacement | metadata={} |
Mutual Exclusivity: You cannot use both metadata and metadata_patch in the same call.
Use the metadata parameter to completely replace all metadata:
# Replace all metadata
update_context(
context_id="0190abcdef1234567890abcdef123456",
metadata={
"status": "completed",
"priority": 10,
"reviewer": "bob"
}
)
# Result: Only these three fields exist (previous metadata is gone)The metadata_patch parameter implements RFC 7396 JSON Merge Patch semantics:
- New keys in patch are ADDED to existing metadata
- Existing keys are REPLACED with new values from patch
- Null values DELETE keys from metadata
# Original metadata: {"status": "pending", "priority": 5, "assignee": "alice"}
# Add a new field
update_context(context_id="0190abcdef1234567890abcdef123456", metadata_patch={"category": "backend"})
# Result: {"status": "pending", "priority": 5, "assignee": "alice", "category": "backend"}
# Update existing field
update_context(context_id="0190abcdef1234567890abcdef123456", metadata_patch={"status": "completed"})
# Result: {"status": "completed", "priority": 5, "assignee": "alice", "category": "backend"}
# Delete a field using null
update_context(context_id="0190abcdef1234567890abcdef123456", metadata_patch={"assignee": None})
# Result: {"status": "completed", "priority": 5, "category": "backend"}
# Multiple operations in one call
update_context(
context_id="0190abcdef1234567890abcdef123456",
metadata_patch={
"status": "archived", # Update existing
"archived_at": "2025-10", # Add new
"category": None # Delete
}
)
# Result: {"status": "archived", "priority": 5, "archived_at": "2025-10"}The metadata_patch can be combined with other update fields:
# Update text and patch metadata in one operation
update_context(
context_id="0190abcdef1234567890abcdef123456",
text="Updated analysis results",
metadata_patch={"status": "reviewed", "reviewer": "bob"}
)
# Update tags and patch metadata
update_context(
context_id="0190abcdef1234567890abcdef123456",
tags=["completed", "verified"],
metadata_patch={"completed": True}
)1. Cannot Set Null Values
Using null in the patch always DELETES the key. If you need to store a null value, use full metadata replacement:
# This DELETES the field, not sets it to null
update_context(context_id="0190abcdef1234567890abcdef123456", metadata_patch={"optional_field": None})
# Result: field is removed
# To store null, use full replacement
current = get_context_by_ids(context_ids=["0190abcdef1234567890abcdef123456"])
new_metadata = current[0]["metadata"]
new_metadata["optional_field"] = None
update_context(context_id="0190abcdef1234567890abcdef123456", metadata=new_metadata)2. Arrays are Replaced Entirely
Array operations are replace-only - no element-wise add/remove:
# Original: {"tags": ["a", "b", "c"]}
# This replaces the entire array
update_context(context_id="0190abcdef1234567890abcdef123456", metadata_patch={"tags": ["x", "y"]})
# Result: {"tags": ["x", "y"]} (not ["a", "b", "c", "x", "y"])
# To append, read current array first
current = get_context_by_ids(context_ids=["0190abcdef1234567890abcdef123456"])
tags = current[0]["metadata"]["tags"]
tags.append("new_tag")
update_context(context_id="0190abcdef1234567890abcdef123456", metadata_patch={"tags": tags})Metadata filtering enables powerful, flexible querying of context entries using structured JSON data. Unlike simple tag-based organization, metadata filtering supports complex queries with 16 operators, nested JSON paths, and performance-optimized indexes for common fields.
All search tools (search_context, semantic_search_context, fts_search_context, and hybrid_search_context) support identical metadata filtering syntax.
Use the metadata parameter for straightforward key-value equality matching:
# Find all active contexts
search_context(
thread_id="project-123",
metadata={"status": "active"}
)
# Multiple conditions (AND logic)
search_context(
thread_id="project-123",
metadata={
"status": "active",
"priority": 5,
"completed": False
}
)Characteristics:
- Case-insensitive string matching by default
- All conditions must match (AND logic)
- Simple syntax for common queries
- Limited to equality comparisons
Use the metadata_filters parameter for complex queries with operators:
# Priority greater than 5
search_context(
thread_id="project-123",
metadata_filters=[
{"key": "priority", "operator": "gt", "value": 5}
]
)
# Multiple advanced filters
search_context(
thread_id="project-123",
metadata_filters=[
{"key": "priority", "operator": "gte", "value": 3},
{"key": "status", "operator": "in", "value": ["active", "pending"]},
{"key": "agent_name", "operator": "starts_with", "value": "executor"}
]
)Characteristics:
- 16 powerful operators
- Numeric comparisons and range queries
- Pattern matching for strings
- Field existence checks
- Supports nested JSON paths
The same metadata filtering syntax works with semantic_search_context:
# Semantic search with metadata filtering
semantic_search_context(
query="authentication implementation",
metadata={"status": "completed"}, # Simple filter
metadata_filters=[ # Advanced filters
{"key": "priority", "operator": "gte", "value": 5},
{"key": "agent_name", "operator": "exists"}
]
)
# Find similar contexts from a specific agent with high priority
semantic_search_context(
query="database optimization",
thread_id="project-123",
metadata_filters=[
{"key": "agent_name", "operator": "eq", "value": "research-agent"},
{"key": "priority", "operator": "gt", "value": 7}
]
)This combines the power of semantic similarity search with precise metadata filtering, enabling queries like "find similar content about authentication from completed, high-priority tasks."
The same metadata filtering syntax works with hybrid_search_context:
# Hybrid search with metadata filtering
hybrid_search_context(
query="authentication implementation",
metadata={"status": "completed"}, # Simple filter
metadata_filters=[ # Advanced filters
{"key": "priority", "operator": "gte", "value": 5},
{"key": "agent_name", "operator": "exists"}
]
)
# Find high-confidence matches from a specific project
hybrid_search_context(
query="database optimization",
thread_id="project-123",
metadata_filters=[
{"key": "status", "operator": "in", "value": ["completed", "reviewed"]},
{"key": "priority", "operator": "gt", "value": 7}
]
)Hybrid search returns results with combined RRF scores from both FTS and semantic search, plus individual rankings from each method.
Simple and advanced filters can be combined in a single query:
search_context(
thread_id="project-123",
source="agent",
metadata={"status": "active"}, # Simple filter
metadata_filters=[ # Advanced filters
{"key": "priority", "operator": "gt", "value": 5},
{"key": "agent_name", "operator": "exists"}
]
)Match exact values. Case-insensitive for strings by default.
# Find completed tasks
{"key": "status", "operator": "eq", "value": "completed"}
# Case-sensitive matching
{"key": "status", "operator": "eq", "value": "Completed", "case_sensitive": True}
# Numeric equality
{"key": "priority", "operator": "eq", "value": 5}
# Boolean equality
{"key": "completed", "operator": "eq", "value": True}Match all values except the specified one.
# Exclude archived entries
{"key": "status", "operator": "ne", "value": "archived"}
# Not a specific priority
{"key": "priority", "operator": "ne", "value": 0}The comparison operators (gt, gte, lt, lte) -- and numeric eq/ne with a numeric value -- match JSON-number-typed values only, identically on SQLite and PostgreSQL. A stored value that is not a JSON number for the filtered key (a string such as "12abc", a boolean, JSON null, or an absent key) never matches a numeric operator on either backend; the query is not aborted and the value is not coerced. So {"key": "score", "operator": "gt", "value": 5} returns only entries whose score is a number greater than 5, never an entry whose score is the string "12abc". Float values are compared at IEEE double precision on both backends, so {"key": "score", "operator": "eq", "value": 0.3} matches a stored 0.3 (and gt 0.3 excludes a stored 0.3) identically on SQLite and PostgreSQL.
The same type-restriction applies to booleans: a boolean value in eq/ne, as an in/not_in member, or in array_contains matches JSON-boolean-typed values only on both backends. A stored numeric 0/1 or a string "true"/"false" never matches a boolean filter (and a boolean array_contains matches only a JSON-boolean array element, never a numeric 1/0); only a real JSON true/false does. Use {"key": "flag", "operator": "eq", "value": true} to match entries whose flag is the JSON boolean true.
The string operators complete the same contract: eq/ne with a string value, the ordered comparisons (gt/gte/lt/lte) with a string value, contains/starts_with/ends_with, and string members of in/not_in match JSON-string-typed values only on both backends. A stored JSON number is never compared as text, because SQLite renders a JSON number through a 64-bit integer / IEEE double (losing the exact text for an out-of-int64 or high-precision value) while PostgreSQL ->> returns the exact original number text -- a text comparison would otherwise diverge across backends. To match a stored number use a numeric value ({"key": "count", "operator": "eq", "value": 5}), not its string form ("5"); to match a stored string use a string value. Numeric members of in/not_in still match stored numbers numerically ({"key": "priority", "operator": "in", "value": [1, 2, 3]} matches a stored number 2), so only the text-vs-number cross-type comparison is removed.
Numeric comparison, exclusive.
# High priority tasks (priority > 5)
{"key": "priority", "operator": "gt", "value": 5}
# Recent scores
{"key": "score", "operator": "gt", "value": 80.5}Numeric comparison, inclusive.
# Priority 5 or higher
{"key": "priority", "operator": "gte", "value": 5}Numeric comparison, exclusive.
# Low priority tasks (priority < 3)
{"key": "priority", "operator": "lt", "value": 3}Numeric comparison, inclusive.
# Priority 3 or lower
{"key": "priority", "operator": "lte", "value": 3}Check if value matches any item in the provided list. Supports both string and integer arrays. in and not_in value lists accept at most 100 members; a longer list is rejected with a validation error before any query runs.
# Active or pending tasks
{"key": "status", "operator": "in", "value": ["active", "pending", "review"]}
# Specific priority levels (integer arrays)
{"key": "priority", "operator": "in", "value": [1, 5, 10]}
# Case-insensitive by default
{"key": "environment", "operator": "in", "value": ["dev", "staging", "prod"]}Exclude values matching any item in the list.
# Exclude archived and deleted
{"key": "status", "operator": "not_in", "value": ["archived", "deleted"]}Check if a metadata field is present (not null, not missing).
# Contexts with agent assignment
{"key": "agent_name", "operator": "exists"}
# Has error information
{"key": "error_message", "operator": "exists"}Check if a metadata field is missing or null.
# Contexts without assignee
{"key": "assignee", "operator": "not_exists"}
# No completion date set
{"key": "completed_at", "operator": "not_exists"}Case-insensitive substring matching by default.
# Description contains "error"
{"key": "description", "operator": "contains", "value": "error"}
# Case-sensitive search
{"key": "log_message", "operator": "contains", "value": "ERROR", "case_sensitive": True}Match strings beginning with specific prefix.
# Agent names starting with "executor"
{"key": "agent_name", "operator": "starts_with", "value": "executor"}
# Task names starting with "test_"
{"key": "task_name", "operator": "starts_with", "value": "test_"}Match strings ending with specific suffix.
# Files ending with .json
{"key": "filename", "operator": "ends_with", "value": ".json"}
# Email addresses ending with company domain
{"key": "email", "operator": "ends_with", "value": "@example.com"}Check if field contains JSON null value (different from missing field).
# Explicitly set to null
{"key": "deleted_at", "operator": "is_null"}Check if field has a non-null value.
# Has a value (not null)
{"key": "assigned_to", "operator": "is_not_null"}Note: exists checks for field presence, is_not_null checks for non-null values. Use exists for missing fields, is_not_null for null values.
Check if a JSON array field contains a specific element value.
# Find entries where technologies array contains "python"
{"key": "technologies", "operator": "array_contains", "value": "python"}
# Find entries where priority_levels array contains 5
{"key": "priority_levels", "operator": "array_contains", "value": 5}
# Case-insensitive string matching
{"key": "technologies", "operator": "array_contains", "value": "PYTHON", "case_sensitive": False}
# Nested array path
{"key": "references.context_ids", "operator": "array_contains", "value": "0190abcdef1234567890abcdef200abc"}Value Requirements:
- Must be a single scalar value (string, number, or boolean)
- Cannot be an array or null
Use Cases:
- Filter by technology stack:
{"key": "technologies", "operator": "array_contains", "value": "python"} - Filter by tag in array:
{"key": "tags", "operator": "array_contains", "value": "urgent"} - Filter by numeric value in array:
{"key": "priority_levels", "operator": "array_contains", "value": 5} - Filter nested arrays:
{"key": "references.context_ids", "operator": "array_contains", "value": "0190abcdef1234567890abcdef200abc"}
Example:
# metadata = {"technologies": ["python", "fastapi", "postgresql"]}
# Find entries where technologies contains "python"
search_context(
metadata_filters=[
{"key": "technologies", "operator": "array_contains", "value": "python"}
]
)
# Case-insensitive search
search_context(
metadata_filters=[
{"key": "technologies", "operator": "array_contains", "value": "PYTHON", "case_sensitive": False}
]
)
# Nested path example
# metadata = {"references": {"context_ids": ["0190abcdef1234567890abcdef100abc", "0190abcdef1234567890abcdef200abc", "0190abcdef1234567890abcdef300abc"]}}
search_context(
metadata_filters=[
{"key": "references.context_ids", "operator": "array_contains", "value": "0190abcdef1234567890abcdef200abc"}
]
)Notes:
- Returns empty results (not error) if the field is not an array or doesn't exist
- Supports case-insensitive matching for string values (set
case_sensitive: false) - Works with nested paths using dot notation
- Type-aware (cross-backend parity): a string member matches only JSON-string array elements, a numeric member only JSON-number elements, and a boolean member only JSON-boolean elements. A string member never matches a numeric or boolean element (and vice versa), so results are identical on SQLite and PostgreSQL. This mirrors the type-aware contract of the scalar operators and avoids divergence from SQLite's numeric-to-text rendering of out-of-int64 / high-precision numbers.
Track task states, priorities, assignments, and deadlines.
# Store task with metadata
store_context(
thread_id="project-alpha",
source="agent",
text="Implement user authentication",
metadata={
"status": "active",
"priority": 8,
"assignee": "alice@example.com",
"due_date": "2025-10-20",
"task_name": "auth-implementation",
"completed": False,
"tags": ["backend", "security"]
}
)
# Find high-priority incomplete tasks
search_context(
thread_id="project-alpha",
metadata_filters=[
{"key": "priority", "operator": "gte", "value": 7},
{"key": "completed", "operator": "eq", "value": False},
{"key": "status", "operator": "in", "value": ["active", "pending"]}
]
)
# Find overdue tasks (assignee exists but not completed)
search_context(
thread_id="project-alpha",
metadata_filters=[
{"key": "assignee", "operator": "exists"},
{"key": "completed", "operator": "eq", "value": False}
]
)
# Find Alice's tasks
search_context(
thread_id="project-alpha",
metadata={"assignee": "alice@example.com"}
)Track agent activities, resource usage, and execution metrics.
# Store agent execution context
store_context(
thread_id="data-pipeline",
source="agent",
text="Processed 1000 records successfully",
metadata={
"agent_name": "data-processor",
"task_name": "batch-processing",
"execution_time": 45.3,
"resource_usage": {
"cpu_percent": 65.2,
"memory_mb": 512
},
"records_processed": 1000,
"status": "completed"
}
)
# Find slow executions (> 60 seconds)
search_context(
thread_id="data-pipeline",
metadata_filters=[
{"key": "execution_time", "operator": "gt", "value": 60}
]
)
# Find specific agent activities
search_context(
thread_id="data-pipeline",
metadata_filters=[
{"key": "agent_name", "operator": "starts_with", "value": "data-"}
]
)
# Find failed or incomplete tasks
search_context(
thread_id="data-pipeline",
metadata_filters=[
{"key": "status", "operator": "in", "value": ["failed", "error", "timeout"]}
]
)Categorize and retrieve information with relevance scoring.
# Store knowledge base entry
store_context(
thread_id="ml-research",
source="agent",
text="GPT-4 architecture uses sparse attention mechanism...",
metadata={
"category": "machine-learning",
"subcategory": "transformers",
"relevance_score": 8.5,
"source_url": "https://arxiv.org/...",
"author": "research-agent",
"year": 2025,
"peer_reviewed": True
}
)
# Find highly relevant ML papers
search_context(
thread_id="ml-research",
metadata_filters=[
{"key": "category", "operator": "eq", "value": "machine-learning"},
{"key": "relevance_score", "operator": "gte", "value": 7.0},
{"key": "peer_reviewed", "operator": "eq", "value": True}
]
)
# Find recent transformer research
search_context(
thread_id="ml-research",
metadata={
"subcategory": "transformers",
"year": 2025
}
)
# Find papers from specific source
search_context(
thread_id="ml-research",
metadata_filters=[
{"key": "source_url", "operator": "contains", "value": "arxiv.org"}
]
)Save error information, stack traces, environment details.
# Store error context
store_context(
thread_id="production-debug",
source="agent",
text="Database connection timeout in payment processor",
metadata={
"error_type": "TimeoutError",
"error_code": "DB_TIMEOUT",
"stack_trace": "File payment.py, line 42...",
"environment": "production",
"version": "v2.3.1",
"severity": "critical",
"timestamp": "2025-10-10T08:30:00Z"
}
)
# Find critical production errors
search_context(
thread_id="production-debug",
metadata_filters=[
{"key": "environment", "operator": "eq", "value": "production"},
{"key": "severity", "operator": "in", "value": ["critical", "high"]}
]
)
# Find timeout-related errors
search_context(
thread_id="production-debug",
metadata_filters=[
{"key": "error_type", "operator": "contains", "value": "timeout", "case_sensitive": False}
]
)
# Find errors in specific version
search_context(
thread_id="production-debug",
metadata_filters=[
{"key": "version", "operator": "starts_with", "value": "v2.3"}
]
)Record user actions, sessions, and event metadata.
# Store analytics event
store_context(
thread_id="user-analytics",
source="agent",
text="User completed checkout process",
metadata={
"user_id": "user_12345",
"session_id": "sess_abc123",
"event_type": "checkout_completed",
"timestamp": "2025-10-10T10:15:30Z",
"revenue": 149.99,
"items_count": 3,
"platform": "web"
}
)
# Find high-value purchases
search_context(
thread_id="user-analytics",
metadata_filters=[
{"key": "event_type", "operator": "eq", "value": "checkout_completed"},
{"key": "revenue", "operator": "gt", "value": 100}
]
)
# Find specific user activity
search_context(
thread_id="user-analytics",
metadata={"user_id": "user_12345"}
)
# Find mobile platform events
search_context(
thread_id="user-analytics",
metadata_filters=[
{"key": "platform", "operator": "in", "value": ["ios", "android"]}
]
)Use dot notation to query nested metadata structures:
# Store context with nested metadata
store_context(
thread_id="config-test",
source="agent",
text="Configuration loaded",
metadata={
"user": {
"id": 123,
"preferences": {
"theme": "dark",
"notifications": True
}
},
"settings": {
"timeout": 30,
"retries": 3
}
}
)
# Query nested fields with dot notation
search_context(
thread_id="config-test",
metadata_filters=[
{"key": "user.preferences.theme", "operator": "eq", "value": "dark"},
{"key": "settings.timeout", "operator": "gt", "value": 20}
]
)Supported Paths:
- Single level:
"status" - Multi-level:
"user.preferences.theme" - Deep nesting:
"config.database.connection.pool.size"
Path Restrictions:
- Only alphanumeric characters, dots, underscores, and hyphens allowed
- No array indexing (e.g.,
items[0].namenot supported) - No special characters or spaces
Use explain_query=True to analyze query performance:
result = search_context(
thread_id="project-123",
metadata={"status": "active"},
metadata_filters=[
{"key": "priority", "operator": "gt", "value": 5}
],
explain_query=True
)
# Returns execution statistics
print(result["stats"])
# {
# "execution_time_ms": 12.34,
# "filters_applied": 3,
# "rows_returned": 15,
# "backend": "sqlite",
# "query_plan": "..."
# }Statistics Provided:
execution_time_ms: Query execution time in millisecondsfilters_applied: Total number of filter conditions the query applied (thread, source, content type, date bounds and the tag subquery, plus every metadata condition)rows_returned: Number of matching entriesbackend: Active storage backend ("sqlite"or"postgresql")query_plan: Backend query execution plan
Every key above is present whenever the stats block is (that is, whenever explain_query=True). A request rejected by parameter validation -- an invalid metadata_filters operator, for example -- returns the same keys with the counters zeroed and query_plan set to null, because no query ran.
Default indexed fields (status, agent_name, task_name, project, report_type) filter significantly faster:
# Fast - uses indexed field
{"key": "status", "operator": "eq", "value": "active"}
# Fast - uses indexed field
{"key": "project", "operator": "eq", "value": "my-app"}
# Slower - non-indexed field
{"key": "custom_field", "operator": "eq", "value": "value"}For array/object fields (technologies, references), PostgreSQL provides efficient GIN-indexed containment queries:
# Fast on PostgreSQL (GIN index) - find entries with specific technology
{"key": "technologies", "operator": "array_contains", "value": "python"}
# Slower on SQLite - requires full scan (no GIN support)
{"key": "technologies", "operator": "array_contains", "value": "python"}Always specify thread_id to reduce the search space:
# Good - scoped to thread
search_context(
thread_id="project-123",
metadata={"status": "active"}
)
# Slower - searches all threads
search_context(
metadata={"status": "active"}
)Place more restrictive filters first for better performance:
# Better - specific agent name filters more
metadata_filters=[
{"key": "agent_name", "operator": "eq", "value": "specific-agent"},
{"key": "status", "operator": "in", "value": ["active", "pending", "review"]}
]
# Less optimal - broad status filter first
metadata_filters=[
{"key": "status", "operator": "in", "value": ["active", "pending", "review"]},
{"key": "agent_name", "operator": "eq", "value": "specific-agent"}
]Pattern matching operations (contains, starts_with, ends_with) are slower than equality:
# Faster
{"key": "status", "operator": "eq", "value": "active"}
# Slower
{"key": "description", "operator": "contains", "value": "active"}Based on test suite results with typical workloads:
| Query Type | Execution Time | Notes |
|---|---|---|
| Single indexed field | < 10ms | Uses database index |
| Multiple indexed fields | < 20ms | AND conditions, indexed |
| Non-indexed field | < 50ms | Full table scan on filtered results |
| Complex pattern match | < 100ms | String operations |
| Nested JSON path | < 50ms | Uses json_extract() |
Acceptable Scale: Up to 100,000 context entries with acceptable performance using indexed fields.
String operations are case-insensitive by default:
# Matches "Active", "active", "ACTIVE"
{"key": "status", "operator": "eq", "value": "active"}Case-insensitive matching folds ASCII letters (A-Z) only -- this is identical on SQLite and PostgreSQL so the two backends always return the same result set. Non-ASCII letters are NOT case-folded: "café" and "CAFÉ" do NOT match case-insensitively (the É/é differ), though "cafe" and "CAFE" do. (SQLite's built-in LOWER() is ASCII-only and cannot fold Unicode without an ICU extension, so ASCII-only folding is the portable cross-backend contract; use case_sensitive: true when you need exact, byte-for-byte matching.)
Set case_sensitive: true for exact case matching:
# Only matches "Active" (exact case)
{"key": "status", "operator": "eq", "value": "Active", "case_sensitive": True}Case sensitivity applies to all string operators:
eq,ne- Equality comparisonsin,not_in- List membershipcontains- Substring searchstarts_with- Prefix matchingends_with- Suffix matching
# Case-insensitive contains (default)
{"key": "description", "operator": "contains", "value": "error"}
# Matches: "Error occurred", "error found", "ERROR: timeout"
# Case-sensitive contains
{"key": "description", "operator": "contains", "value": "ERROR", "case_sensitive": True}
# Matches: "ERROR: timeout", "FATAL ERROR"
# No match: "error found", "Error occurred"Both search_context and semantic_search_context return identical error response formats when metadata filter validation fails. This unified error handling ensures consistent client-side processing.
Invalid filters return error responses with validation details:
# Invalid operator
result = search_context(
thread_id="test",
metadata_filters=[
{"key": "status", "operator": "invalid_op", "value": "active"}
]
)
# Returns:
# {
# "results": [],
# "count": 0,
# "error": "Metadata filter validation failed",
# "validation_errors": [
# "Invalid metadata filter {...}: ... 'invalid_op' is not a valid MetadataOperator"
# ]
# }
# When explain_query=True, a zeroed "stats" key is also present.# Invalid - empty list
{"key": "status", "operator": "in", "value": []}
# Error: "Operator in requires a non-empty list"# Invalid - special characters
{"key": "status'; DROP TABLE", "operator": "eq", "value": "x"}
# Error: "Invalid metadata key: ... Only alphanumeric characters, dots, underscores, and hyphens are allowed"# Invalid - string value for IN operator
{"key": "status", "operator": "in", "value": "active"}
# Error: "Operator in requires a list value"All validation errors are collected and returned together:
metadata_filters=[
{"key": "status", "operator": "invalid_op", "value": "active"},
{"key": "priority", "operator": "in", "value": []},
{"key": "type", "operator": "another_invalid", "value": 5}
]
# Returns validation_errors array with all 3 errorsDesign your metadata schema to leverage indexed fields:
# Good - uses default indexed fields
metadata={
"status": "active", # Indexed (B-tree)
"agent_name": "processor", # Indexed (B-tree)
"task_name": "data-export", # Indexed (B-tree)
"project": "backend-api", # Indexed (B-tree)
"report_type": "implementation", # Indexed (B-tree)
"technologies": ["python"], # GIN indexed in PostgreSQL
}
# Avoid - custom fields for frequently queried data
metadata={
"my_custom_status": "active", # Not indexed
"my_priority_level": 5 # Not indexed (add via METADATA_INDEXED_FIELDS if needed)
}Tip: If you need to index custom fields (like priority or completed), configure METADATA_INDEXED_FIELDS environment variable.
- Simple filtering for exact matches
- Advanced filtering for complex queries
- Combine both for mixed requirements
# Simple - exact match queries
search_context(metadata={"status": "active", "priority": 5})
# Advanced - range and pattern matching
search_context(metadata_filters=[
{"key": "priority", "operator": "gte", "value": 5},
{"key": "agent_name", "operator": "starts_with", "value": "exec"}
])
# Combined - best of both
search_context(
metadata={"status": "active"},
metadata_filters=[{"key": "priority", "operator": "gt", "value": 5}]
)Maintain consistency across agents for better filtering:
# Good - consistent schema
metadata={
"status": "active", # Always lowercase
"priority": 5, # Always number 1-10
"agent_name": "planner" # Always lowercase, hyphen-separated
}
# Avoid - inconsistent values
metadata={
"status": "Active", # Mixed case
"priority": "high", # String instead of number
"agent_name": "Planner" # Mixed case
}Ensure metadata conforms to your schema:
# Define schema validation
def validate_task_metadata(metadata):
required_fields = ["status", "priority"]
valid_statuses = ["pending", "active", "completed", "failed"]
for field in required_fields:
if field not in metadata:
raise ValueError(f"Missing required field: {field}")
if metadata["status"] not in valid_statuses:
raise ValueError(f"Invalid status: {metadata['status']}")
if not isinstance(metadata["priority"], int) or not (1 <= metadata["priority"] <= 10):
raise ValueError("Priority must be an integer between 1 and 10")
# Use validation before storing
metadata = {"status": "active", "priority": 5}
validate_task_metadata(metadata)
store_context(thread_id="task", source="agent", text="...", metadata=metadata)Profile slow queries to identify optimization opportunities:
result = search_context(
thread_id="large-project",
metadata_filters=[...],
explain_query=True
)
# Check execution time
if result["stats"]["execution_time_ms"] > 100:
print("Slow query detected!")
print(result["stats"]["query_plan"])Use pagination for large result sets:
# Fetch in batches
search_context(
thread_id="large-thread",
metadata={"status": "active"},
limit=50, # Results per page
offset=0 # Start position
)Leverage all filtering capabilities for precise queries:
search_context(
thread_id="project-123", # Thread filter
source="agent", # Source filter
tags=["urgent", "backend"], # Tag filter (OR logic)
content_type="text", # Content type filter
metadata={"status": "active"}, # Simple metadata filter
metadata_filters=[ # Advanced metadata filters
{"key": "priority", "operator": "gte", "value": 5}
],
start_date="2025-11-01", # Date range filter
end_date="2025-11-30"
)Date Filtering:
The start_date and end_date parameters filter entries by creation timestamp using ISO 8601 format:
# Find entries from a specific day
search_context(thread_id="project-123", start_date="2025-11-29", end_date="2025-11-29")
# Find entries from a date range with metadata filters
search_context(
thread_id="project-123",
metadata={"status": "active"},
start_date="2025-11-01",
end_date="2025-11-30"
)# Producer agent stores tasks
async def create_task(task_data):
await store_context(
thread_id="task-queue",
source="agent",
text=task_data["description"],
metadata={
"status": "pending",
"priority": task_data["priority"],
"task_name": task_data["name"],
"created_at": datetime.now().isoformat(),
"completed": False
}
)
# Consumer agent fetches high-priority pending tasks
async def get_next_task():
result = await search_context(
thread_id="task-queue",
metadata_filters=[
{"key": "status", "operator": "eq", "value": "pending"},
{"key": "priority", "operator": "gte", "value": 5}
],
limit=1
)
return result["results"][0] if result["results"] else None
# Mark task as completed (preserves existing metadata like priority, task_name)
async def complete_task(context_id):
await update_context(
context_id=context_id,
metadata_patch={
"status": "completed",
"completed": True,
"completed_at": datetime.now().isoformat()
}
)# Planner agent creates tasks
await store_context(
thread_id="data-pipeline",
source="agent",
text="Extract data from API",
metadata={
"agent_name": "planner",
"task_name": "extract-api-data",
"status": "assigned",
"assigned_to": "extractor-agent",
"priority": 7
}
)
# Extractor agent finds assigned tasks
my_tasks = await search_context(
thread_id="data-pipeline",
metadata={
"assigned_to": "extractor-agent",
"status": "assigned"
}
)
# Monitor agent finds stalled tasks
stalled = await search_context(
thread_id="data-pipeline",
metadata_filters=[
{"key": "status", "operator": "in", "value": ["assigned", "in_progress"]},
{"key": "assigned_to", "operator": "exists"}
]
)Problem: Filter returns empty results when data exists.
Solutions:
-
Check case sensitivity:
# Try case-insensitive (default) {"key": "status", "operator": "eq", "value": "active"} # If that fails, try explicit case-insensitive {"key": "status", "operator": "eq", "value": "active", "case_sensitive": False}
-
Verify metadata field names:
# Use exact field names result = search_context(thread_id="test", include_images=False) print(result["results"][0]["metadata"]) # Check actual field names
-
Check for null vs missing fields:
# Field doesn't exist {"key": "field", "operator": "not_exists"} # Field exists but is null {"key": "field", "operator": "is_null"}
Problem: Queries take longer than expected.
Solutions:
-
Add thread_id filter:
# Slow - searches all threads search_context(metadata={"status": "active"}) # Fast - scoped to specific thread search_context(thread_id="specific-thread", metadata={"status": "active"})
-
Use indexed fields:
# Slow - custom field {"key": "my_status", "operator": "eq", "value": "active"} # Fast - indexed field {"key": "status", "operator": "eq", "value": "active"}
-
Profile with explain_query:
result = search_context(..., explain_query=True) print(result["stats"]["execution_time_ms"]) print(result["stats"]["query_plan"])
Problem: Receiving metadata filter validation errors.
Solutions:
-
Check operator spelling:
# Invalid {"key": "priority", "operator": "greater_than", "value": 5} # Valid {"key": "priority", "operator": "gt", "value": 5}
-
Verify value types:
# Invalid - string for IN operator {"key": "status", "operator": "in", "value": "active"} # Valid - list for IN operator {"key": "status", "operator": "in", "value": ["active"]}
-
Check key format:
# Invalid - special characters {"key": "status@field", "operator": "eq", "value": "active"} # Valid - alphanumeric, dots, underscores, hyphens only {"key": "status_field", "operator": "eq", "value": "active"}
Configure which metadata fields are indexed for faster filtering.
- Type: String (comma-separated list)
- Default:
status,agent_name,task_name,project,report_type,references:object,technologies:array - Format:
field1,field2:type,field3:type
Supported Type Hints:
string(default): Standard string indexinteger: Cast to integer for numeric comparisonsboolean: Cast to booleanfloat: Cast to numeric for decimal comparisonsarray: Array field (PostgreSQL GIN only, skipped in SQLite)object: Nested object field (PostgreSQL GIN only, skipped in SQLite)
Type hints are enforced at the write boundary. A field declared integer, boolean, or float is indexed through a SQL cast on PostgreSQL, so a value that cast cannot accept is REJECTED when it is stored -- on BOTH backends, and before any embedding or summary generation runs. Previously such a value stored fine on SQLite (whose index applies no cast) and aborted the PostgreSQL transaction after a full generation pass, with a raw driver error.
What counts as acceptable mirrors PostgreSQL's own input parsers applied to the text the index stores: integer accepts an optional sign and a decimal run, plus PostgreSQL 16's non-decimal literals (0x10, 0o17, 0b101) and _ digit separators, and must fit the 32-bit INTEGER range; boolean accepts the usual true/false/yes/no/on/off spellings and their unambiguous prefixes; float accepts decimal and exponent forms plus NaN/Infinity. Only ASCII whitespace is trimmed, because that is all PostgreSQL trims. A list or object under a typed field is rejected outright -- there is no scalar for the cast to produce.
Turning an existing field into a typed one is an operator decision that the server verifies at startup: if rows already stored hold a value the new cast rejects, CREATE INDEX cannot succeed, and the server stops with a configuration error naming the field and quoting the offending value rather than crash-looping on a raw driver error.
Examples:
# Default indexed fields
METADATA_INDEXED_FIELDS=status,agent_name,task_name,project,report_type,references:object,technologies:array
# Add priority with integer type hint for range queries
METADATA_INDEXED_FIELDS=status,agent_name,task_name,project,report_type,priority:integer
# Minimal indexing for simple use cases
METADATA_INDEXED_FIELDS=status,agent_nameControl how the server handles index mismatches at startup.
- Type: String (enum)
- Default:
additive - Options:
strict,auto,warn,additive
| Mode | Missing Indexes | Extra Indexes | Behavior |
|---|---|---|---|
strict |
Fail startup | Fail startup | Requires exact match - use for production safety |
auto |
Create | Drop | Full synchronization - may drop indexes |
warn |
Log warning | Log warning | Continue startup with warnings |
additive |
Create | Keep (log info) | Add missing, never drop - safest for upgrades |
Examples:
# Default: Add missing indexes, keep existing ones
METADATA_INDEX_SYNC_MODE=additive
# Production: Fail if indexes don't match configuration
METADATA_INDEX_SYNC_MODE=strict
# Automatic cleanup: Sync indexes exactly to configuration
METADATA_INDEX_SYNC_MODE=auto- API Reference: API Reference - complete tool documentation
- Database Backends: Database Backends Guide - database configuration
- Semantic Search: Semantic Search Guide - vector similarity search
- Full-Text Search: Full-Text Search Guide - linguistic search with stemming
- Hybrid Search: Hybrid Search Guide - combined FTS + semantic search with RRF fusion
- Docker Deployment: Docker Deployment Guide - containerized deployment
- Authentication: Authentication Guide - HTTP transport authentication
- Main Documentation: README.md - overview and quick start
- Contributing: CONTRIBUTING.md - architecture and development guidelines
app/metadata_types.py- MetadataFilter and MetadataOperator definitionsapp/query_builder.py- SQL query construction with security validationapp/repositories/context_repository.py- Database operations for metadata filtering
tests/core/test_metadata_filtering.py- Comprehensive operator tests and integration examplestests/tools/test_metadata_error_handling.py- Error handling and validation tests
For advanced performance optimization:
- Review indexed fields in
app/schemas/sqlite_schema.sql - Check query execution plans with
explain_query=True - Monitor execution time in query statistics
- Optimize metadata schema design for your use case