Initialize and Configure the Vector Store and Ingestion Process
Overview
This guide describes how to set up and configure the Bloomreach Content AI Vector Store and the Ingestion process.
The Vector Store uses an external vector database. For on-premises projects, Redis and Postgres (with the PgVector extension) are supported. You can evaluate both locally using containerized deployments. Refer to the Redis documentation and PgVector container page for setup instructions.
Vector Store and Ingestion features are not available for Bloomreach Cloud implementations at this time.
A managed vector store for Search Agent is planned for Bloomreach Cloud in Q2 2026. Contact your Account Manager for updates.
The Ingestion Process is a background service running in a CMS pod. It listens for workflow events and updates the Vector Store as needed. The process supports two modes:
previewmode: Updates or removes vectors in the store when a content item is saved or deleted. The store indexes the unpublished variant, so it contains preview content.livemode: Updates or removes vectors when a content item is published or taken offline. Only the published variant is indexed, so the store contains only published content.
When you first set up the Vector Store, it is empty. No content is indexed until you configure the ingestion process. For details on starting content indexing, see the section on Ingestion Process Options, including how to use the brxm.ai.ingest.mode and brxm.ai.ingest.types properties.
Info: The Vector Store is extensible. You can integrate custom vector store backends by implementing the
VectorStoreFactorySPI. This allows you to use any VectorStore supported by Spring AI. See the AI Module Extensibility Guide for more information.
During ingestion, the configured embedding model generates embeddings for each document, and the Vector Store is updated with these vectors. Embedding generation incurs cost and latency at the model provider, but does not impact system performance.
Info: All AI usage costs for embedding generation are logged, regardless of model provider. Usage by the background ingestion process is attributed to the user ID
system-ingestion. For audit logging details, see "Conversation token usage" in the Content Assistant developer guide.
Info: Embeddings are vectors with specific dimensions determined by the model. If you change the embedding model to one with different dimensions, existing embeddings in the vector store become incompatible and must be regenerated.
Installation
The Vector Store requires a running external Redis or Postgres instance. The Ingestion process is installed automatically when you install the AI module.
Configure via Properties Files
Configure the Vector Store and Ingestion process for production using properties files. The system checks the following locations in order of precedence (higher in the list overrides lower):
- System properties passed on the command line
- A properties file named
xm-ai-service.propertiesavailable on the classpath - The project's
platform.propertiesfile
Info: For more information on managing properties files and system properties:
- Bloomreach Cloud: Set Environment Configuration Properties
- On-premises: HST-2 Container Configuration
Hint: When configuring via JCR, use the same property names as in properties files, and set them as String properties in JCR.
If no chat or embedding model is registered, the Vector Store and Ingestion process will not initialize.
Multiplicity of Configurations
You can configure multiple model providers and vector stores using properties, but only one configuration can be active at a time.
Global Configuration Options
Set the active Vector Store using the brxm.ai.vectorstore property. Supported values:
RedisPgVector
Info: If you set this property to an empty value, both the Vector Store and Ingestion process are disabled.
Ingestion Process Options
Control the Ingestion process with the following properties:
| Property | Required | Type | Description | Default | Example |
|---|---|---|---|---|---|
brxm.ai.ingest.mode | yes | enum | Ingestion operating mode. In preview mode, unpublished content is indexed on save/rename/copy/move. In live mode, only published content is indexed during publication (including scheduled publication). | preview or live | |
brxm.ai.ingest.types | no | list of doctypes | Comma-separated list of fully qualified document types. Only these types are indexed. If not set, no content is ingested. Removal from the store ignores this filter. | No types allowed | myproject:bannerdocument, hippogallery:exampleAssetSet |
brxm.ai.ingest.include-dirs | no | list of paths | Comma-separated list of absolute paths. Documents outside these paths are skipped during ingestion. Leave empty to allow ingestion from any directory. Removal from the store ignores this filter. | Any path is included | /content/documents/myproject/banners/, /content/documents/myproject/news/ |
brxm.ai.ingest.exclude-dirs | no | list of paths | Comma-separated list of absolute paths. Documents under these paths are skipped during ingestion. Removal from the store ignores this filter. | No path is excluded | /content/documents/myproject/taxonomies/, /content/documents/myproject/private/ |
brxm.ai.ingest.initial-delay | no | integer (seconds) | Delay, in seconds, after system startup before the Ingestion process starts. | 300 | 1000 |
brxm.ai.ingest.interval | no | integer (seconds) | Frequency, in seconds, at which the process runs. | 12 | 60 |
brxm.ai.ingest.batch-size | no | integer | Number of documents processed per batch. Reduce if your Vector Store is overloaded. | 5 | 2 |
brxm.ai.ingest.delay | no | integer (milliseconds) | Back-off time after each batch. Increase if your Vector Store is overloaded. | 1000 | 10000 |
Info: The ingestion process collects events in an in-memory queue and processes them in batches. Each batch is retried up to 5 times, with the delay between attempts starting at
brxm.ai.ingest.delayand increasing exponentially after each failure.
The ingestion filter always excludes content items not under the absolute path /content/. This built-in filter is always applied during ingestion, but not during removal. Similarly, content under /content/attic/ or /content/taxonomies/ is always excluded from ingestion, but not from removal.
Redis Options
Info: Ensure a Redis instance is running and accessible. See the Redis Spring AI documentation for more details.
| Property | Required | Type | Description | Default | Example |
|---|---|---|---|---|---|
brxm.ai.vectorstore.redis.host | yes | url | Hostname of the Redis instance. | myredis, 127.0.0.1 | |
brxm.ai.vectorstore.redis.port | yes | integer | Port number for Redis. | 6379 | |
brxm.ai.vectorstore.redis.index | yes | string | Index name used in Redis to store embeddings. | myindex | |
brxm.ai.vectorstore.redis.prefix | yes | string | Prefix for each entry in the index, used for identification. | my_prefix | |
brxm.ai.vectorstore.redis.user | no | string | Username for Redis if authentication is enabled. | ||
brxm.ai.vectorstore.redis.password | no | string | Password for Redis if authentication is enabled. | ||
brxm.ai.vectorstore.redis.client-name | no | string | Name of the connecting application, used for identification. | myAppName | |
brxm.ai.vectorstore.redis.timeout-millis | no | integer | Connection timeout to Redis, in milliseconds. | Redis default | 5000 |
PgVector Options
Info: Ensure a Postgres instance with the vector extension is running and accessible. See the PgVector Spring AI documentation for details.
| Property | Required | Type | Description | Default | Example |
|---|---|---|---|---|---|
brxm.ai.vectorstore.pgvector.url | yes | url | Connection string for PgVector. | jdbc:postgresql://myhost:5432/mydbname | |
brxm.ai.vectorstore.pgvector.username | yes | string | Username for PgVector. | ||
brxm.ai.vectorstore.pgvector.password | yes | string | Password for PgVector. | ||
brxm.ai.vectorstore.pgvector.dimensions | yes | string | Embedding vector dimensions. Set during initial table creation. Changing dimensions requires recreating the vector_store table. | 1536 | |
brxm.ai.vectorstore.pgvector.index-type | no | string | Nearest neighbor search index type. Options: NONE (exact search), IVFFlat (list-based), HNSW (multilayer graph). | HNSW | |
brxm.ai.vectorstore.pgvector.distance-type | no | string | Search distance type. Default is COSINE_DISTANCE. Use EUCLIDEAN_DISTANCE or NEGATIVE_INNER_PRODUCT for normalized vectors. | COSINE_DISTANCE | |
brxm.ai.vectorstore.pgvector.remove-existing-vector-store-table | no | boolean | If true, deletes the existing vector_store table on startup. | false | true |
brxm.ai.vectorstore.pgvector.initialize-schema | no | boolean | If true, initializes the required schema. | false | true |
brxm.ai.vectorstore.pgvector.schema-name | no | string | Schema name for the vector store. | public | myschema |
brxm.ai.vectorstore.pgvector.table-name | no | string | Table name for the vector store. | vector_store | my_vector_table |
brxm.ai.vectorstore.pgvector.schema-validation | no | boolean | Enables validation of schema and table names. | false | true |
brxm.ai.vectorstore.pgvector.max-document-batch-size | no | integer | Maximum number of documents per batch. | 10000 |
Set brxm.ai.vectorstore.pgvector.initialize-schema to true the first time you connect to your PgVector instance. This creates the required schema in your PgVector database. You must complete this step before using the vector store.
Logging
To troubleshoot ingestion and processing issues, adjust the log level for com.bloomreach.xm.ai.service.impl.vector.ingest:
info: Logs ingestion events and queue usage.debug: Adds detailed logs for ingestion events, scheduler runs, and queue usage.trace: Includes logs from CMS event listeners and ingestion triggers.