Initialize and Configure the Vector Store and Ingestion Process

Overview

This guide describes how to set up and configure the Bloomreach Content AI Vector Store and the Ingestion process.

The Vector Store uses an external vector database. For on-premises projects, Redis and Postgres (with the PgVector extension) are supported. You can evaluate both locally using containerized deployments. Refer to the Redis documentation and PgVector container page for setup instructions.

Vector Store and Ingestion features are not available for Bloomreach Cloud implementations at this time.

A managed vector store for Search Agent is planned for Bloomreach Cloud in Q2 2026. Contact your Account Manager for updates.

The Ingestion Process is a background service running in a CMS pod. It listens for workflow events and updates the Vector Store as needed. The process supports two modes:

  • preview mode: Updates or removes vectors in the store when a content item is saved or deleted. The store indexes the unpublished variant, so it contains preview content.
  • live mode: Updates or removes vectors when a content item is published or taken offline. Only the published variant is indexed, so the store contains only published content.

When you first set up the Vector Store, it is empty. No content is indexed until you configure the ingestion process. For details on starting content indexing, see the section on Ingestion Process Options, including how to use the brxm.ai.ingest.mode and brxm.ai.ingest.types properties.

Info: The Vector Store is extensible. You can integrate custom vector store backends by implementing the VectorStoreFactory SPI. This allows you to use any VectorStore supported by Spring AI. See the AI Module Extensibility Guide for more information.

During ingestion, the configured embedding model generates embeddings for each document, and the Vector Store is updated with these vectors. Embedding generation incurs cost and latency at the model provider, but does not impact system performance.

Info: All AI usage costs for embedding generation are logged, regardless of model provider. Usage by the background ingestion process is attributed to the user ID system-ingestion. For audit logging details, see "Conversation token usage" in the Content Assistant developer guide.

Info: Embeddings are vectors with specific dimensions determined by the model. If you change the embedding model to one with different dimensions, existing embeddings in the vector store become incompatible and must be regenerated.

Installation

The Vector Store requires a running external Redis or Postgres instance. The Ingestion process is installed automatically when you install the AI module.

Configure via Properties Files

Configure the Vector Store and Ingestion process for production using properties files. The system checks the following locations in order of precedence (higher in the list overrides lower):

  1. System properties passed on the command line
  2. A properties file named xm-ai-service.properties available on the classpath
  3. The project's platform.properties file

Info: For more information on managing properties files and system properties:

Hint: When configuring via JCR, use the same property names as in properties files, and set them as String properties in JCR.

If no chat or embedding model is registered, the Vector Store and Ingestion process will not initialize.

Multiplicity of Configurations

You can configure multiple model providers and vector stores using properties, but only one configuration can be active at a time.

Global Configuration Options

Set the active Vector Store using the brxm.ai.vectorstore property. Supported values:

  • Redis
  • PgVector

Info: If you set this property to an empty value, both the Vector Store and Ingestion process are disabled.

Ingestion Process Options

Control the Ingestion process with the following properties:

PropertyRequiredTypeDescriptionDefaultExample
brxm.ai.ingest.modeyesenumIngestion operating mode. In preview mode, unpublished content is indexed on save/rename/copy/move. In live mode, only published content is indexed during publication (including scheduled publication).preview or live
brxm.ai.ingest.typesnolist of doctypesComma-separated list of fully qualified document types. Only these types are indexed. If not set, no content is ingested. Removal from the store ignores this filter.No types allowedmyproject:bannerdocument, hippogallery:exampleAssetSet
brxm.ai.ingest.include-dirsnolist of pathsComma-separated list of absolute paths. Documents outside these paths are skipped during ingestion. Leave empty to allow ingestion from any directory. Removal from the store ignores this filter.Any path is included/content/documents/myproject/banners/, /content/documents/myproject/news/
brxm.ai.ingest.exclude-dirsnolist of pathsComma-separated list of absolute paths. Documents under these paths are skipped during ingestion. Removal from the store ignores this filter.No path is excluded/content/documents/myproject/taxonomies/, /content/documents/myproject/private/
brxm.ai.ingest.initial-delaynointeger (seconds)Delay, in seconds, after system startup before the Ingestion process starts.3001000
brxm.ai.ingest.intervalnointeger (seconds)Frequency, in seconds, at which the process runs.1260
brxm.ai.ingest.batch-sizenointegerNumber of documents processed per batch. Reduce if your Vector Store is overloaded.52
brxm.ai.ingest.delaynointeger (milliseconds)Back-off time after each batch. Increase if your Vector Store is overloaded.100010000

Info: The ingestion process collects events in an in-memory queue and processes them in batches. Each batch is retried up to 5 times, with the delay between attempts starting at brxm.ai.ingest.delay and increasing exponentially after each failure.

The ingestion filter always excludes content items not under the absolute path /content/. This built-in filter is always applied during ingestion, but not during removal. Similarly, content under /content/attic/ or /content/taxonomies/ is always excluded from ingestion, but not from removal.

Redis Options

Info: Ensure a Redis instance is running and accessible. See the Redis Spring AI documentation for more details.

PropertyRequiredTypeDescriptionDefaultExample
brxm.ai.vectorstore.redis.hostyesurlHostname of the Redis instance.myredis, 127.0.0.1
brxm.ai.vectorstore.redis.portyesintegerPort number for Redis.6379
brxm.ai.vectorstore.redis.indexyesstringIndex name used in Redis to store embeddings.myindex
brxm.ai.vectorstore.redis.prefixyesstringPrefix for each entry in the index, used for identification.my_prefix
brxm.ai.vectorstore.redis.usernostringUsername for Redis if authentication is enabled.
brxm.ai.vectorstore.redis.passwordnostringPassword for Redis if authentication is enabled.
brxm.ai.vectorstore.redis.client-namenostringName of the connecting application, used for identification.myAppName
brxm.ai.vectorstore.redis.timeout-millisnointegerConnection timeout to Redis, in milliseconds.Redis default5000

PgVector Options

Info: Ensure a Postgres instance with the vector extension is running and accessible. See the PgVector Spring AI documentation for details.

PropertyRequiredTypeDescriptionDefaultExample
brxm.ai.vectorstore.pgvector.urlyesurlConnection string for PgVector.jdbc:postgresql://myhost:5432/mydbname
brxm.ai.vectorstore.pgvector.usernameyesstringUsername for PgVector.
brxm.ai.vectorstore.pgvector.passwordyesstringPassword for PgVector.
brxm.ai.vectorstore.pgvector.dimensionsyesstringEmbedding vector dimensions. Set during initial table creation. Changing dimensions requires recreating the vector_store table.1536
brxm.ai.vectorstore.pgvector.index-typenostringNearest neighbor search index type. Options: NONE (exact search), IVFFlat (list-based), HNSW (multilayer graph).HNSW
brxm.ai.vectorstore.pgvector.distance-typenostringSearch distance type. Default is COSINE_DISTANCE. Use EUCLIDEAN_DISTANCE or NEGATIVE_INNER_PRODUCT for normalized vectors.COSINE_DISTANCE
brxm.ai.vectorstore.pgvector.remove-existing-vector-store-tablenobooleanIf true, deletes the existing vector_store table on startup.falsetrue
brxm.ai.vectorstore.pgvector.initialize-schemanobooleanIf true, initializes the required schema.falsetrue
brxm.ai.vectorstore.pgvector.schema-namenostringSchema name for the vector store.publicmyschema
brxm.ai.vectorstore.pgvector.table-namenostringTable name for the vector store.vector_storemy_vector_table
brxm.ai.vectorstore.pgvector.schema-validationnobooleanEnables validation of schema and table names.falsetrue
brxm.ai.vectorstore.pgvector.max-document-batch-sizenointegerMaximum number of documents per batch.10000

Set brxm.ai.vectorstore.pgvector.initialize-schema to true the first time you connect to your PgVector instance. This creates the required schema in your PgVector database. You must complete this step before using the vector store.

Logging

To troubleshoot ingestion and processing issues, adjust the log level for com.bloomreach.xm.ai.service.impl.vector.ingest:

  • info: Logs ingestion events and queue usage.
  • debug: Adds detailed logs for ingestion events, scheduler runs, and queue usage.
  • trace: Includes logs from CMS event listeners and ingestion triggers.
Share Feedback
Page: /content-ai/vector-store-search-agent/initialize-and-configure-the-vector-store-and-ingestion
Section: Content AI
Category *