Elasticsearch Data Store
Overview
The Experiments and Trends features in the Relevance module require Elasticsearch.
Info: In Bloomreach Content versions up to 14.0.0, the Relevance module always included Experiments and Trends, so Elasticsearch was required. Starting with versions 14.0.1 and 14.1.0, Experiments and Trends are optional. Elasticsearch is only required if you install these features.
brXM 14 supports Elasticsearch versions 6.x and 7.x (7.x supported from version 14.3.0).
brXM 15 supports Elasticsearch versions 6.x, 7.x, and 8.1.
Refer to the System Requirements page for the exact supported versions.
Install Elasticsearch
Download and install Elasticsearch.
Info: For installation, configuration, deployment, and administration instructions, refer to the Elasticsearch documentation. Bloomreach Content does not require any custom setup steps for Elasticsearch. For production environments, deploy a cluster with at least two Elasticsearch nodes to ensure high availability.
Select a Stale Data Removal Strategy
To manage the size of the Elasticsearch index, choose one of the following strategies:
- Scheduled cleanup job (Relevance Module):
The relevance engine automatically deletes entries older than a specified number of days. Configure the maximum age in the targeting datasource within the application context configuration. - Rollover Index API (Elasticsearch):
Elasticsearch manages index rollover and aliases. When the application connects, it uploads an index template for the visit type. Elasticsearch applies this mapping to new indices and moves the alias to the new index during rollover. Configure the template name and alias name in the targeting datasource in the application context configuration.
Configure Visits Data Store
The Relevance Elasticsearch Data Store connects to Elasticsearch using a JNDI data source lookup. You must define this at the container level (for example, in Apache Tomcat).
Depending on your stale data removal strategy, add one of the following environment entries to conf/context.xml in your project.
For the scheduled cleanup job strategy:
<Environment name="elasticsearch/targetingDS" type="java.lang.String" value="{'indexName':'visits','maxAgeDays':'60', 'locations':['url-1','url-2]',...]}" />
For the rollover index strategy:
<Environment name="elasticsearch/targetingDS" type="java.lang.String" value="{'templateName':'myproject-hippo_relevance_visit', 'aliasName':'visits', 'locations':['url-1','url-2]',...]}" />
This configuration registers a JNDI environment resource at java/comp/env/elasticsearch/targetingDS when the site web application starts. The JSON string defines the properties required to connect to your Elasticsearch cluster.
Replace ['url-1','url-2]',...] with the URLs of your Elasticsearch cluster nodes. For local development, set locations to ['http://localhost:9200'].
JSON Field Reference
| Field | Type | Default | Description |
|---|---|---|---|
indexName¹ | String | n/a | Name of the Elasticsearch index (required for the scheduled cleanup job strategy). |
templateName² | String | n/a | Name of the index template (required for the rollover index strategy). Use a descriptive name to avoid collisions. |
aliasName² | String | n/a | Name of the alias (required for the rollover index strategy). |
locations³ | String array | n/a | URLs of Elasticsearch cluster nodes. One location is sufficient, but specifying multiple nodes improves robustness. |
username | String | n/a | Optional. Username for authenticated access to Elasticsearch. |
password | String | n/a | Optional. Password for authenticated access to Elasticsearch. |
maxConnections | Long | 20 | Optional. Maximum number of client threads in the connection pool. |
maxAgeDays | Long | 397 | Records older than this value (in days) are deleted. Set to 0 to disable deletion. |
cleanupJobCronTrigger⁴ | String | n/a | Optional. Cron expression for scheduling cleanup jobs. Cleanup jobs run only if maxAgeDays is greater than 0. If not set, jobs run every hour. |
¹ Required for the scheduled cleanup job strategy.
² Required for the rollover index strategy.
³ Required for both strategies.
⁴ Available since version 13.4.0.
Configure the JNDI Resource in the Visits Store
Below are example configurations for the visits store in your Bloomreach Content project.
Elasticsearch 6:
/targeting:targeting/targeting:datastores/targeting:visits: targeting:storefactoryclass: com.onehippo.cms7.targeting.storage.elastic6.ElasticStoreFactory dataSource: elasticsearch/targetingDS
Elasticsearch 7 (brXM 14.3.0 and later):
/targeting:targeting/targeting:datastores/targeting:visits: targeting:storefactoryclass: com.onehippo.cms7.targeting.storage.elastic7.ElasticStoreFactory dataSource: elasticsearch/targetingDS
Configure Elasticsearch
Scheduled Cleanup Job Strategy
Create the index specified by the indexName property in Elasticsearch. For example, use the following curl command:
curl -s -S -XPUT http://localhost:9200/visits
If the index does not exist when the CMS starts, the application will attempt to create it.
Rollover Index Strategy
Create the aliased index specified by the aliasName property in Elasticsearch. For example, use the following curl command:
curl -XPUT 'localhost:9200/%3Cvisits-%7Bnow%2Fd%7D-000001%3E' -d
'{
"aliases": {
"visits": {}
}
}'
This command creates an initial index named visits-YYYY.MM.dd-000001, where YYYY.MM.dd is the current date. The alias for this index is visits. Ensure that the index name starts with the alias, as the application uses ${alias}* for queries.
The index must be readable and writable by the users configured through the authentication property. The method for configuring access depends on your Elasticsearch deployment. Consult your system administrator for details on creating and securing the index in your environment.