Reaggregating Visits
This guide describes how to rebuild the visit store after schema changes. Use this process when the original requests are still available in the request log store.
Preparation
- Configure an
INFOlevel logger forcom.onehippo.cms7.targeting.dataflowto monitor the reaggregation process. - Ensure that no experiments are running. Complete all active experiments in Experience Manager. After completion and channel publication, verify that there are no child nodes under:
/targeting:targeting: /targeting:experiments: - Stop the dataflow jobs by setting the
runningproperty tofalseon both themodelTrainerandvisitsAggregatornodes:/targeting:targeting: /targeting:dataflow: /modelTrainer: running: false /visitsAggregator: running: false - Wait until both jobs log that they have been disabled. This typically occurs within 10 seconds.
Recreate the Index
You can either create a new index with a different name or replace the existing index. Creating a new index with a unique name is recommended for the following reasons:
- The original visits data remains available if you encounter issues.
- Changing the index name forces the Visit Store to restart and upload the correct mapping (schema).
The Elasticsearch visits index name is defined in the indexName property at:
/targeting:targeting: /targeting:datastores: /visits:
Important: Ensure that no running CMS instance is writing data to Elasticsearch during this process. On startup, the CMS must initialize the search index mappings correctly.
Option 1: Create a New Index with a Different Name
- Create the new index in Elasticsearch. For example:
curl -s -S -XPUT http://elastic.host:9200/newindexname - Update the
indexNameproperty in the console to the new index name and save the changes.
Option 2: Replace the Existing Index
- Delete the old index in Elasticsearch. For example:
curl -s -S -XDELETE http://elastic.host:9200/indexname - Create the new index in Elasticsearch. For example:
curl -s -S -XPUT http://elastic.host:9200/indexname - In the console, temporarily add a
dummyproperty to the/targeting:targeting/targeting:datastores/visitsnode and save:
This triggers a restart of the Visit Store./targeting:targeting: /targeting:datastores: /visits: dummy: bla - After a few minutes, remove the
dummyproperty.
Restart the Data Flow Jobs
- Remove the
processedUntilproperty from thevisitsAggregatornode, but leave it unchanged on themodelTrainernode. Set therunningproperties totrue:/targeting:targeting: /targeting:dataflow: /modelTrainer: running: true // processedUntil unchanged /visitsAggregator: running: true // processedUntil REMOVED - Within a few seconds, both jobs will log that they are enabled. The VisitsAggregator will log that it is aggregating requests into visits in batches of up to 10,000.