Develop a Custom Collector

Overview

This guide explains how to create a custom collector for the Relevance Module to collect domain-specific targeting data about visitors.

When to Use

Use a custom collector when you need to gather visitor data that is not covered by the default collectors provided by the Relevance Module.

Prerequisites

  • Access to the Bloomreach Content project source code.
  • Familiarity with Maven and Java.
  • Permission to modify the CMS (platform) application.

Maven Dependencies

Custom collectors must be included in the cms (platform) application. For maintainability, add your collector to a dedicated Maven module. Make the myproject-cms-dependencies module depend on your collector module. If you deploy the platform without the CMS, ensure the platform module depends on your collector module.

Add the following dependency to the module containing your custom collector:

<dependency> <groupId>com.onehippo.cms7</groupId> <artifactId>hippo-addon-targeting-api</artifactId> </dependency>

If you plan to extend the AbstractCollector base class, add this dependency as well:

<dependency> <groupId>com.onehippo.cms7</groupId> <artifactId>hippo-addon-targeting-collectors</artifactId> </dependency>

Collector Configuration

  1. Open the Console.
  2. Add a node of type targeting:collector under /targeting:targeting/targeting:collectors.
  3. Name the node with your collector ID. This name is used as the collector's identifier.
  4. Add a String property targeting:className to the node. Set its value to the fully qualified class name of your custom collector.

Example configuration:

/targeting:targeting/targeting:collectors: /mycollector: jcr:primaryType: targeting:collector targeting:className: org.example.MyCollector

Implement the Collector Class

Create a class that implements com.onehippo.cms7.targeting.Collector. The constructor must accept a String (collector ID) and a JCR Node (the configuration node for the collector).

Alternatively, extend AbstractCollector to reuse default JSON serialization logic. This approach reduces the need to implement serialization methods manually.

Example implementation:

import javax.jcr.Node; import com.onehippo.cms7.targeting.collectors.AbstractCollector; public class MyCollector extends AbstractCollector<MyTargetingDataImpl, MyRequestData> { public MyCollector(String id, Node node) throws RepositoryException { super(id); // Read collector-specific configuration from the node if needed } /** * Returns request data collected for the current HTTP request. * @param request The HTTP request to inspect. * @param newVisitor True if this is a new visitor. * @param newVisit True if this is a new visit. * @param previousTargetingData Previously collected data for this collector, or null. * @return Processed request data, or null if no relevant data is available. */ MyRequestData getTargetingRequestData(HttpServletRequest request, boolean newVisitor, boolean newVisit, MyTargetingDataImpl previousTargetingData) { // TODO: implement } /** * Updates the visitor's targeting data with the collected request data. * @param targetingData The current targeting data, or null if this is the first call. * @param requestData The data collected from the current request, or null. * @return The updated targeting data, or null if no data is available. */ MyTargetingDataImpl updateTargetingData(MyTargetingDataImpl targetingData, MyRequestData requestData) throws IllegalArgumentException { // TODO: implement } }

Data Classes

Your collector will typically define its own classes for storing targeting data and request data. In this example, these are MyTargetingDataImpl and MyRequestData.

Targeting Data Bean

The targeting data bean holds all targeting data for a visitor. It must implement the TargetingData interface, which defines the getCollectorId method. The bean must be serializable to and from JSON using Jackson.

Example:

public class MyTargetingDataImpl extends AbstractTargetingData { @JsonCreator public MyTargetingDataImpl(@JsonProperty("collectorId") String collectorId, ...) { super(collectorId); } // Add custom fields, getters, and setters. }

Request Data Bean

The request data bean contains data collected from a single HTTP request. This data is stored in the targeting engine's request log.

Example:

public class MyRequestData { public MyRequestData(...) { } // Add custom getters and setters. }

JSON Serialization

Targeting and request data are serialized to JSON for communication with the CMS UI and for persistence. By default, Jackson is used for serialization, and you can customize it with annotations such as @JsonCreator and @JsonProperty.

If you need more control over serialization, or need to support data serialized in an older format, override the serialization methods in your collector:

T convertJsonToTargetingData(ObjectNode root, ObjectMapper objectMapper) throws IOException; JsonNode convertTargetingDataToJson(T data, ObjectMapper objectMapper) throws IOException; U convertJsonToRequestData(JsonNode root, ObjectMapper objectMapper) throws IOException; JsonNode convertRequestDataToJson(U data, ObjectMapper objectMapper) throws IOException;

The AbstractCollector base class provides default implementations for these methods. In the targeting data bean, pass the collector ID to the AbstractTargetingData base class. You can set additional properties using setters or with @JsonProperty annotations to support immutable data structures.

Access Collector Targeting Data in the HST Site Webapp

In most cases, you do not need to access targeting data in the HST site web application. If you do, retrieve the collector's targeting data from the TargetingProfile:

TargetingProfile profile = TargetingStateProvider.get().getProfile(); Map<String, TargetingData> targetingData = profile.getTargetingData(); GeoIPTargetingData geoIpTargetingData = (GeoIPTargetingData)targetingData.get("geo");

Since TargetingData implementations are part of the CMS (platform) webapp (from version 13.0.0 onward), you cannot cast to implementation classes in the site webapp. The example above works because GeoIPTargetingData is an interface in the shared library (hippo-addon-targeting-shared-api). If you want your collector's targeting data to be accessible in the site webapp, define an interface for it and include that interface in a Maven module that is part of the shared library. Place all JSON annotations on the implementation class, not the interface, as the shared library does not include Jackson dependencies.

Replace Dots in Field Names with Underscores

Elasticsearch 2, which stores visit data, does not allow dots in field names. Collectors must ensure that serialized targeting data does not include field names containing dots.

If you need to collect data with dots in the field name (for example, a cookie named myapp.rememberme.cookie), replace all dots with underscores. All default collectors in the Relevance Module follow this approach and are 'dot safe'.

If a field name with a dot is encountered during serialization, the dot is replaced by an underscore and a warning is logged.

Prevent Elasticsearch Mapping Explosion

If you use Experiments or Trends, request data is stored in Elasticsearch. If your model contains objects with arbitrary keys (such as timestamps or query parameters), you risk causing a mapping explosion in Elasticsearch.

Elasticsearch dynamically updates its mapping (schema) for each new field detected, creating an index for every new field. This consumes system resources. To avoid unbounded growth in the number of fields:

  1. Avoid using Map<> properties in your model beans.
  2. If you must use a Map<> property and the key set is small and controlled (for example, document types or a fixed configuration list), annotate the getter with @LimitedKeySet. This signals that the property is safe for Elasticsearch.
  3. If your collector must support arbitrary visitor-controlled map keys, implement custom serialization by overriding convertJsonToTargetingData(), convertTargetingDataToJson(), convertJsonToRequestData(), and convertRequestDataToJson().

Example using @LimitedKeySet:

public class MyRequestData { @JsonProperty @LimitedKeySet("Keys are bounded because there are only 5 categories") private Map<String,Integer> categoryViewPercentage; ... }

Enable Alter Ego Support

To allow users to override data collected by your custom collector using the Alter Ego feature in Experience Manager, you must also develop a collector plugin.


Related topics:

Share Feedback
Page: /build/enterprise-plugins/targeting-relevance/develop-a-custom-collector
Section: Build
Category *
Develop a Custom Collector | Bloomreach Content Documentation