Searchindex Configuration

Overview

The SearchIndex element is part of the Workspace configuration in Bloomreach Content. For additional background, refer to the Jackrabbit Wiki: Search and Indexing Configuration.

This page describes recommended parameters for configuring the SearchIndex element.

A minimal SearchIndex configuration is shown below:

<SearchIndex class="org.hippoecm.repository.FacetedNavigationEngineImpl"> <param name="indexingConfiguration" value="indexing_configuration.xml"/> <param name="indexingConfigurationClass" value="org.hippoecm.repository.query.lucene.ServicingIndexingConfigurationImpl"/> </SearchIndex>

The Jackrabbit documentation lists all available <SearchIndex> parameters. The following example shows default values used in the Bloomreach repository that have been found to work well in production environments:

<param name="useCompoundFile" value="true"/> <param name="minMergeDocs" value="1000"/> <param name="volatileIdleTime" value="10"/> <param name="maxMergeDocs" value="1000000000"/> <param name="mergeFactor" value="5"/>

Parameter details:

  • Setting maxMergeDocs too low or mergeFactor too high increases the number of Lucene indexes, which can significantly degrade query performance.
  • volatileIdleTime specifies the idle time (in seconds) before the volatile index is moved to a persistent index, even if minMergeDocs has not been reached.

Analyzer Configuration

You can configure the text analyzer used by the search index. By default, Bloomreach uses the following analyzer:

<param name="analyzer"
       value="org.hippoecm.repository.query.lucene.StandardHippoAnalyzer"/> 
  • analyzer: The default is org.hippoecm.repository.query.lucene.StandardHippoAnalyzer. You can replace this with a language-specific analyzer, such as org.apache.lucene.analysis.Analyzer.GermanAnalyzer for German content.

Consistency Check Parameters

The following parameters control search index consistency checks. All are enabled by default. For more information, see Checking and fixing search index inconsistencies.

<param name="forceConsistencyCheck" value="true"/> <param name="enableConsistencyCheck" value="true"/> <param name="autoRepair" value="true"/>

Highlighting Support

If your implementation does not require highlighting in search results (the default), disable highlighting to reduce Lucene index size:

<param name="supportHighlighting" value="false"/>

Similarity Support

By default, similarity on text between documents is enabled, while similarity on binaries is disabled. You can control this behavior with the following parameters:

<param name="supportSimilarityOnStrings" value="true"/> <param name="supportSimilarityOnBinaries" value="false"/>

Query Result Size Accuracy

Bloomreach Content supports fast retrieval of query result sizes by checking authorization in Lucene. In some authorization configurations, not all checks can be performed in Lucene. In these cases, the #getSize method of the QueryResult may return a value larger than the number of accessible results (unauthorized nodes are never returned, as authorization is always enforced when fetching nodes).

If you require exact result sizes and are willing to trade performance for accuracy, set the following parameter to true:

<param name="slowAlwaysExactSizedQueryResult" value="false"/>
Share Feedback
Page: /build/search/searchindex-configuration
Section: Build
Category *