Searchindex Configuration
Overview
The SearchIndex element is part of the Workspace configuration in Bloomreach Content. For additional background, refer to the Jackrabbit Wiki: Search and Indexing Configuration.
This page describes recommended parameters for configuring the SearchIndex element.
A minimal SearchIndex configuration is shown below:
<SearchIndex class="org.hippoecm.repository.FacetedNavigationEngineImpl"> <param name="indexingConfiguration" value="indexing_configuration.xml"/> <param name="indexingConfigurationClass" value="org.hippoecm.repository.query.lucene.ServicingIndexingConfigurationImpl"/> </SearchIndex>
Recommended Parameters
The Jackrabbit documentation lists all available <SearchIndex> parameters. The following example shows default values used in the Bloomreach repository that have been found to work well in production environments:
<param name="useCompoundFile" value="true"/> <param name="minMergeDocs" value="1000"/> <param name="volatileIdleTime" value="10"/> <param name="maxMergeDocs" value="1000000000"/> <param name="mergeFactor" value="5"/>
Parameter details:
- Setting
maxMergeDocstoo low ormergeFactortoo high increases the number of Lucene indexes, which can significantly degrade query performance. volatileIdleTimespecifies the idle time (in seconds) before the volatile index is moved to a persistent index, even ifminMergeDocshas not been reached.
Analyzer Configuration
You can configure the text analyzer used by the search index. By default, Bloomreach uses the following analyzer:
<param name="analyzer"
value="org.hippoecm.repository.query.lucene.StandardHippoAnalyzer"/>
analyzer: The default isorg.hippoecm.repository.query.lucene.StandardHippoAnalyzer. You can replace this with a language-specific analyzer, such asorg.apache.lucene.analysis.Analyzer.GermanAnalyzerfor German content.
Consistency Check Parameters
The following parameters control search index consistency checks. All are enabled by default. For more information, see Checking and fixing search index inconsistencies.
<param name="forceConsistencyCheck" value="true"/> <param name="enableConsistencyCheck" value="true"/> <param name="autoRepair" value="true"/>
Highlighting Support
If your implementation does not require highlighting in search results (the default), disable highlighting to reduce Lucene index size:
<param name="supportHighlighting" value="false"/>
Similarity Support
By default, similarity on text between documents is enabled, while similarity on binaries is disabled. You can control this behavior with the following parameters:
<param name="supportSimilarityOnStrings" value="true"/> <param name="supportSimilarityOnBinaries" value="false"/>
Query Result Size Accuracy
Bloomreach Content supports fast retrieval of query result sizes by checking authorization in Lucene. In some authorization configurations, not all checks can be performed in Lucene. In these cases, the #getSize method of the QueryResult may return a value larger than the number of accessible results (unauthorized nodes are never returned, as authorization is always enforced when fetching nodes).
If you require exact result sizes and are willing to trade performance for accuracy, set the following parameter to true:
<param name="slowAlwaysExactSizedQueryResult" value="false"/>