Configure the Lucene Analyzer
Overview
Bloomreach Content uses the org.hippoecm.repository.query.lucene.StandardHippoAnalyzer as the default Lucene Analyzer for content indexing. This analyzer removes stop words for English, German, Dutch, French, Spanish, Brazilian Portuguese, and Czech. It also applies an ISO Latin 1 accent filter, which replaces accented characters such as ç with c and ï with i.
Customizing Stop Words in StandardHippoAnalyzer
Info: Customizing stop words is supported in brXM 12.0 and later.
The StandardHippoAnalyzer uses language-specific stop word lists, each stored in a separate classpath resource file. You can override these files to customize the stop words for your implementation.
The following resource files define the stop words for each supported language:
- Default:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer.properties - English:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_en.properties - Spanish:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_es.properties - French:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_fr.properties - German:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_de.properties - Dutch:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_nl.properties - Brazilian Portuguese:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_pt_BR.properties - Czech:
classpath:org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_cs.properties
To customize the stop words for a language, create a file with the same name and path in your project’s classpath. The custom file will override the default.
For example, the English stop words configuration looks like this:
# The delimiters to use when splitting stopwords.split.tokens value. stopwords.split.delimiters=, # Whether or not to preserve all the tokens including empty string token. stopwords.split.preserveAllTokens=true # Stopwords tokens. stopwords.split.tokens=a,and,are,as,at,be,but,by,for,if,in,into,is,it,no,not,of,on,or,s,such,t,that,the,their,then,there,these,they,this,to,was,will,with,,www
To add additional stop words such as "etc" or "ie", append them to the stopwords.split.tokens property, separated by commas. For example, to customize English stop words in a project where cms/ contains the repository instance, add your custom file at cms/src/main/resources/org/hippoecm/repository/query/lucene/StandardHippoAnalyzer_en.properties.
Using a Custom Lucene Analyzer
You can configure a custom analyzer to add features such as stemming. Note that using a custom analyzer may disable wildcard search functionality. This limitation is due to how Lucene handles tokenization and inverted indexes. If you require wildcard search, use the StandardHippoAnalyzer.
Configure the Analyzer Class
The analyzer class is defined in the repository.xml configuration file.
To use a different analyzer, update the following parameter:
<param name="analyzer" value="org.hippoecm.repository.query.lucene.StandardHippoAnalyzer"/>
Replace the value with the fully qualified class name of your custom analyzer.
For instructions on customizing and deploying repository.xml, see Repository deployment settings.