Write an Updater Script
Overview
This page describes how to write a Groovy Updater Script to perform bulk updates on repository content in Bloomreach Content.
When to Use
Use an updater script when you need to make bulk changes to existing content in a running repository. Updater scripts provide full access to the JCR API and are executed in the context of the repository.
Prerequisites
- Access to the CMS as an administrator.
- Familiarity with Groovy and the JCR API.
- Authorization to run updater scripts in your environment.
Security Considerations
Updater scripts can modify significant portions of your repository. Only trusted developers and administrators should have access to create and run these scripts.
Scripts are executed using a custom Groovy ClassLoader that blocks some dangerous operations (such as System.exit()), but this does not provide a secure sandbox. Groovy updater scripts can technically execute external programs and may compromise the server environment if misused. Always restrict script execution to trusted users.
Create a New Script
- Log in to the CMS as an administrator.
- Navigate to
Setup>System. - Select
Updater Editor. - Click
Newto create a new script. - Enter a name for the script.
- Configure execution options as needed. For details, see Run an Updater Script.
Implementing NodeUpdateVisitor
Updater scripts must be written in Groovy and implement the NodeUpdateVisitor interface:
/** * Visitor for updating repository content. Replaces * {@link org.hippoecm.repository.ext.UpdaterModule}s for all update tasks * except backward incompatible node type changes. */ public interface NodeUpdateVisitor { /** * Allows initialization of this updater. Called before any other method is * called. * * @param session a JCR {@link Session} with system credentials * @throws RepositoryException when thrown, the updater will not be run by * the framework */ void initialize(Session session) throws RepositoryException; /** * Update the given node. * * @param node the {@link Node} to be updated * @return <code>true</code> if the node was changed, <code>false</code> * if not * @throws RepositoryException if an exception occurred while updating * the node */ boolean doUpdate(Node node) throws RepositoryException; /** * Revert the given node. This method is intended to be the reverse of the * {@link #doUpdate} method. * It allows update runs to be reverted in case a problem arises due to the * update. The method should throw an {@link UnsupportedOperationException} * when it is not implemented. * * @param node the node to be reverted. * @return <code>true</code> if the node was changed, <code>false</code> * if not * @throws RepositoryException if an exception occurred while reverting * the node * @throws UnsupportedOperationException if the method is not implemented */ boolean undoUpdate(Node node) throws RepositoryException, UnsupportedOperationException; /** * Allows cleanup of resources held by this updater. Called after an * updater run was completed. */ void destroy(); }
Most updater scripts extend the BaseNodeUpdateVisitor base class. This class provides a logger and default (no-op) implementations for initialize and destroy.
The updater engine uses the visitor pattern. For each node, it calls the script's doUpdate method. If the script modifies the node, return true from doUpdate to notify the engine.
The following template logs the path of each visited node:
package org.hippoecm.frontend.plugins.cms.admin.updater import org.onehippo.repository.update.BaseNodeUpdateVisitor import javax.jcr.Node import javax.jcr.RepositoryException import javax.jcr.Session class UpdaterTemplate extends BaseNodeUpdateVisitor { boolean logSkippedNodePaths() { return false // don't log skipped node paths } boolean skipCheckoutNodes() { return false // return true for readonly visitors and/or updates unrelated to versioned content } Node firstNode(final Session session) throws RepositoryException { return null // implement when using custom node selection/navigation } Node nextNode() throws RepositoryException { return null // implement when using custom node selection/navigation } boolean doUpdate(Node node) { log.debug "Updating node ${node.path}" return false } boolean undoUpdate(Node node) { throw new UnsupportedOperationException('Updater does not implement undoUpdate method') } }
The node parameter is a javax.jcr.Node object, which provides full access to the repository.
For a basic implementation, see example 1 (Add a property) at Groovy Updater Scripts Examples.
Optional Features
Using Parameters
To make updater scripts reusable, define parameters instead of hard-coding values. Specify parameters in the execution options as a JSON string mapping parameter names to values.
Access parameters in your script using the parametersMap variable. For example, if you set:
{ "basePath": "/content/documents/myproject/news", "tag" : "gogreen" }
You can access these parameters in your script:
def basePath = parametersMap["basePath"] def tag = parametersMap["tag"] log.debug "basePath: ${basePath}, tag: ${tag}"
Supporting Undo
To allow undoing script changes, implement the undoUpdate method. This method should revert the node to its previous state before doUpdate was called.
See example 1 (Add a property) at Groovy Updater Scripts Examples for an implementation.
Custom Node Visiting Logic
Info: Available in brXM v12.1.1 and later (also backported to v12.0.4, v11.2.5, and v10.2.9).
By default, the nodes to visit are specified by an XPath query or repository path in the execution options. Alternatively, you can override the following methods in BaseNodeUpdateVisitor to implement custom node selection:
/** * Initiates the retrieval of the nodes when using custom, instead of path or xpath (query) based, node * selection/navigation, returning the first node to visit. Intended to be overridden, default implementation returns null. * @param session * @return first node to visit, or null if none found * @throws RepositoryException */ public Node firstNode(final Session session) throws RepositoryException { return null; } /** * Return a following node, when using custom, instead of path or xpath (query) based, node selection/navigation. * Intended to be overridden, default implementation returns null. * @return next node to visit, or null if none left * @throws RepositoryException */ public Node nextNode() throws RepositoryException { return null; }
Example: To visit all nodes of type hippo:document (similar to using the XPath query //element(*, hippo:document)):
private NodeIterator nodeIterator; Node firstNode(final Session session) throws RepositoryException { final javax.jcr.query.QueryManager queryManager = session.getWorkspace().getQueryManager(); final javax.jcr.query.Query jcrQuery = queryManager.createQuery("//element(*, hippo:document)", "xpath"); nodeIterator = jcrQuery.execute().getNodes(); return nextNode(); } Node nextNode() throws RepositoryException { return nodeIterator.hasNext() ? nodeIterator.next() : null; }
When using a repository path or XPath query, the updater engine first collects all nodes before calling doUpdate. With custom node visiting logic, doUpdate is called during iteration, which can improve efficiency and allows cancellation during processing.
Avoid selecting the
rep:rootnode with an XPath query and performing all processing in a singledoUpdatecall. This approach prevents cancellation and is not recommended.
Overriding Default Behavior
You can override two boolean methods in BaseNodeUpdateVisitor to change default behavior:
skipCheckoutNodes()
By default, this method returns false, and nodes are checked out before visiting to allow updates. If your script only queries or updates non-versioned content, override this method to return true to avoid unnecessary checkouts.
/** * Overridable boolean function to indicate if node checkout can be skipped (default false) * @return true if node checkout can be skipped (e.g. for readonly visitors and/or updates unrelated to versioned content) */ public boolean skipCheckoutNodes() { return false; }
Info: Available in brXM v12.1.1 and later (also backported to v12.0.4, v11.2.5, and v10.2.9).
logSkippedNodePaths()
By default, this method returns true, and all skipped node paths (where doUpdate returns false) are logged as an audit trail. If many nodes are skipped and logging is unnecessary, override this method to return false.
/** * Overridable boolean function to indicate if skipped node paths should be logged (default true) * @return true if skipped node paths should be logged */ public boolean logSkippedNodePaths() { return true; }
Info: Available in brXM v12.1.1 and later (also backported to v12.0.4, v11.2.5, and v10.2.9).
Manually Reporting Updated, Skipped, or Failed Nodes
By default, the updater engine tracks updated, skipped, and failed nodes for each doUpdate(Node) call. If your script's update logic does not align with the default node iteration (for example, if you perform manual queries and iteration), the automatic reporting may not reflect actual changes. In these cases, use the visitorContext variable (org.onehippo.repository.update.NodeUpdateVisitorContext) to manually report node status.
Example: Manually reporting updated nodes during custom iteration:
/** * ExampleNewsDocumentDateFieldUpdateDemoVisitor is a script that does manual node iteration * in an original iteration cycle and reports updated node manually in order to be aligned * with the built-in batch commit/revert feature of the updater engine for demonstration purpose. */ package org.hippoecm.frontend.plugins.cms.admin.updater import org.onehippo.repository.update.BaseNodeUpdateVisitor import java.util.* import javax.jcr.* import javax.jcr.query.* class ExampleNewsDocumentDateFieldUpdateDemoVisitor extends BaseNodeUpdateVisitor { boolean doUpdate(Node node) { log.debug "Visiting node at ${node.path} just as an entry point in this demo." // new date field value from the current time def now = Calendar.getInstance() // do manual query and node iteration def query = node.session.workspace.queryManager.createQuery("//element(*,demosite:newsdocument)", "xpath") def result = query.execute() for (NodeIterator nodeIt = result.getNodes(); nodeIt.hasNext(); ) { def newsNode = nodeIt.nextNode() newsNode.setProperty("demosite:date", now) // report updated to the engine manually here. visitorContext.reportUpdated(newsNode.path) } return false } boolean undoUpdate(Node node) { throw new UnsupportedOperationException('Updater does not implement undoUpdate method') } }
In this example, visitorContext.reportUpdated(path) is called after updating each node. This allows the updater engine to track updated nodes and manage batch processing according to the configured batch size.
Additional Notes
Default Imports
The script classloader automatically imports the main JCR API packages: javax.jcr, javax.jcr.nodetype, javax.jcr.security, and javax.jcr.version. You do not need to import these packages explicitly.
Restrictions
The following restrictions apply to updater scripts:
- File system access is disabled. The following classes are not available:
java.io.File,java.io.FileDescriptor,java.io.FileInputStream,java.io.FileOutputStream,java.io.FileWriter,java.io.FileReader. - The following packages are not available:
java.nio.file,java.net,javax.net,javax.net.ssl. - Reflection is not allowed. You cannot use
Class.forNameor thejava.lang.reflectpackage. - Calling
System.exitis blocked.
Additional limitations may apply when running updater scripts automatically at startup, depending on the environment. In a delivery-tier-only environment, only the Hippo Repository functionality may be available on the classpath.
Portability
When executed from the Updater Editor, scripts use the CMS application context classloader. All libraries packaged with your CMS application are available. To maximize portability and allow reuse across projects, avoid dependencies on project-specific libraries. Prefer using libraries and APIs available in the shared class loader. Libraries such as commons-collections and guava are generally available.
For scripts executed automatically at startup, only repository context classes may be available in a delivery-tier-only environment.
Hint: For maximum portability, write updater scripts that depend only on repository libraries, not on CMS libraries.