Check and Fix Repository Inconsistencies

Overview

This page describes how to identify and repair inconsistencies between the database that backs the content repository and the JCR (Java Content Repository) model in Bloomreach Content.

Purpose

Use the consistency checker to detect and resolve discrepancies between the database and the JCR model.

Context

Bloomreach Content uses a JCR repository based on Apache Jackrabbit. Under certain conditions, the underlying database may become inconsistent with the JCR model. These inconsistencies can cause errors in the CMS or site, and may prevent the repository from starting. To analyze and repair the database, you can run a consistency check.

Choosing a Consistency Checker

Run consistency checks using the stand-alone Checker tool. Although Apache Jackrabbit supports running consistency checks during repository startup, the stand-alone checker provides more control and flexibility.

Running the Checker in Check Mode

First, download and configure the Checker tool.

To perform a consistency check, run:

java -jar hippo-addon-checker-<version>.jar check

This command checks the consistency of all configured workspaces, as specified in the checker.properties file. If you set the check.history property, the checker also verifies the version history. To check referential integrity, set the check.references property.

For initial analysis, run an integrity check on the default workspace. This is the most important and fastest check. Only check the version history and referential integrity if necessary. Version history checks can take longer, and referential integrity checks require loading both the workspace and version history, since workspaces may reference nodes in the version history. The checker does not perform referential integrity checks within the version history itself, as the JCR specification does not require it.

To check specific nodes, provide one or more UUIDs after the check command:

java -jar hippo-addon-checker-<version>.jar check [UUID1 UUID2 UUIDn] [--recursive]

Add the --recursive option to check all descendants of the specified nodes. This approach also applies when running the checker in fix mode.

Running the Checker in Fix Mode

Before running the checker in fix mode on a database used by active repository instances, review the following:

  • Fixing inconsistencies in a live cluster is supported from version 1.02.00 of the checker tool onward.
  • If you use an older version, you must shut down the entire cluster before running a consistency fix.

To repair a database in a live cluster, configure clustering for both the checker and all repository instances that use the same database. The sample checker-repository.xml file, created according to the generic configuration instructions, includes an example cluster configuration. This configuration ensures that the checker notifies other repository instances about changes made during repairs. Without this notification, local caches in other instances may become outdated, leading to data loss when those instances overwrite changes made by the checker.

To run the checker in repair mode:

java -jar hippo-addon-checker-<version>.jar fix [UUID1 UUID2 UUIDn] [--recursive]

For information about the causes of data corruption in Jackrabbit and details on types of inconsistencies, see the sections below.

Understanding the Bundle Table

Each workspace and the version history store nodes in a separate bundle table. Each table has two columns: one for the node ID and one for the bundle data (a binary large object). The bundle data includes property names and values, type information, and structural information such as parent and child node IDs. Both parent and child relationships are stored, which can lead to inconsistencies if these references become misaligned.

Orphaned Nodes

Orphaned nodes are nodes whose parent no longer exists in the bundle table. This can occur if a parent node is removed but the operation is incomplete, leaving child nodes in the database. The checker can move orphans to a dedicated folder if you configure it accordingly. To do this:

  • Create a lost+found node for each workspace, of type nt:unstructured.
  • Specify the UUID of this node in checker.properties using check.default.lostnfound=uuid, where default matches the value of check.workspaces.

The checker reports orphaned nodes with the message:

NodeState '{nodeId}' references inexistent parent id '{parentNodeId}'

Abandoned Nodes

Abandoned nodes reference an existing parent, but the parent does not list them as a child. This often results from an incomplete remove operation. The checker fixes this by adding a child node entry to the parent. Since the node name is not stored in the bundle, a generated name is assigned. The checker reports abandoned nodes with the message:

NodeState '{nodeState}' is not referenced by its parent node '{parentNodeId}'

Missing Nodes

Missing nodes occur when a node lists a child node entry that no longer exists. This is also often due to incomplete remove operations. The checker resolves this by removing the child node entry. The checker reports missing nodes with the message:

NodeState '{nodeId}' references inexistent child '{childNodeId}'

Disconnected Nodes

Disconnected nodes have a child node that does not reference them as a parent. This can result from incomplete move operations. The checker resolves this by removing the child node entry from the parent. The checker reports disconnected nodes with the message:

Node has invalid parent id: '{parentNodeId}' (instead of '{nodeId}')

Orphaned and missing nodes may result from earlier corruption, such as a node first becoming disconnected and then deleted. Missing and disconnected nodes typically cause the most immediate problems, as the repository may fail during traversal.

Manual Inspection After Fix

After running the checker:

  • Orphaned nodes are moved to the lost and found folder. Inspect these nodes using the console.
  • Abandoned nodes are reattached to their parent with a generated name. If the node is a handle or document node, verify and adjust the name as needed.

Review all fixed nodes to ensure they are correctly integrated into the repository structure.

Resolving UUIDs in the Checker Report to JCR Paths

The fix-mode console log identifies repaired nodes by UUID. To map these UUIDs to JCR paths, use the CheckerLogResolver Updater script, available in the CMS from release 17.2 onward.

Procedure

  1. Capture the fix output to a file:

    java -jar hippo-addon-checker-<version>.jar fix > hippo-fix.log
    
  2. Upload the log file as an asset in the CMS (for example, /content/assets/test/hippo-fix-1.log).

  3. In the Updater script editor, select the registered CheckerLogResolver script and set the parameters:

    { "logFile": "/content/assets/test/hippo-fix-1.log", "lostFoundUuid": "" }
    

The script generates a report between the BEGIN CSV REPORT and END CSV REPORT markers in the run log. Each row contains path,issues,info, where info is one of: abandoned, missing, disconnected, orphan, orphan-related, or orphan-untraceable.

Share Feedback
Page: /deploy/administration/maintenance/checking-and-fixing-repository-inconsistencies
Section: Deploy
Category *
Check and Fix Repository Inconsistencies | Bloomreach Content Documentation