Repository Maintenance
Removing Old Revisions from the Repository Journal
Jackrabbit 2 uses a database-backed journal to coordinate cluster messaging. Over time, this journal can grow significantly, depending on the volume of write operations in the repository. To manage database size and performance, periodically remove outdated journal entries.
The REPOSITORY_LOCAL_REVISIONS table tracks the latest journal revision processed by each cluster node. You can safely delete entries from the REPOSITORY_JOURNAL table that are older than the lowest revision recorded in REPOSITORY_LOCAL_REVISIONS.
To view the current revisions in the database, run:
mysql> SELECT * FROM REPOSITORY_LOCAL_REVISIONS;
+---------------+-------------+
| JOURNAL_ID | REVISION_ID |
+---------------+-------------+
| cms-node1 | 2302839 |
| cms-node2 | 2302839 |
| site-node1 | 2302839 |
| site-node2 | 2302839 |
+---------------+-------------+
4 rows in set (0.00 sec)
mysql> SELECT * FROM REPOSITORY_GLOBAL_REVISION;
+-------------+
| REVISION_ID |
+-------------+
| 2302839 |
+-------------+
1 row in set (0.00 sec)
Each JOURNAL_ID value corresponds to a Jackrabbit cluster node ID.
You can remove old journal entries in two ways:
- By executing SQL statements directly
- By using Jackrabbit's built-in cleanup mechanism
Remove Old Revisions Using SQL
To manually delete outdated entries from the journal table, execute the following SQL statements:
MySQL example for cleaning up the REPOSITORY_JOURNAL table:
mysql> DELETE FROM REPOSITORY_JOURNAL WHERE REVISION_ID < ANY (SELECT min(REVISION_ID)
FROM REPOSITORY_LOCAL_REVISIONS);
Query OK, 0 rows affected (0.00 sec)
mysql> OPTIMIZE TABLE REPOSITORY_JOURNAL;
+-----------------------------+----------+----------+----------------------------+
| Table | Op | Msg_type | Msg_text |
+-----------------------------+----------+----------+----------------------------+
| hippocms.REPOSITORY_JOURNAL | optimize | note | Table does not support |
| | | | optimize, doing recreate + |
| | | | analyze instead |
| hippocms.REPOSITORY_JOURNAL | optimize | status | OK |
+-----------------------------+----------+----------+----------------------------+
2 rows in set (0.03 sec)
Considerations
Adding a Node to the Cluster
When you add a new node to the cluster, it attempts to read the entire REPOSITORY_JOURNAL table during startup. If the table contains a large number of entries, this may cause out-of-memory errors. To avoid this, clean up old revisions before adding new cluster nodes.
Removing a Node from the Cluster
When you remove a cluster node, its entry in REPOSITORY_LOCAL_REVISIONS remains. To ensure safe cleanup of old revisions, manually remove entries for inactive nodes:
DELETE FROM REPOSITORY_LOCAL_REVISIONS WHERE JOURNAL_ID NOT IN (%s) AND JOURNAL_ID NOT LIKE '_HIPPO_EXTERNAL%%'
Replace %s with a comma-separated list of active JCR cluster node IDs.
Lucene Index Export
Deleting old revisions as described above invalidates all but the most recent Lucene index export files. The latest export remains valid because of the dedicated entry '_HIPPO_EXTERNAL_REPO_SYNC_index-backup' in the local revisions table.