1 · The situation
Your app runs on Postgres. The daily-active-users chart now takes 5 seconds, and the weekly report times out. So you copy the data into ClickHouse.
On our own test data, a default Postgres took 5.4 s for "daily active users, last 30 days" and 22 s for "events per kind per week". The same queries took 0.1 to 1.6 seconds on ClickHouse. So the copy is worth doing. This page is about doing it without losing data.
2 · What goes wrong
The copy looks right. Then it quietly drifts.
We set up a table of 20,000 rows and replicated it to two ClickHouse replicas with Postgres’s default settings. Then, in Postgres, we updated 4,000 rows and deleted 1,000.
- Deleted rows are still in ClickHouse. Postgres has 19,000 rows. Both replicas have 20,000.
- Long text arrives empty. In one range of ids, the text column holds 124,215 bytes in Postgres and 1,335 bytes in ClickHouse.
- Nothing reports an error. Replication shows as running. The only thing that noticed was a check that compared both sides.
chlift migrate check fails. Text version: 07-corruption-repro.log.3 · Why it happens
Postgres sends less than you think.
Logical replication sends each change with the table's replica identity. With the default identity, a delete carries only the primary key, and an update leaves out large values (TOASTed text and JSON) that did not change. ClickHouse tables are usually sorted by time first, so a delete that carries only the id cannot find the row it should remove, and the row stays. An update without the large value writes an empty one in its place.
4 · What to do (with or without chlift)
Three rules.
- Send the whole row. On every table that gets updates or deletes, before you start the copy:
Postgres then writes the full old row to its WAL on each update and delete, so WAL grows. Plan disk space for it.ALTER TABLE public.user_events REPLICA IDENTITY FULL; - Compare, don't count. A total row count hides errors that cancel out. Compare per day and per primary-key range: counts, ids, and checksums of the columns, on every replica.
- Cut over only on a match. Switch readers to ClickHouse only when every day and every key range is equal on both sides.
5 · How chlift does it
The same rules, as commands.
chlift is a command-line tool. It installs a ClickHouse cluster on your own Linux hosts over SSH (replicas, Keeper, and PeerDB for change data capture), then runs and proves the migration.
chlift install- Sets up the ClickHouse replicas, Keeper and PeerDB on your hosts. Running it again changes nothing unless the config changed.
chlift migrate plan- Picks the event tables and finds the risky ones (time-first sort keys, updated tables with long text or JSON). It writes the plan, including the
REPLICA IDENTITY FULLlines, for you to review. chlift migrate start --apply-postgres- Applies those Postgres settings, copies a snapshot to every replica, then keeps it current with live change data capture.
chlift migrate check --checksums- Compares row counts for every day and checksums for every primary-key range, on every replica.
chlift migrate cutover- Waits until replication has caught up, runs the check again, and stops replication only if 0 ranges differ. ClickHouse keeps the data.
With the setting in place, the same kind of changes gave 18,000 rows in Postgres and on both replicas, and 0 of 64 key ranges differed.