WAL - How PostgreSQL remembers everything it was about to do
Part 5 of 5 in Reading PostgreSQL
Ever wondered how PostgreSQL can say COMMIT in a millisecond, lose power a moment later, and still have the row when it comes back? It never wrote the table to disk at commit time. It wrote a few hundred bytes to a log, and that was enough. That log is the Write-Ahead Log (WAL), and crash recovery, streaming replication, point-in-time recovery and logical replication are all just different readers of it.
In this post, we are going to follow one change; from the moment a backend modifies a page to the moment a crashed server replays it. Here is the map of everything we are going to cover:
WAL INTERNALS
1. WHY & WHAT
a. WAL Rule
b. LSN
c. RMGRs
d. Full-Page Writes
2. WRITE PATH
a. Record Format
b. Segments/Pages
c. Insert APIs
d. Insert Locks
e. WAL Buffers
f. Critical Sections
3. DURABILITY
a. XLogFlush
b. WAL Writer
c. Group Commit
d. synchronous_commit
e. Commit record
4. CHECKPOINT & RECOVERY
a. Checkpointer
b. Redo Pointer
c. pg_control
d. Startup Process
e. Redo Loop
f. Timelines
5. CONSUMERS
a. Streaming Replication
b. Archiving/PITR
c. Logical Decoding
d. pg_waldumpReference: Everything here we are going to discuss in this blog was read from the PostgreSQL 18 source (18.4). Sizes assume the default build: 8 kB WAL pages and 16 MB segments. Code blocks marked with the
=> { }style are simplified pseudocode showing the call chain; the rest are quoted from the source file named above them.
Part 1: What is WAL and Why does it exist?
The WAL Rule
PostgreSQL changes the data pages in shared memory and writes them to disk later, in whatever order the buffer manager finds convenient. That is only safe because of one rule, written in src/backend/access/transam/README:
log entries must reach stable storage i.e., disk, before the data-page changes they describe.
How is it enforced? With a single field. Every data page starts with the LSN of the last WAL record that touched it. Before the buffer manager writes a dirty page out, it flushes the WAL up to that LSN.
src/backend/storage/buffer/bufmgr.c, in FlushBuffer():
recptr = BufferGetLSN(buf);
...
if (buf_state & BM_PERMANENT)
XLogFlush(recptr);So a commit does not need to write any data page. It only needs its WAL on disk. The pages follow lazily, and if the server crashes before that, the log is replayed to rebuild them.
Note: This check lives only in the shared buffer manager. That is why temp tables are never WAL-logged, and why unlogged tables (not
BM_PERMANENT) skip the flush.
LSN - an address in an endless file
An LSN (Log Sequence Number) is just a uint64 byte offset into the WAL stream. It is declared as XLogRecPtr in src/include/access/xlogdefs.h and printed as two hex halves like 0/17AD650.
Since it is a plain byte offset, we can work out the file and the position inside by ourself:
LSN 0/17AD650 (16 MB segments)
segment number = 0x17AD650 / 0x1000000 = 1
offset in file = 0x7AD650
file name = 000000010000000000000001
└───┘└────────┘
timeline segment numberThe LSN does three jobs across the system:
Ordering - a bigger LSN happened later.
Durability bookkeeping - "flushed up to X" is one number.
Safe replay - if a page's LSN is already at or past a record's LSN, that record has been applied.
Resource Managers
The WAL core knows nothing about heaps or B-trees. Every record carries a resource manager ID, and the core simply dispatches through a struct of function pointers. If you read the previous post on TableAmRoutine, this will look familiar.
src/include/access/xlog_internal.h :
typedef struct RmgrData
{
const char *rm_name;
void (*rm_redo) (XLogReaderState *record);
void (*rm_desc) (StringInfo buf, XLogReaderState *record);
const char *(*rm_identify) (uint8 info);
void (*rm_startup) (void);
void (*rm_cleanup) (void);
void (*rm_mask) (char *pagedata, BlockNumber blkno);
void (*rm_decode) (struct LogicalDecodingContext *ctx,
struct XLogRecordBuffer *buf);
} RmgrData;src/include/access/rmgrlist.h lists the 22 built-in ones: XLOG, Transaction, Storage, CLOG, Heap, Heap2, Btree and so on. Keep three callbacks in mind, because each one comes back later in this post:
Callback | Who calls it | What it does |
|---|---|---|
| startup process | replays the record |
|
| prints the record |
| logical decoding | turns the record into a row change |
Full-page writes
Replay normally applies a small change on top of an existing page. But what if the page on disk is broken? If the server crashes while the kernel is halfway through writing an 8 kB page, the disk holds half old and half new bytes. This is a torn page, and applying a small change on top of it gives garbage.
PostgreSQL's answer: the first time a page is modified after a checkpoint, the WAL record carries a complete image of the page. Replay restores the image instead of applying the change. The whole decision is one comparison.
src/backend/access/transam/xloginsert.c, in XLogRecordAssemble():
XLogRecPtr page_lsn = PageGetLSN(regbuf->page);
needs_backup = (page_lsn <= RedoRecPtr);RedoRecPtr is where the latest checkpoint began. If the page's LSN is at or before it, the page has not been logged since that checkpoint, so it needs an image.
Note: This is why WAL volume jumps right after every checkpoint, and it is the main reason not to checkpoint too often.
Part 2: The write path
Record Format
Let's see what a WAL record looks like. The layout is described in src/include/access/xlogrecord.h:
┌──────────────── ──┐
│ XLogRecord (24-byte header) │
├──────── ──────────┤
│ XLogRecordBlockHeader #0 │ one per page the record touches
│ XLogRecordBlockHeader #1 │
│ ... │
├──── ──────────────┤
│ XLogRecordDataHeader │
├──── ──────────────┤
│ block data #0 │ per-page payload/full-page image
│ block data #1 │
│ ... │
├─────── ───────────┤
│ main data │ record-specific, always kept
└───── ─────────────┘And the fixed header:
typedef struct XLogRecord
{
uint32 xl_tot_len; /* total len of entire record */
TransactionId xl_xid; /* xact id */
XLogRecPtr xl_prev; /* ptr to previous record in log */
uint8 xl_info; /* flag bits, see below */
RmgrId xl_rmid; /* resource manager for this record */
/* 2 bytes of padding here, initialize to zero */
pg_crc32c xl_crc; /* CRC for this record */
} XLogRecord;xl_prevchains each record to the one before it. A record whose back-link does not match is rejected, and that is one of the ways PostgreSQL finds the end of valid WAL.xl_crcis a CRC-32C over the whole record.The high four bits of
xl_infobelong to the resource manager and say which operation this is, for exampleXLOG_HEAP_INSERT.
Segments and Pages
WAL files live in pg_wal/. Each file (a segment) is 16 MB by default, and is divided into 8 kB pages. Every page starts with a small header.
src/include/access/xlog_internal.h:
typedef struct XLogPageHeaderData
{
uint16 xlp_magic; /* magic value for correctness checks */
uint16 xlp_info; /* flag bits, see below */
TimeLineID xlp_tli; /* TimeLineID of first record on page */
XLogRecPtr xlp_pageaddr; /* XLOG address of this page */
uint32 xlp_rem_len; /* total len of remaining data for record */
} XLogPageHeaderData; one 16 MB segment file
│ page 0 │ page 1 │ page 2 │ ... │ page N │ => 8 kB eachA record can be bigger than what is left on the page. In that case it continues on the next page, whose header gets the XLP_FIRST_IS_CONTRECORD flag, and xlp_rem_len says how many bytes are still to come.
The Insert API
Any code that changes a WAL-logged page follows the same seven steps:
Pin and exclusive-lock the buffer.
START_CRIT_SECTION()Change the page.
MarkBufferDirty()Build and insert the WAL record, then stamp the page with the returned LSN.
END_CRIT_SECTION()Unlock and unpin.
Now let's see step 5 for a plain INSERT, in heap_insert() of src/backend/access/heap/heapam.c:
XLogBeginInsert();
XLogRegisterData(&xlrec, SizeOfHeapInsert);
...
XLogRegisterBuffer(0, buffer, REGBUF_STANDARD | bufflags);
XLogRegisterBufData(0, &xlhdr, SizeOfHeapHeader);
XLogRegisterBufData(0,
(char *) heaptup->t_data + SizeofHeapTupleHeader,
heaptup->t_len - SizeofHeapTupleHeader);
...
recptr = XLogInsert(RM_HEAP_ID, info);
PageSetLSN(page, recptr);If we observe, the tuple is registered with XLogRegisterBufData() and not XLogRegisterData(). That is deliberate. Data attached to a buffer is dropped from the record when a full-page image is taken, because the image already contains it. Main data is always kept.
Insertion Locks - Reserve one by one, copy in Parallel
What happens inside XLogInsert()? Here is the path:
XLogInsert()
│
XLogRecordAssemble() build the record, no locks held
│
XLogInsertRecord()
│
├─ take one of the 8 WAL insertion locks
├─ did a checkpoint move RedoRecPtr, so that a page
│ now needs a full-page image?
│ └── yes ──► release, go back and assemble again
├─ ReserveXLogInsertLocation() spinlock, bump CurrBytePos
├─ fill xl_prev, finish the CRC
├─ CopyXLogRecordToWAL() memcpy into the WAL buffers
└─ release the insertion lock
│
return end LSN ──► caller does PageSetLSN()The interesting part is that the insert is split into two steps.
Step 1, reserve. This is the only part every backend must do one at a time, so it is kept tiny.
In ReserveXLogInsertLocation() of src/backend/access/transam/xlog.c:
SpinLockAcquire(&Insert->insertpos_lck);
startbytepos = Insert->CurrBytePos;
endbytepos = startbytepos + size;
prevbytepos = Insert->PrevBytePos;
Insert->CurrBytePos = endbytepos;
Insert->PrevBytePos = startbytepos;
SpinLockRelease(&Insert->insertpos_lck);CurrBytePos counts only "usable" bytes, which means it ignores page headers. So reserving space is a single addition. Converting it into a real LSN happens after the spinlock is released.
Step 2, copy. Each backend copies its record into the space it reserved. There are 8 insertion locks (NUM_XLOGINSERT_LOCKS) and holding any one of them is enough, so up to eight backends can copy at the same time.
Each lock also carries an insertingAt value. Whoever wants to flush the WAL has to be sure that all copies into that region are finished. WaitXLogInsertionsToFinish() walks the eight locks and waits only for the inserters that are still behind the flush target.
WAL Buffers
The records are copied into XLogCtl->pages, a ring of 8 kB pages in shared memory. Its size is wal_buffers, and the default (-1) means auto-tune.
src/backend/access/transam/xlog.c, in XLOGChooseNumBuffers():
xbuffers = NBuffers / 32;
if (xbuffers > (wal_segment_size / XLOG_BLCKSZ(8KB)))
xbuffers = (wal_segment_size / XLOG_BLCKSZ);
if (xbuffers < 8)
xbuffers = 8;So it is about 3% of shared_buffers, with a minimum of 8 pages and a maximum of one segment.
What if the ring is full? Then the inserting backend has to write an old page to disk itself, right in the middle of its insert, inside AdvanceXLInsertBuffer(). PostgreSQL counts this as wal_buffers_full in pg_stat_wal. If that number keeps growing on your system, wal_buffers is too small.
Critical Sections
Between changing the page and inserting the record, the shared buffer holds a change that the log knows nothing about. If the backend threw a normal error at that point, the page could reach disk without its WAL.
START_CRIT_SECTION() turns any error into a PANIC. The server restarts and recovers from the log, which is always consistent. This is also why the code checks everything that can fail (free space, memory) before entering the critical section.
Part 3: Durability
At this point our record is sitting in shared memory. Nothing is durable yet.
Three Pointers
PostgreSQL tracks three positions in the WAL, always in this order:
on disk, fsynced written to OS copied to WAL buffers
────┼────────────┼─────────────┼──────► LSN
logFlushResult logWriteResult logInsertResultOur commit is durable once logFlushResult has passed the end of our commit record.
The Commit Record
Now let's see the path of the code when we run COMMIT:
CommitTransaction() => RecordTransactionCommit() {
XactLogCommitRecord(); // insert the commit record into WAL
XLogFlush(XactLastRecEnd); // make it durable
=> {
WaitXLogInsertionsToFinish(); // let in-flight copies finish
XLogWrite(); => {
pg_pwrite(); // WAL buffers -> segment file
issue_xlog_fsync(); // fdatasync / fsync
}
}
TransactionIdCommitTree(); // only now: mark committed in pg_xact
SyncRepWaitForLSN(); // wait for standby, if configured
}The order is the whole point: insert the commit record, flush the WAL, and only then mark the transaction committed. The real branch in src/backend/access/transam/xact.c looks like this:
if ((wrote_xlog && markXidCommitted &&
synchronous_commit > SYNCHRONOUS_COMMIT_OFF) ||
forceSyncCommit || nrels > 0)
{
XLogFlush(XactLastRecEnd);
...
}
else
{
XLogSetAsyncXactLSN(XactLastRecEnd);
...
}Note:
nrels > 0means the transaction is dropping relation files. Such a commit is always flushed synchronously, whatever your setting is, so that files are never deleted before the commit is on disk.
XLogFlush and group commit
fsync is the expensive part. Group commit means one fsync covers many transactions, and PostgreSQL gets it from a few lines in XLogFlush():
for (;;)
{
/* done already? */
RefreshXLogWriteResult(LogwrtResult);
if (record <= LogwrtResult.Flush)
break;
...
insertpos = WaitXLogInsertionsToFinish(WriteRqstPtr);
if (!LWLockAcquireOrWait(WALWriteLock, LW_EXCLUSIVE))
continue;
...
/* try to write/flush later additions to XLOG as well */
WriteRqst.Write = insertpos;
WriteRqst.Flush = insertpos;
XLogWrite(WriteRqst, insertTLI, false);
LWLockRelease(WALWriteLock);
break;
}Two things make this work:
The leader flushes more than it needs. It writes up to
insertpos, which is everything any backend has finished inserting, not only its own record.The followers wait without taking the lock.
LWLockAcquireOrWait()sleeps until the lock is free and returns without acquiring it. The follower loops, finds its record already flushed, and leaves.
backend A ─ XLogFlush ─ gets WALWriteLock ── write + fsync ─ done
backend B ─ XLogFlush ─ waits ............. already flushed ─ done
backend C ─ XLogFlush ─ waits ............. already flushed ─ done
--------------------------------------------------------------------
one fsync, three commitssynchronous_commit
Setting |
|
|---|---|
| the commit record is in the WAL buffers |
| local flush |
| local flush, and the standby has written it |
| local flush, and the standby has flushed it |
| local flush, and the standby has replayed it |
With no synchronous standby configured, the three remote levels i.e., remote_write/on/remote_apply behave like local.
Is off dangerous? We can lose the last few commits after a crash. We cannot get a corrupted database, because the WAL rule still holds for every data page.
The WAL Writer
WalWriterMain() in src/backend/postmaster/walwriter.c is a background process that calls XLogBackgroundFlush() every wal_writer_delay (200 ms by default). It does three jobs:
Writes the completed WAL pages, so that backends rarely hit a full ring.
Flushes asynchronous commits. The source comment gives the bound: an async commit reaches disk after at most three
wal_writer_delaycycles.Prepares the next WAL buffer pages in advance.
Part 4: Checkpoints
The WAL cannot grow forever, and recovery cannot start from the beginning of time. A checkpoint creates a point from which replaying is enough.
What triggers a checkpoint?
The checkpointer process (src/backend/postmaster/checkpointer.c) starts one when:
checkpoint_timeouthas passed (default 300 seconds), orenough WAL has been written, or
someone asks for it (
CHECKPOINT, shutdown, base backup, end of recovery).
"Enough WAL" comes from max_wal_size. In CalculateCheckpointSegments():
target = (double) ConvertToXSegs(max_wal_size_mb (1GB), wal_segment_size (16MB)) /
(1.0 + CheckPointCompletionTarget(0.9));With the defaults that is 64 segments / 1.9, rounded down to 33 segments. So roughly every 528 MB of WAL i.e., 33x16MB.
The redo pointer
A checkpoint is not instant. Writing every dirty buffer can take minutes, and the system keeps running meanwhile. So a checkpoint has two positions:
checkpoint in progress
◄─ dirty buffers being written─►
───┬────────────────────┬────────────► WAL
│ │
XLOG_CHECKPOINT_REDO XLOG_CHECKPOINT_ONLINE
(redo pointer) (checkpoint record)
recovery starts replaying here. pg_control points hereRecovery starts at the redo pointer, not at the checkpoint record. Here is the simplified path of CreateCheckPoint() in src/backend/access/transam/xlog.c:
CreateCheckPoint()
=> {
XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO); // its LSN = new RedoRecPtr
// full-page images start again
// wait for backends that are between
// "commit record inserted" and "pg_xact updated"
CheckPointGuts()
=> {
CheckPointCLOG(); // and the other SLRUs
CheckPointBuffers(); // write all dirty shared buffers
ProcessSyncRequests(); // fsync the data files
}
XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_ONLINE); // carries the CheckPoint struct
XLogFlush();
UpdateControlFile(); // pg_control now points to this checkpoint
RemoveOldXlogFiles(); // delete or recycle old segments
}The buffer writing is spread over checkpoint_completion_target (0.9) of the interval, so that the checkpoint does not flood your disks all at once.
Note: Old segments are not always removed.
KeepLogSeg()holds back what replication slots andwal_keep_sizestill need. A stuck replication slot is the usual reasonpg_wal/fills up a disk.
pg_control
global/pg_control is the one file PostgreSQL reads before it can read any WAL. The fields that matter here, from src/include/catalog/pg_control.h:
DBState state;
XLogRecPtr checkPoint; /* last check point record ptr */
CheckPoint checkPointCopy; /* copy of last check point record */
XLogRecPtr minRecoveryPoint;The file is 8 kB on disk, but the struct is asserted to fit in 512 bytes. That way an update is a single disk sector write, which is as close to atomic as a disk gets.
Part 5: Recovery
The startup process
When the server starts, the startup process runs StartupXLOG(). Let's see the path:
StartupXLOG()
=> {
InitWalRecovery()
=> {
read_backup_label(); // restored from a base backup?
// else take the checkpoint location from pg_control
ReadCheckpointRecord(); // and get its redo pointer
// decide: is recovery needed?
}
PerformWalRecovery()
=> {
RmgrStartup();
do {
recoveryStopsBefore(); // PITR target reached?
ApplyWalRecord()
=> rm_redo() // e.g. heap_redo()
=> heap_xlog_insert()
=> XLogReadBufferForRedo()
recoveryStopsAfter();
record = ReadRecord();
} while (record != NULL);
RmgrCleanup();
}
FinishWalRecovery();
PerformRecoveryXLogAction(); // end-of-recovery checkpoint or record
}Recovery is needed if the redo pointer is before the checkpoint record, or if state i.e., DB State in pg_control is anything other than DB_SHUTDOWNED. A server that was shut down cleanly skips the redo loop completely.
The redo loop
The whole dispatch is one line inside ApplyWalRecord() of src/backend/access/transam/xlogrecovery.c:
GetRmgr(record->xl_rmid).rm_redo(xlogreader);This is the rm_redo callback from Part 1. The resource manager does the real work, and almost every redo function begins by asking what to do with each page. The answer comes from XLogReadBufferForRedoExtended() in src/backend/access/transam/xlogutils.c:
Result | Meaning | What redo does |
|---|---|---|
| the record carried a full-page image | image is copied over the page, done |
| record LSN <= page LSN | already applied, skip |
| page is older than the record | apply the change, set the page LSN |
| page does not exist | skip, the table was dropped or truncated later |
BLK_DONE is what makes replay safe to repeat. Recovery can crash and start again from the same redo pointer any number of times.
And where does it stop? Crash recovery has no stored end point. It stops when ReadRecord() returns NULL because the next record is not valid: a bad CRC, a wrong xl_prev link, an invalid length or a bad page header.
Note: A standby replays forever and cannot write checkpoints of its own. It does restartpoints instead: it flushes its buffers at a checkpoint record it has already replayed, so that its own restart does not begin from scratch.
Timelines
Suppose we restore a backup, and recovered up to yesterday 14:00, and start writing. Our new WAL uses the same LSN range as the WAL the original server wrote after 14:00. Now two different histories share the same LSNs.
timeline 1
──────────●───────────────► (original history)
│
│ recovered up to here
│
timeline 2 └───────────────► (new history)
files: 00000001000000000000000A 00000002000000000000000A
└───┘ └───┘
timeline 1 timeline 2The timeline ID keeps them apart. It is the first eight hex digits of every segment file name. In StartupXLOG():
newTLI = endOfRecoveryInfo->lastRecTLI;
if (ArchiveRecoveryRequested)
{
newTLI = findNewestTimeLine(recoveryTargetTLI) + 1;
...
XLogInitNewTimeline(EndOfLogTLI, EndOfLog, newTLI);
...
writeTimeLineHistory(newTLI, recoveryTargetTLI,
EndOfLog, endOfRecoveryInfo->recoveryStopReason);
}A plain crash recovery stays on the same timeline.
Archive recovery and standby promotion always move to a new one.
The
0000000N.historyfile records the family tree, one line per fork: parent timeline, switch LSN and the reason.
Part 6: Who reads the WAL?
Once the WAL is flushed, several independent readers use it. All of them share the same reader code in src/backend/access/transam/xlogreader.c.
Flushed WAL will be read by the following:
1. WalSender: for sending raw bytes to the standby
2. archiver: for copying finished segments and store away
3. logical decoding: to give a row change to the subscriber
4. pg_waldump: prints records for humansStreaming Replication
The walsender on the primary ships raw WAL bytes in XLogSendPhysical() of src/backend/replication/walsender.c:
It sends only up to the flush pointer. A standby never receives WAL that the primary could still lose.
It sends at most 16 pages (128 kB) per message.
On the standby, the walreceiver writes those bytes into its own pg_wal/, fsyncs them, and wakes up the startup process, which is running the very same redo loop from Part 5.
Archiving and PITR
When a segment is completed, XLogArchiveNotifySeg() creates an empty file pg_wal/archive_status/<segment>.ready. The archiver process picks it up, hands the segment to your archive_command (or archive module), and renames the marker to .done.
Point-in-time recovery is the same redo loop with two differences: the segments come from the archive, and recoveryStopsBefore() / recoveryStopsAfter() end the loop at your target. Then it moves to a new timeline.
Logical Decoding
Physical WAL says "put these bytes in block 1234 of file 16384". Logical decoding turns that back into "a row was inserted into orders". It needs wal_level = logical.
src/backend/replication/logical/decode.c, in LogicalDecodingProcessRecord():
rmgr = GetRmgr(XLogRecGetRmid(record));
if (rmgr.rm_decode != NULL)
rmgr.rm_decode(ctx, &buf);This is the rm_decode callback from Part 1. Only a few resource managers have one (Heap, Heap2, Transaction and a couple more). Index records are ignored, because the subscriber maintains its own indexes.
XLogSendLogical()
=> LogicalDecodingProcessRecord()
=> heap_decode()
=> DecodeInsert()
=> ReorderBufferQueueChange() // hold the change, keyed by xid
...
=> xact_decode()
=> DecodeCommit()
=> ReorderBufferCommit() // replay that xact's changes in order
=> begin / change / commit callbacks of the output pluginChanges of concurrent transactions are mixed together in the WAL, so they are held in the reorder buffer until the commit record shows up. An aborted transaction's changes are simply thrown away.
pg_waldump
src/bin/pg_waldump/pg_waldump.c is the smallest reader and the best learning tool. It loops over XLogReadRecord() and prints each record using the rm_desc callback.
Part 7: Hands on
Now let's watch all of this happen on a real cluster. The output below is from a fresh PostgreSQL 18.4 cluster with default settings (wal_level = replica, full_page_writes = on, wal_compression = off).
CREATE TABLE t (id int, v text);
SELECT pg_current_wal_insert_lsn(); -- 0/17AD510
INSERT INTO t VALUES (1, 'first');
SELECT pg_current_wal_insert_lsn(); -- 0/17AD580
CHECKPOINT;
SELECT pg_current_wal_insert_lsn(); -- 0/17AD650
INSERT INTO t VALUES (2, 'hello');
SELECT pg_current_wal_insert_lsn(); -- 0/17AD720
INSERT INTO t VALUES (3, 'world');
SELECT pg_current_wal_insert_lsn(); -- 0/17AD790
SELECT pg_walfile_name('0/17AD650'); -- 000000010000000000000001Each LSN is the boundary between two steps, so we can dump the WAL of every step separately with pg_waldump -s <start> -e <end>.
The first insert into an empty table
pg_waldump -p $PGDATA/pg_wal -s 0/17AD510 -e 0/17AD580rmgr: Heap len (rec/tot): 65/ 65, tx: 753, lsn: 0/017AD510, prev 0/017AD0C8, desc: INSERT+INIT off: 1, flags: 0x00, blkref #0: rel 1663/5/16384 blk 0
rmgr: Transaction len (rec/tot): 34/ 34, tx: 753, lsn: 0/017AD558, prev 0/017AD510, desc: COMMIT 2026-10-04 10:57:02.232238 ISTTwo records: the insert and its commit. If we observe each line closely, everything from Part 2 is there:
rmgris the resource manager,txisxl_xid,previsxl_prev. The commit'sprev(0/017AD510) is exactly the insert'slsn. That is the back-link chain.blkref #0: rel 1663/5/16384 blk 0is the block reference: tablespace / database / relfilenode, and the block number.INSERT+INITmeans this tuple is the first one on its page.heap_insert()registers the buffer withREGBUF_WILL_INIT, so redo rebuilds the page from scratch. No page image is needed, and that is as safe as having one.
The checkpoint
pg_waldump -p $PGDATA/pg_wal -s 0/17AD580 -e 0/17AD650rmgr: XLOG len (rec/tot): 30/ 30, tx: 0, lsn: 0/017AD580, prev 0/017AD558, desc: CHECKPOINT_REDO wal_level replica
rmgr: Standby len (rec/tot): 50/ 50, tx: 0, lsn: 0/017AD5A0, prev 0/017AD580, desc: RUNNING_XACTS nextXid 754 latestCompletedXid 753 oldestRunningXid 754
rmgr: XLOG len (rec/tot): 114/ 114, tx: 0, lsn: 0/017AD5D8, prev 0/017AD5A0, desc: CHECKPOINT_ONLINE redo 0/17AD580; tli 1; prev tli 1; fpw true; wal_level replica; xid 0:754; oid 24576; multi 1; offset 0; oldest xid 744 in DB 1; oldest multi 1 in DB 1; oldest/newest commit timestamp xid: 0/0; oldest running xid 754; onlineThis is the two-position checkpoint from Part 4, exactly as described:
CHECKPOINT_REDOat 0/17AD580 marks where the checkpoint started.CHECKPOINT_ONLINEat 0/17AD5D8 is written when it finished, and its body saysredo 0/17AD580, pointing back to the first record.The
RUNNING_XACTSrecord in between is the snapshot of running transactions that a hot standby needs.
And pg_control agrees:
pg_controldata $PGDATADatabase cluster state: in production
Latest checkpoint location: 0/17AD5D8
Latest checkpoint's REDO location: 0/17AD580
Latest checkpoint's REDO WAL file: 000000010000000000000001
Latest checkpoint's TimeLineID: 1On this idle cluster the two locations are only 88 bytes apart. On a busy system the gap is all the WAL that was written while the checkpoint was running.
The first insert after the checkpoint
pg_waldump -p $PGDATA/pg_wal -s 0/17AD650 -e 0/17AD720rmgr: Heap len (rec/tot): 54/ 166, tx: 754, lsn: 0/017AD650, prev 0/017AD5D8, desc: INSERT off: 2, flags: 0x00, blkref #0: rel 1663/5/16384 blk 0 FPW
rmgr: Transaction len (rec/tot): 34/ 34, tx: 754, lsn: 0/017AD6F8, prev 0/017AD650, desc: COMMIT 2026-10-04 10:57:02.261017 ISTThere is the FPW at the end of the line. The page's LSN was before the new redo pointer, so needs_backup became true and the record carries a full-page image.
Look at len (rec/tot): 54/166:
totgrew to 166 bytes. The extra 112 bytes are the page image.An 8 kB page fits in 112 bytes because the page has only two small tuples, and the hole between
pd_lowerandpd_upperis left out.recshrank from 65 to 54. The tuple was registered withXLogRegisterBufData(), so it is dropped when an image is taken, just as Part 2 said.
The next insert
pg_waldump -p $PGDATA/pg_wal -s 0/17AD720 -e 0/17AD790rmgr: Heap len (rec/tot): 65/ 65, tx: 755, lsn: 0/017AD720, prev 0/017AD6F8, desc: INSERT off: 3, flags: 0x00, blkref #0: rel 1663/5/16384 blk 0
rmgr: Transaction len (rec/tot): 34/ 34, tx: 755, lsn: 0/017AD768, prev 0/017AD720, desc: COMMIT 2026-10-04 10:57:02.274960 ISTNo FPW any more. The page has already been logged once since the checkpoint, so the record is back to 65 bytes. It will stay that way until the next checkpoint moves the redo pointer again.
The totals
pg_waldump -p $PGDATA/pg_wal -s 0/17AD510 -e 0/17AD790 --statsType N (%) Record size (%) FPI size (%) Combined size (%)
---- - --- ----------- --- -------- --- ------------- ---
XLOG 2 ( 22.22) 144 ( 30.00) 0 ( 0.00) 144 ( 24.32)
Transaction 3 ( 33.33) 102 ( 21.25) 0 ( 0.00) 102 ( 17.23)
Standby 1 ( 11.11) 50 ( 10.42) 0 ( 0.00) 50 ( 8.45)
Heap 3 ( 33.33) 184 ( 38.33) 112 (100.00) 296 ( 50.00)
-------- -------- -------- --------
Total 9 480 [81.08%] 112 [18.92%] 592 [100%]Note: The real output lists all 22 resource managers. The rows with zero records are removed here.
Three one-row inserts and a checkpoint cost 592 bytes of WAL, and 19% of that is a single page image. On a real workload with full pages, the FPI share right after a checkpoint is far higher. Run --stats on one of your own busy segments to see it.
Closing Notes
This post covered why WAL exists, how a record is built and inserted, what makes a commit durable, how checkpoints bound the recovery work, how the redo loop replays, and who else reads the log. What it didn't cover is two-phase commit, hint-bit logging, WAL summarization for incremental backups, and how hot standby resolves conflicts with running queries; those deserve their own deep dives. For now, src/backend/access/transam/README and xlog.c are the two best files to keep open while you read, and pg_waldump is the best tool to keep running next to them.
Enjoyed this post?
1 reaction