BUG #19530: Crash (SIGSEGV/SIGBUS) in parallel B-tree index vacuum during plain VACUUM
BUG #19530: Crash (SIGSEGV/SIGBUS) in parallel B-tree index vacuum during plain VACUUM
От:
PG Bug reporting form <noreply@postgresql.org>
Дата:
The following bug has been logged on the website: Bug reference: 19530 Logged by: Maciej Email address: webmaster@algedi.pl PostgreSQL version: 17.7 Operating system: FreeBSD 14.4-RELEASE, x86-64 Description: PostgreSQL 17.7 intermittently crashes a parallel maintenance worker during routine VACUUM, taking the server through automatic crash recovery each time. Summary: Recurring backend / parallel-worker terminations with signal 11 (SIGSEGV) and signal 10 (SIGBUS). Each crash is followed by "terminating any other active server processes" and crash recovery (~5 s). It is intermittent: about 16 process terminations over 5 days, with fully clean days in between; the same nightly maintenance succeeds on most days. Hardware was checked and is clean (no MCE/ECC/I/O errors). Environment: PostgreSQL 17.7, FreeBSD 14.4-RELEASE, x86-64, physical server, 256 GB RAM. dynamic_shared_memory_type=posix, huge_pages=try. max_parallel_workers=8, max_parallel_maintenance_workers=4, max_parallel_workers_per_gather=4. Extensions: pg_stat_statements, pg_trgm, unaccent, plpgsql only — no custom C extensions, no PL/Perl/Python. What the crashing process was doing: The crashing process is a parallel maintenance worker. A preserved core dump shows the process title "parallel worker for PID ", with memory contexts "BTree Vacuum State" and "AutoVacuum Data" present and no executor/query state — i.e. parallel index vacuum, not parallel query. The trigger is a scheduled VACUUM (SKIP_DATABASE_STATS, VERBOSE, ANALYZE) run via vacuumdb over several large tables, each carrying many (~10–30) btree indexes, so parallel index vacuuming is engaged. The SIGBUS variant is consistent with a DSM (POSIX shared memory) problem in the parallel path. Expected vs actual: Expected: VACUUM completes normally. Actual: a parallel index-vacuum worker crashes (SIGSEGV/SIGBUS) and the server goes through crash recovery. Possibly related: This appears to reach the same btbulkdelete / btvacuumscan code as the report "Segmentation fault in PostgreSQL 17.7 during REINDEX TABLE CONCURRENTLY" (https://www.postgresql.org/message-id/VI0PR07MB10718A4DC292E068C51404AB2974EA@VI0PR07MB10718.eurprd07.prod.outlook.com), but reached via the plain-VACUUM index-vacuum path rather than validate_index / ReindexRelationConcurrently. That suggests the fault is in the (parallel) B-tree index-vacuum code, reachable from both paths. Workaround: Setting max_parallel_maintenance_workers = 0 (sequential index vacuuming) appears to stop the crashes (observation ongoing).
Re: BUG #19530: Crash (SIGSEGV/SIGBUS) in parallel B-tree index vacuum during plain VACUUM
От:
John Naylor <johncnaylorls@gmail.com>
Дата:
On Mon, Jun 22, 2026 at 3:19 PM PG Bug reporting form wrote: > PostgreSQL 17.7 intermittently crashes a parallel maintenance worker during > routine VACUUM, taking the server through automatic crash recovery each > time. > What the crashing process was doing: > The crashing process is a parallel maintenance worker. A preserved core > dump shows the process title "parallel worker for PID ", with memory > contexts "BTree Vacuum State" and "AutoVacuum Data" present and no This doesn't really give us any actionable info. Are you able to get a backtrace from the core dump? (Autovacuum in v17 cannot use parallel workers, so I'm not sure where that came from) > Possibly related: > This appears to reach the same btbulkdelete / btvacuumscan code as the > report "Segmentation fault in PostgreSQL 17.7 during REINDEX TABLE > CONCURRENTLY" > (https://www.postgresql.org/message-id/VI0PR07MB10718A4DC292E068C51404AB2974EA@VI0PR07MB10718.eurprd07.prod.outlook.com), > but reached > via the plain-VACUUM index-vacuum path rather than validate_index / > ReindexRelationConcurrently. That suggests the fault is in the (parallel) > B-tree > index-vacuum code, reachable from both paths. That operation was writing TIDs to a file, which is not what vacuum does, and it's not clear from that thread whether there's a bug in the first place. -- John Naylor Amazon Web Services