layout: doc lang: en translation_key: BENCHMARKS title: PostgreSQL cache benchmarks, 3.1.0 description: “Two-machine pg_local_cache 3.1.0 results for pinned and full-server reads, write overhead, and reads during writes, with raw JSONL and reproduction steps.” section: Benchmarks permalink: /docs/BENCHMARKS.html redirect_from: - /docs/benchmarks-go.html - /docs/benchmarks-node.html - /docs/benchmarks-go/ - /docs/benchmarks-node/

last_modified_at: “2026-10-07”

PostgreSQL cache benchmarks: 3.1.0

These measurements compare pg_local_cache RESP MGET, Valkey, and prepared SQL on two private-network VMs. Every result below uses final 3.1.0 build c431bcc. Repeated-run medians are shown for the pinned and all-core read matrices and write-overhead tests; both pinned one-core Valkey io-threads=1 cases have two runs (other pinned cases have five). Mixed, 120-second stale-probe, and one-hour soak results are single runs. These workload samples are not production capacity promises.

What the final-build runs show

  • Per pinned server vCPU pair, the best measured Valkey setting was io-threads=1 at one pair (180k reads/s vs 210k for pg_local_cache); at two pairs it was io-threads=4 (266k vs 337k). At four pairs, throughput reached a client-bound plateau near 361k vs 357k, while pg_local_cache used 5.24 vs 7.95 server vCPU. Valkey io-threads=1 used the least server CPU per request. Prepared SQL was far behind in the pinned single-key tests.
  • With all 16 vCPUs available, observed median MGET 16/64 throughput differences with Valkey io-threads=8 ranged from about −2.15% to +1.26%. The observed ranges did not overlap in any of the four 64-key comparisons, so these samples do not establish statistical equivalence. Wide-MGET runs were client/network-bound. For random-key MGET 64 at 256 clients, pg_local_cache had higher p99 in one row (43.52 vs 40.37 ms).
  • Write overhead versus plain table transactions was 3.6–8.2% for UPDATE and 3.7–5.5% for INSERT. Plain UPDATE plus Valkey DEL was 28.9–63.7% below plain UPDATE throughput.
  • During writes, stale-entry checks found zero pg_local_cache entries. The 120-second Valkey cache-aside probes ended with 2 Uniform and 1 Zipf stale entries. The one-hour soak completed 252M MGET and 72M UPDATE with worker RssAnon growth capped at 44 kB.

Reads per pinned server vCPU pair

The Go/pgx client used 256 clients and one key per MGET. Each case ran for 20 seconds. The two one-core Valkey io-threads=1 cases ran twice; all other pinned cases ran five times. PostgreSQL and Valkey were pinned to N vCPU pairs (2N vCPU). Valkey used io-threads=2N and io-threads=1; prepared SQL used the same key sets.

vCPU pairs Keys pg_local_cache req/s (p99 ms, vCPU) Valkey io-threads=2N (req/s, p99 ms, vCPU) Valkey io-threads=1 (req/s, p99 ms, vCPU) Prepared SQL (req/s, p99 ms, vCPU)
1 Hot 209,696 (2.46, 1.99) 91,478 (3.24, 2.00) 180,141 (2.59, 0.85) 23,200 (24.90, 2.00)
1 100k random 195,950 (2.65, 2.00) 88,672 (3.38, 2.00) 175,853 (2.65, 0.92) 22,433 (25.95, 2.00)
2 Hot 336,691 (2.02, 3.78) 266,340 (1.69, 4.00) 162,230 (2.85, 0.66) 46,386 (16.38, 4.00)
2 100k random 326,125 (1.98, 3.83) 267,826 (1.72, 4.00) 160,220 (2.92, 0.69) 45,211 (17.56, 4.00)
4 Hot 360,968 (2.33, 5.24) 357,161 (2.20, 7.95) 160,633 (2.85, 0.66) 88,604 (6.75, 8.00)
4 100k random 353,941 (2.33, 5.27) 351,344 (2.20, 7.95) 159,414 (2.92, 0.68) 86,790 (6.88, 8.00)

At one vCPU pair, single-threaded Valkey was the strongest Valkey configuration. At two pairs, io-threads=4 was strongest; Valkey p99 was lower despite lower throughput. At four pairs, pg_local_cache and io-threads=8 were near the same client-bound throughput, with lower pg_local_cache server CPU.

Raw runs: Valkey io-threads=2N, Valkey io-threads=1.

Reads with all 16 vCPUs available

These runs were unpinned and repeated three times. The load node reached about 7–13 CPU cores in client-bound cases. With Valkey io-threads=8, observed median MGET 16/64 throughput differences ranged from about −2.15% to +1.26%; the observed throughput ranges did not overlap in any of the four 64-key comparisons. These samples do not establish statistical equivalence. With io-threads=1, random-key MGET 16 at 256 clients measured 116,130 vs 97,489 req/s (19.1% higher for pg_local_cache). Wide MGETs were client/network-bound; the io-threads=1 column has no 64-key result.

Keys MGET keys Clients pg_local_cache (req/s, p99 ms, vCPU) Valkey io-threads=8 (req/s, p99 ms, vCPU) Valkey io-threads=1 (req/s, p99 ms, vCPU) Prepared SQL (req/s, p99 ms, vCPU)
Hot 1 64 217,724 (0.70, 2.89) 206,634 (0.70, 7.85) 167,929 (0.71, 0.66) 132,433 (0.96, 11.32)
Hot 1 256 372,683 (2.39, 4.21) 365,947 (2.33, 7.92) 161,455 (2.92, 0.68) 173,571 (4.26, 13.14)
Hot 16 64 82,858 (3.38, 3.13) 83,098 (3.38, 6.41) 79,179 (3.38, 0.69) 73,512 (2.05, 8.78)
Hot 16 256 122,496 (9.83, 4.07) 122,817 (9.83, 7.42) 111,790 (8.06, 0.83) 100,966 (6.36, 11.49)
Hot 64 64 30,241 (8.78, 3.40) 30,845 (9.04, 3.57) — 27,790 (6.88, 9.05)
Hot 64 256 39,531 (43.52, 4.26) 39,995 (47.71, 5.66) — 34,519 (30.67, 12.23)
100k random 1 64 215,310 (0.70, 2.89) 205,081 (0.70, 7.85) 162,778 (0.73, 0.67) 131,563 (0.97, 11.34)
100k random 1 256 366,860 (2.39, 4.31) 360,894 (2.33, 7.92) 158,444 (2.98, 0.68) 169,884 (4.26, 13.19)
100k random 16 64 79,505 (3.51, 3.45) 79,556 (3.44, 6.68) 74,270 (3.38, 0.79) 69,262 (2.13, 9.12)
100k random 16 256 116,130 (10.35, 4.42) 114,683 (9.83, 7.53) 97,489 (7.67, 0.95) 91,572 (7.01, 11.96)
100k random 64 64 29,000 (8.78, 3.81) 29,464 (9.04, 4.61) — 26,169 (7.14, 9.39)
100k random 64 256 37,356 (43.52, 4.93) 38,175 (40.37, 7.58) — 32,756 (31.20, 13.03)

For random-key MGET 64 at 256 clients, pg_local_cache p99 was 43.52 ms versus 40.37 ms for Valkey. Prepared SQL throughput was substantially lower in the single-key pinned tests; full results vary by batch size and latency metric.

Raw runs: Valkey io-threads=8, Valkey io-threads=1.

Write overhead

Three 15-second repetitions compare plain table transactions, transactions through pg_local_cache, and plain transactions followed by Valkey DEL. Percentages compare throughput with the matching plain-table case.

Operation synchronous_commit Clients Plain table tx/s (DB µs/tx) pg_local_cache tx/s (DB µs/tx, change) Plain + Valkey DEL tx/s (DB µs/tx, change)
Update on 32 46,747 (109) 45,065 (119, −3.6%) 33,234 (204, −28.9%)
Update on 64 63,479 (108) 61,027 (121, −3.9%) 37,659 (238, −40.7%)
Update off 32 119,156 (75) 110,987 (87, −6.9%) 53,374 (158, −55.2%)
Update off 64 148,358 (73) 136,194 (83, −8.2%) 53,853 (189, −63.7%)
Insert on 32 48,831 (97) 47,020 (105, −3.7%) 34,863 (184, −28.6%)
Insert on 64 65,688 (97) 63,011 (108, −4.1%) 39,642 (218, −39.7%)
Insert off 32 123,679 (63) 117,868 (72, −4.7%) 58,168 (135, −53.0%)
Insert off 64 138,454 (68) 130,805 (77, −5.5%) 58,361 (165, −57.8%)

Across UPDATE cases, pg_local_cache was 3.6–8.2% below plain writes; INSERT overhead was 3.7–5.5%. Plain UPDATE plus Valkey DEL was 28.9–63.7% below plain UPDATE throughput; the INSERT comparison was 28.6–57.8% below.

Raw runs: write overhead.

Reads during writes

Each case ran once for 30 seconds with 64 readers issuing one-key reads over 60k rows; writers used synchronous_commit=off. Valkey used cache-aside (SQL write followed by DEL) with io-threads=8. Stale is the end-of-run sentinel check; n/a means the stack has no cache entries to inspect.

Stack Key distribution Writer Reads/s Read p99 ms Hit ratio Writes/s Stale entries
pg_local_cache Uniform None 214,508 0.71 1.000 0 0
pg_local_cache Uniform Update 10,000/s 195,384 0.86 0.948 10,000 0
pg_local_cache Uniform Update 30,000/s 161,919 1.04 0.837 29,999 0
pg_local_cache Uniform Update unlimited 82,640 1.69 0.432 111,013 0
pg_local_cache Zipf None 216,297 0.70 1.000 0 0
pg_local_cache Zipf Update 10,000/s 198,454 0.86 0.964 10,000 0
pg_local_cache Zipf Update 30,000/s 168,875 1.01 0.923 30,000 0
pg_local_cache Zipf Update unlimited 96,297 1.46 0.834 113,505 0
Valkey cache-aside Uniform None 223,537 0.61 1.000 0 0
Valkey cache-aside Uniform Update 10,000/s 156,983 1.92 0.936 10,000 0
Valkey cache-aside Uniform Update 30,000/s 42,294 7.14 0.577 29,998 0
Valkey cache-aside Uniform Update unlimited 29,692 8.78 0.444 38,481 1
Valkey cache-aside Zipf None 222,575 0.63 0.999 0 0
Valkey cache-aside Zipf Update 10,000/s 168,911 1.65 0.960 10,000 0
Valkey cache-aside Zipf Update 30,000/s 75,642 4.78 0.887 29,999 0
Valkey cache-aside Zipf Update unlimited 48,993 6.88 0.852 40,735 0
Prepared SQL Uniform None 129,714 0.97 — 0 n/a
Prepared SQL Uniform Update 10,000/s 121,567 1.13 — 10,000 n/a
Prepared SQL Uniform Update 30,000/s 106,949 1.33 — 30,000 n/a
Prepared SQL Uniform Update unlimited 77,574 2.02 — 93,955 n/a
Prepared SQL Zipf None 129,937 0.97 — 0 n/a
Prepared SQL Zipf Update 10,000/s 122,182 1.13 — 10,000 n/a
Prepared SQL Zipf Update 30,000/s 107,852 1.33 — 29,999 n/a
Prepared SQL Zipf Update unlimited 78,601 2.08 — 95,199 n/a
pg_local_cache Uniform Insert unlimited 117,531 1.20 1.000 114,066 0
Valkey cache-aside Uniform Insert unlimited 70,302 4.92 1.000 46,622 0
Prepared SQL Uniform Insert unlimited 84,294 1.98 — 95,094 n/a

pg_local_cache ended every listed mixed run with zero stale entries. Valkey ended one UPDATE case with one stale sentinel entry: Uniform keys at an unlimited update rate. At 30k UPDATE/s on Uniform keys, read rates were 161,919/s for pg_local_cache, 42,294/s for Valkey cache-aside, and 106,949/s for prepared SQL.

Raw runs: mixed reads and writes.

120-second stale probes

Each probe ran once with 64 readers and 64 unlimited UPDATE writers, then performed a full stale-entry check. Valkey used io-threads=8.

Stack Key distribution Reads/s Read p99 ms Writes/s Stale entries Note
pg_local_cache Uniform 82,565 1.687552 108,580 0 —
pg_local_cache Zipf 0 n/a 101,797 0 See note below.
Valkey cache-aside Uniform 30,210 8.781824 36,508 2 —
Valkey cache-aside Zipf 28,568 8.781824 35,789 1 —

The pg_local_cache Zipf reader failed its startup equivalence check while concurrent writes were active (harness issue), so it reported zero reads. Its stale check still ran and found zero stale entries.

Raw runs: 120-second stale probes.

One-hour soak

Single run: MGET 16 over 100k random keys, 64 clients, plus 20,000 UPDATE/s. RESP worker RssAnon was sampled every minute.

  • Reads: 251,987,814 MGET over 3,600 seconds (69,997/s), p99 3.57 ms.
  • Writes: 71,999,993 UPDATE (20,000/s), p99 1.13 ms, zero errors.
  • Worker RssAnon stayed flat: maximum per-worker growth was 44 kB across 60 samples.

Raw runs: one-hour soak.

Stand and method

Each VM reported 16 vCPU arranged as 2 sockets × 4 cores per socket × 2 threads per core, 62 GB RAM, and Debian 13 with kernel 6.12.111+deb13-amd64; this is the VM-visible topology. The DB VM ran PostgreSQL 16.15 (max_connections=1000, shared_buffers=1048576) and pg_local_cache 3.1.0 c431bcc (enabled, memory_budget_mb=1024, workers=8). Its table held 100,000 JSON rows averaging 166 bytes. Valkey 8.1.1 used io-threads=8 during write-overhead, mixed, and stale-probe runs. From the load VM to the DB VM, iperf3 measured 10.4 Gbit/s and ping reported min/avg/max RTT of 0.090/0.143/0.871 ms. The Go/pgx load client ran on the second VM. Reads used 20-second runs; write overhead used 15 seconds, mixed reads/writes 30 seconds, stale probes 120 seconds, and the soak 3,600 seconds.

The pinned matrix used 256 clients; every case had five runs except the two one-core Valkey io-threads=1 cases, which had two. The all-core matrix and write-overhead cases had three runs each; mixed, stale-probe, and soak tests each ran once. NIC interrupts were not pinned. The 64-key results are client/network-bound. Every table uses final build c431bcc. Throughput is requests/s unless the table says transactions/s; latency is p99, CPU is server vCPU.

Environment and raw JSONL: environment, stand evidence, pinned reads, Valkey io-threads=1 pinned reads, all-core reads, all-core Valkey io-threads=1 reads, write overhead, reads during writes, 120-second stale probes, and one-hour soak.

Reproduce with bench/

The scripts need Python 3 and SSH aliases for the database and load VMs. The DB VM needs PostgreSQL 16, Valkey, mpstat, sar, ip, and systemd; the load VM needs the compiled Go client in /root/bench-client. See bench/README.md and the Go/pgx client.

Run the default 20-second read matrix (three repeats, 1/16/64 keys, 16/64/256 clients, hot and 100k-key spaces):

python3 bench/run_bench.py reads.jsonl

For a pinned one-vCPU-pair, one-key comparison with 256 clients and five repeats, replace 0-1 with the DB VM logical CPU IDs for one vCPU pair:

PGLC_SERVER_CPUS=0-1 REPEATS=5 BATCHES=1 CLIENTS=256 KEY_SPACES=0,100000 python3 bench/run_bench.py pinned.jsonl

The runner applies AllowedCPUs to both PostgreSQL and Valkey, then clears it afterward. It records server CPU, latency, network RX/TX, and client CPU. Configure Valkey’s io-threads on the DB VM before the run to match the configuration being reproduced.

Run write-overhead and 30-second mixed workloads:

REPEATS=3 WRITE_SECONDS=15 WRITE_CLIENTS=32,64 python3 bench/write_bench.py writes.jsonl
REPEATS=1 MIXED_SECONDS=30 MIXED_RATES=none,10000,30000,unlimited MIXED_KEY_DISTS=uniform,zipf python3 bench/write_bench.py mixed.jsonl

The write runner records database CPU per transaction; mixed rows also record read latency, hit/miss counts, writer rate, and the end-of-run stale-entry check. The scripts append JSONL rows and preserve completed samples on failure.

Historical 2.x laptop results

These older JSON files are historical Apple M3 Max laptop measurements using PostgreSQL 16; their environment and client topology differ from the 3.1.0 two-VM stand above. They include the 2.x SQL mget API, removed in 3.0.0. Use the 2.0.4 documentation for that API; current cached reads use RESP MGET.