Contents

Changelog

All notable changes to pg_turbovec are documented in this file. The format follows Keep a Changelog and the project adheres to Semantic Versioning.

[2.11.0] — 2026-10-06

MINOR: adopt upstream turbovec 1.1.1 (staged 2/4-bit search), plus fixes for two long-standing bugs found by this release’s soak test: one silently corrupted index entries under concurrent writes, the other leaked memory per scan (see Fixed; both are in v2.10.3 and earlier). No SQL surface change, no GUC change, no wire-format change (MetaPageData::version stays 8), index bytes unchanged. ALTER EXTENSION pg_turbovec UPDATE plus a restart is sufficient; no REINDEX.

Minor rather than patch for one reason: on hosts where the new search engages, the candidate set an index scan returns can differ slightly from 2.10.3’s (scores are exact; recall@10 measured unchanged). Details below.

Changed

  • turbovec 1.0.0 → 1.1.1. For an index of ≥ 32,768 rows on aarch64 (dotprod: Graviton2+, Ampere, Apple) or x86 with AVX-512 VBMI+VNNI (Ice Lake server+, Zen 4+), turbovec now keeps its in-memory search cache as separate bit planes and searches in stages: a sign-plane pass for a shortlist, ranking on the lower planes, then an exact rescore. On AVX2-only x86 (and below 32,768 rows) the scan is unchanged.

    Measured on Graviton4 (c8gd.8xlarge, Debian 13, PG 16.15 non-assert), 1M × 1024-d real Cohere embeddings, flat index, warm, whole-query EXPLAIN ANALYZE p50, 3 alternated A/B rounds on one index:

    search_k=32 100 256 1024
    4-bit: 2.10.3 → 2.11.0 7.88 → 6.72 ms (1.17×) 10.88 → 9.52 (1.14×) 18.32 → 15.96 (1.15×) 62.17 → 54.57 (1.14×)
    2-bit: 2.10.3 → 2.11.0 7.22 → 6.86 (1.05×) 10.45 → 9.78 (1.07×) 17.81 → 16.51 (1.08×) 61.67 → 55.11 (1.12×)

    Recall@10 unchanged in every cell (1.000, or 0.993 at 2-bit k=32 for both). p95 improves by the same margin.

    The turbovec kernel itself is 3–6.5× faster on this host (pure kernel, same corpus: 4-bit single-query k=10 1.84 → 0.62 ms at 32 threads, 22.7 → 6.6 ms at one thread; k=1024 8.29 → 1.28 ms). End-to-end gains are smaller because the kernel is now a minor share of a flat-index query: the saving in milliseconds is the same, but the rest — the per-candidate heap fetch and exact recheck, ~48 µs per candidate — is unchanged. For flat indexes the recheck, not the scan, is now the bottleneck.

  • Cold-backend latency kept (and slightly improved) via a new fork carry. Stock 1.1.1 builds the planes cache with a serial planes_repack, bypassing the parallel cold-open repack that gave v2.10.3 its 3.1× cold-scan cut — it takes 279 ms (4-bit) / 1,092 ms (2-bit) at 1M × 1024-d. Fork carry #4 parallelizes it (18–21 ms / 32 ms), byte-identical to the serial body (parallel_planes_repack_is_byte_identical_to_serial, x86_64 + aarch64). Cold-backend p50 on Graviton4, 1M × 1024-d: 4-bit 513.7 → 476.9 ms, 2-bit 327.0 → 319.8 ms.

Fixed

  • Silent index-entry corruption under sustained concurrent writes + VACUUM (present since v1.29.1). The deferred insert flush tracks which rows a transaction upserted (touched_ids) and splices only those onto the current on-disk index. That list was never cleared after a successful flush, and the backend’s cached index outlives the transaction, so a long-lived writer connection re-spliced every row it had ever written on every later commit, taking their codes from its stale in-memory copy. Once VACUUM had removed such a row and PostgreSQL reused its heap slot for a different row, that re-splice either overwrote the new row’s index entry with the old row’s codes (the row is then searched by the wrong vector, so it can be missed) or put the vacuumed entry back (a stale pointer to a dead or reused tuple). turbovec_check stayed clean throughout (ids remain unique) — only a byte-level comparison against a fresh REINDEX shows it. Measured: after 40 min of 3 writers + VACUUM every 2 min on Graviton4, v2.10.3 had 530 live rows with another row’s codes and 2,369 stale entries in a ~660k-row index (v2.11.0 before the fix: 721 / 2,564). Fixed: touched_ids is emptied once flushed. Repro test flush_does_not_resplice_previous_txn_ids (fails before: a vacuumed id is resurrected; passes after). Re-run A/B, same workload side by side: v2.10.3 corrupted 406 entries and left 2,073 stale; v2.11.0 matched a fresh CREATE INDEX of its heap byte-for-byte on all 658,221 entries (40 min, 12,387 commits, 40 mid-flush kills, 19 VACUUMs). Details: benches/results/tv111_arm_20261005/FINDINGS.md §6.

    Affected: flat/IVF TurboQuant (2/¾-bit) indexes that take writes from connections living across many transactions while VACUUM runs. Short-lived connections (one transaction per connection) were not affected. Recovery: REINDEX INDEX once after upgrading an index that has taken such writes — the bug corrupted entries already on disk, and upgrading stops new damage but does not repair old. 1-bit (bit_width = 1) and graph indexes use other write paths and are not affected.

  • Per-scan memory leak that could OOM-kill a backend (and restart the cluster) under concurrent writes. Present since v1.8.0; found by this release’s sustained-insert soak. amendscan was a no-op, so every finished index scan leaked its Rust-side scan state, including a reference to the backend’s cached copy of the whole in-memory index. While nothing changes that costs little, but after another session commits, the next scan replaces the cached copy and the leaked references keep every superseded copy alive: a long-lived connection (a pooled one, say) that keeps querying an index other sessions keep writing to grows by about one in-memory index per commit it observes. Measured: +200 MB per commit on a 200k × 1024-d 4-bit index, on v2.10.3 and v2.11.0 alike; in the soak a scanning backend reached 52 GB and the OOM killer took it, and with it the postmaster (crash recovery). Fixed: the same workload now holds at 241 → 248 MB over 30 commits. Repro test: amendscan_releases_cached_index_handle (fails before the fix with 6 live references after 5 scans; passes with 1). A scan that aborts with an ERROR still skips amendscan, so it leaks once; that is bounded and noted in the code.

    If you run v2.10.x or earlier with long-lived connections and concurrent writes, this alone is a reason to upgrade. Short-lived connections were not affected: the memory is freed when the backend exits.

Behaviour change: approximate candidate set

The staged search returns bit-exact scores but may miss a candidate whose sign bits alone rank it outside the shortlist. On the Cohere corpus above, staged vs whole-index scan returns the identical top-10 id set for 98.5–100% of queries (≥99.6% mean overlap), and the identical top-100 for 73.5–99% (≥99.6% mean overlap — a tail candidate or two swapped). The exact heap recheck re-ranks what turbovec returns, and recall@10 against exact ground truth is unchanged at every measured search_k. To restore the whole-index scan, set TURBOVEC_4BIT_PLANES=0 and/or TURBOVEC_2BIT_PLANES=0 in the postmaster’s environment (read once per process).

Safety

The new risk is that after an INSERT/DELETE, turbovec reconstructs the packed codes we persist from the planes cache. Gates:

  • Byte-identical persisted index: the same 1M × 1024-d heap indexed by 2.10.3 and by 2.11.0 on Graviton4 has identical meta + codes/scales/ids chains (sha256).
  • New #[pg_test] persist_is_byte_exact_through_planes_layout: 40k rows (past the planes gate), real remove/add + PreCommit flush, every row’s in-memory and persisted bytes checked at 2 and 4 bits. Confirmed to take the planes layout on Graviton4.
  • cargo pgrx test pg16: 448 passed / 0 failed / 8 ignored on x86_64 (446 on aarch64 Graviton4 before the two fix tests were added).
  • Sustained-insert soaks on Graviton4 with mid-flush pg_terminate_backend, periodic VACUUM and UPDATE churn, then a byte-level comparison of every persisted entry against a fresh CREATE INDEX of the same heap. These found both bugs under Fixed; the A/B re-run against v2.10.3 is in benches/results/tv111_arm_20261005/FINDINGS.md §6.

turbovec’s own suite: 523 passed / 0 failed on Graviton4.

Fork

Pinned to gburd/turbovec@pgtv-2.11.0-port (455549f): upstream 1.1.1 + carries

1–#3 cherry-picked unchanged (pub repack, IdMapIndex parts API, parallel

repack; upstream issues #545–#547) + new carry #4 (parallel planes repack).

Not measured

x86 AVX-512 VBMI+VNNI hosts (no measurement here; upstream reports similar kernel gains); IVF and ColBERT latency; indexes under 32,768 rows (unchanged by construction).

Build note

turbovec 1.1 needs LLVM ≥ 22 for its AVX-512 VNNI intrinsics; the nix-packaged rustc 1.97.0 (LLVM 21) fails with intrinsic signature mismatch. Any rustup toolchain ≥ 1.89 works (CI uses rustup stable). On aarch64 Linux, cargo pgrx test (debug profile) fails to assemble gemm-common’s fp16 inline asm without RUSTFLAGS="-C target-feature=+fp16"; the --release build (cargo pgrx install --release, what ships) compiled cleanly without it. This is pre-existing (gemm 0.18.2, unchanged by this release).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.11.0'; and restart PostgreSQL. The format is unchanged, so no REINDEX is required, but REINDEX any 2/¾-bit flat or IVF index that has taken writes from long-lived connections while VACUUM was running to clear entries the touched_ids bug (above) already corrupted. Evidence: benches/results/tv111_arm_20261005/.

[2.10.3] — 2026-09-25

PATCH: cold-scan latency. No SQL surface change, no GUC change, no wire-format change (MetaPageData::version stays 8), index bytes unchanged. ALTER EXTENSION pg_turbovec UPDATE is sufficient; no REINDEX.

Changed

  • Cold-scan latency cut ~3× by parallelizing the per-backend blocked-layout rebuild. Every cold backend rebuilds the SIMD-blocked code layout from the row-major packed codes at index-open (v7+ persists only the codes, halving the on-disk footprint). That repack was single-threaded and was the dominant term in cold latency. It now runs in parallel across block-aligned ranges.

    Measured A/B on identical hardware/corpus/index (c7i.4xlarge, 16 vCPU, AVX-512; 1M × 1024-d 4-bit flat, 534 MB): cold-backend p50 1766 ms → 566 ms (3.1×); warm-backend p50 unchanged (30.4 → 30.6 ms, within noise — a warm backend never repays the repack). See benches/results/rebench_20260925/COLDSCAN_FINDINGS.md.

    The parallel repack (turbovec fork carry #3, rev 47a26a3) produces byte-identical output to the serial version, pinned by turbovec’s parallel_repack_is_byte_identical_to_serial (bit-widths 2/¾, sub/above the parallel threshold, tail-padding shapes). This is a speed change only; the persisted format is untouched.

Docs

  • RETRACTED the “we LOSE ~490×” latency scoreboard in docs/PARITY_GAPS.md. A corrected end-to-end benchmark (top-level EXPLAIN(ANALYZE) Execution Time, literal query vectors — a query-vector subquery in the ORDER BY had added ~90 ms of InitPlan overhead to BOTH engines and produced the bogus 2552 ms figure; one warm psql session per arm) on 1M × 1024-d Cohere-wiki shows flat-bw4 at 5.2 ms / R@10 = 1.000 BEATS pgvector HNSW at R@10 ≥ 0.95 (HNSW 8.6 ms) and is 3× faster at ≥ 0.98. IVF-bw1 is within 2.4–2.8× of HNSW at 58× smaller storage; IVF-bw4 is the weakest arm (confirms flat > IVF for bit_width ≥ 2 at this scale). Full iso-recall table + corrected harness in benches/results/rebench_20260925/.
  • Documented turbovec.ivf_max_delta_pct + “fewer, larger transactions” guidance for continuous high-ingest into an already-large IVF index (docs/PARITY_GAPS.md), and recorded the sparse-ANN and bulk-INSERT design memos (benches/results/parity_20260925/).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.10.3'; — that’s all. No REINDEX, no downtime, index bytes unchanged.

[2.10.2] — 2026-09-24

PATCH: documentation plus one build-time NOTICE. No SQL surface change, no GUC change, no wire-format change (MetaPageData::version stays 8), index bytes unchanged. No REINDEX.

Added

  • A flat build over 100 k rows now emits a NOTICE naming WITH (lists = N). lists defaults to 0, so a plain CREATE INDEX ... USING turbovec builds a flat exact scan — and nothing told the user that an approximate, cell-pruned IVF layer was one reloption away. An evaluator concluded from exactly that experience that pg_turbovec “does not support ANN” and chose a different extension. That is a discoverability failure on our side, not a misreading.

    Deliberately a NOTICE, not a WARNING: flat is frequently the better choice, so it must not read as a fault to fix. Suppressed for IVF, graph and ColBERT builds (already non-flat) and below 100 k rows, where flat is unambiguously right and the message would be noise.

  • README: “Common objections, answered with measurements”. Four things evaluators say, each answered from our own benchmarks — including the two where the objection is correct:

    • “doesn’t support ANN” — it does; the default is exact, which is why it looked absent.
    • “HNSW has a high memory footprint and is slow to build” — agreed, which is why we didn’t build on it. Measured 10 M × 1536-d: our index is 4.4× smaller (14.9 vs 65.5 GiB) and builds 2.6× faster (1 h 24 m vs 3 h 38 m), and our own log shows HNSW slowing super-linearly past 5 M rows.
    • “pg_turbovec’s own build memory was worse than HNSW’s” — true, and now fixed. That benchmark measured us at 121 GiB peak + 60 GiB swap against HNSW’s 16.9 GiB. Root cause fixed in v2.10.1; measured after: 2 M × 1024-d peak 12.16 → 3.45 GiB, and 10 M × 1024-d went from OOM-killing a 61 GiB host to completing at 11.20 GiB. Scope stated plainly in the README: the post-fix numbers are 1024-d, and 10 M × 1536-d (the dimension the 121 GiB figure used) has not been re-measured.
    • “HNSW is faster on query latency” — true, and we don’t dispute it.
  • README: “Choosing lists — the ANN tuning knob”. A full tuning guide, because the honest answer is more subtle than “turn on ANN”:

    • Whether you want IVF at all. At 1 M × 1024-d with the default bit_width = 4, flat is 6.08 ms at recall@10 = 1.000 while lists = 1024 is 62 % slower and capped at recall 0.959 — a per-probe ceiling that is CPU-independent and that no amount of tuning removes. IVF is a trade, not a free speedup.
    • The measured decision table (by bit_width, scale, recall target, and whether the index exceeds RAM).
    • √n is a ceiling, not a target — at 1 M, lists = 4096 measured worse than 1024 on every axis: 11× the build time and ~50 % higher latency.
    • Tune probes at query time, not lists (which is baked in at build).
    • Measure on your own data, at matched recall.

Unchanged, deliberately

  • lists still defaults to 0 (flat). Changing the default to √n was considered and rejected on our own measurements: at the default bit_width = 4 it would ship a configuration that is 62 % slower and recall-capped at 0.959 up to at least 1 M rows. The defect was discoverability, not the default value.

Tests

445 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane. New: flat_build_hints_at_the_ann_option, which also pins the silence (IVF builds and sub-threshold indexes must not be nagged).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.10.2'; — nothing else. No REINDEX.

[2.10.1] — 2026-09-23

PATCH: build-time memory profile only. No SQL surface change, no GUC change, no wire-format change (MetaPageData::version stays 8), and the on-disk index bytes are identical — guarded by the IVF byte-identity tests. No REINDEX.

Fixed

  • Phase Z6 — CREATE INDEX / REINDEX peak memory cut ~3.5×. The build callback had no memory-context management at all: no switch, no reset, no pfree. Vector is a PostgresType stored as CBOR, so FromDatum::from_datum palloc’s a decoded buffer for every row — and those buffers accumulated in the long-lived ambuild context for the entire scan. The CorpusSpill was faithfully streaming the corpus to disk while PostgreSQL held a decoded copy of all of it in RAM, defeating the spill’s whole purpose.

    build_callback is now a thin wrapper that switches into a per-tuple context, calls the unchanged inner callback, restores, and resets. The inner function has six early-return paths, so doing this inline would have been six chances to leak the switch.

    One hazard came with it: BufFileCreateTemp palloc’s in the current context and the spill is opened lazily on the first row — inside the context now being reset, which would have left a dangling BufFile on row two. CorpusSpill::new_in(cxt, dim) plus an explicit BuildState.build_cxt keep it in the long-lived context by construction at all three lazy-open sites (single-vector, BQ/graph, ColBERT).

    Measured on 2M × 1024-d, lists = 1414 (benches/results/z6_buildmem_20260922/):

    before after
    at drain entry 12.07 GiB 2.39 GiB (5.05× less)
    whole-build peak 12.16 GiB 3.45 GiB (3.5× less)
    build wall time 1055 s 889 s (16 % faster)

    This is what made 10M × 1024-d builds OOM-kill a 61 GiB host. Since measured (benches/results/z6_10m_20260924/): on the same c7i.8xlarge / 61 GiB host with the identical config that died before (mwm = 8GB, 16 parallel workers, lists = 3162), the 10M × 1024-d build now completes in 69.8 min at 11.20 GiB peak private — 18 % of the host, zero OOM events, against anon-rss 58.47 GiB at the kill. The projection above was 11.2 GiB. Index verified sound (wire v8, 10M/10M slots, is_corrupt = false, scan_fraction = 0.00506, 51 ms warm).

  • The Lloyd cross matrix was quadratic in lists. train_kmeans allocated n_sample × lists where n_sample = lists × 256 — 9.54 GiB at lists = 3162. Now chunked over sample rows under a fixed ~256 MiB budget. Bit-identity is tested, not assumed: kmeans_cross_chunking_is_bit_identical compares the whole-sample call against chunk sizes 1/7/64/199/500.

Changed

  • turbovec.build_parallelism is a memory knob, and its description said the opposite (“only build wall-clock changes”). Measured: 16 threads → 16.24 GiB peak private, 1 thread → 12.16 GiB — ≈ 0.27 GiB/thread of thread-local GEMM packing buffers, which maintenance_work_mem does not bound. Corrected, with the advice to lower it when CREATE INDEX is memory-constrained.
  • trace_stage! (under TURBOVEC_BUILD_TRACE) now reports private memory per stage, plus a new 0_at_drain_entry marker. Peak had been misattributed for three sessions because the only signal was a process-wide total — which also includes shared_buffers. That single marker localised this bug in one build.

Documentation

  • A constraint is now pinned by test: rotate_corpus_into is not invariant to row-block shape (Parallelism::Rayon(0) makes the GEMM’s reduction order depend on m), so the reservoir rotation cannot be chunked to save memory. A blocked-rotation optimisation was written and CI caught it before it could change index bytes; rotate_corpus_is_not_row_block_shape_invariant records why. The pre-existing rotate_corpus_bit_identical_across_pool_sizes does not cover this — it varies thread count at a fixed shape.

Tests

444 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.10.1'; — nothing else. No REINDEX.

[2.10.0] — 2026-09-22

MINOR: one new GUC and a changed IVF insert/scan behaviour. No SQL-surface change and no wire-format change (MetaPageData::version stays 8) — existing indexes decode byte-identically and no REINDEX is required.

Changed

  • Phase Z5 Route A — an IVF index no longer loses its cell layout on the first INSERT. aminsert cannot place a row into its cell without an O(n) reshuffle, so appended rows land at the tail, outside every cell. Previously the deferred-commit flush dropped the coarse-centroid and cell-directory chains outright, so one commit turned the index into an O(n) flat scan until REINDEX — an operational cliff with no cheap recovery. The flush now writes those chains back unchanged, and the scan additionally sweeps the tail exhaustively. Results stay exact; only latency is affected.

    No new meta field was needed. The delta length is derivable as n_live - cell_directory.total_vectors(), because existing rows keep their slot (an UPDATE writes in place) and inserts append at the tail. An earlier assessment in docs/PARITY_GAPS.md called this “blocked on a wire-format change”; that was wrong, and the correction is recorded there.

    Bounded by the new turbovec.ivf_max_delta_pct (default 10, range 0..=100): past the bound the index degrades to flat and reports it exactly as before, so the tail cannot grow until the optimisation is an O(n) scan with extra bookkeeping. Setting it to 0 restores pre-2.10.0 behaviour byte-for-byte.

Fixed

  • Out-of-core path could make appended rows unreachable. The OOC gather iterates only probed cells, so a tail outside every cell was never read — those rows existed on disk and could never be returned. That is silent loss, not slowness. The delta is now fed through the same tombstone-aware run-splitting as cells, so a tombstoned appended row cannot resurrect (the v2.7.0 class of bug).

Documentation

  • shared_preload_libraries = 'pg_turbovec' is now documented as required (README + first section of PRODUCTION.md). Every turbovec.* GUC is registered by _PG_init, i.e. at library load; without preloading a backend has zero of them, SET turbovec.probes = 16 is accepted and silently ignored, and every query runs at compiled-in defaults. Indexes still build and queries still return correct results, which is exactly why it misdiagnoses as “tuning has no effect” or “IVF pruning doesn’t work” — it cost a full benchmark round before being spotted. Includes the one-line check (SELECT count(*) FROM pg_settings WHERE name LIKE 'turbovec.%') and the LOAD / session_preload_libraries alternative for managed providers (verified: every GUC is Userset, with no Postmaster/Sighup context).

Measured (EC2, benches/results/z5_delta_20260922/)

c7i.4xlarge (16 vCPU, AVX-512), PostgreSQL 16.15, 1M × 256-d, bit_width=4, lists=1024, 1000 inserts, 30 warm queries per arm in one session, fresh index per arm, autovacuum disabled, index health verified after the inserts:

arm in-memory out-of-core
pre-Z5 (pct=0) → degraded, scan_fraction 1.0 4.51 ms 4.43 ms
Z5 delta (pct=10) → healthy, scan_fraction 0.0156 4.03 ms 3.46 ms

The win is 11 % in memory and 22 % out-of-core — not the 64× that was modelled. The model assumed latency scales with rows scanned, but at 1M × 256-d a full 4-bit scan is only 1.6× a 1-cell scan (5.97 vs 3.66 ms): the 139 MB index is RAM-resident and per-query fixed costs are the same order as the SIMD sweep. This agrees with our own published 1M bw4 result (flat 6.08 ms beats IVF 16.04 ms; v2.8.3 “flat wins at every target”). This release ships on the functional contract — no cliff on insert, plus the OOC correctness fix — not on the latency delta. The >RAM regime, where a full scan is disk I/O and the win could be materially larger, is explicitly unmeasured.

Corruption validation (HARD MANDATE)

6 concurrent writers + 4 readers + VACUUM every 30 s for 5 minutes, with autovacuum enabled: is_corrupt = false, no duplicate id, n_vectors == slot_count exactly, wire still v8, and the cell layout survived (degraded = false, scan_fraction = 0.0156). Two apparent discrepancies were investigated rather than assumed: the index holding 1072 more rows than the heap (lazy vacuum reclaim of 10072 dead tuples) and a 1000-row sample returning 254 (turbovec.search_k = 32 caps candidates).

Tests

442 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane. New: the delta-bound predicate (accept/reject, 0 disables, inconsistent layouts), plus two integration tests asserting both halves — cells preserved and every appended slot swept even when a single cell is probed, while unprobed cells stay excluded (otherwise the pruning is gone and it is a flat scan in disguise).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.10.0'; — nothing else. No REINDEX. To keep the previous behaviour exactly: SET turbovec.ivf_max_delta_pct = 0.

[2.9.0] — 2026-09-22

MINOR: adds one SQL function. Wire format unchanged from 2.8.x (MetaPageData::version stays 8), so existing indexes decode byte-identically and no REINDEX is required — the upgrade is in place.

Added

  • turbovec.index_degradation(regclass) — quantifies an IVF degradation instead of merely flagging it. Returns degraded, lists, n_vectors, scan_fraction, est_slowdown and a recovery string. Phase Z1 made degradation observable and Phase Z4 made the planner cost it correctly; neither told an operator the size of the problem, which is what decides whether to act — a degraded 10k-row index is a non-event, a degraded 10M-row index is an outage. A degraded index reports scan_fraction = 1.0 (it reads everything) and est_slowdown = lists / probes, plus the REINDEX command naming the index. A flat index is explicitly not reported as degraded — it scans everything by design, and faulting it would train operators to ignore the signal. Reads only the meta page (one buffer hit), so it is safe to poll from monitoring.

Changed

  • Phase Z4 — amcostestimate is now probe- and filter-aware. Three defects, all of which made the planner blind to what the AM does:
    • IVF was costed as a full-corpus scan although the scan clamps to turbovec.probes cells, so an index probing 1 of 1024 cells was costed identically to a flat scan of everything. Cost now scales with probes / lists, floored at one cell’s worth. A degraded IVF index is costed as flat, since that is the path it takes.
    • index_selectivity was hardcoded to 0.0 for every query. It now derives from the planner’s own rel->rows / rel->tuples, so we agree with the planner by construction instead of second-guessing it.
    • A pre-existing unit error: the ns→cost conversion divided seconds by cpu_operator_cost, making a 1M × 1024-d flat scan cost ~23 while PostgreSQL costs the equivalent sequential scan at ~73,000 — about 3000× too cheap, which let an ANN path beat plans that are genuinely faster. Now ~1688. The ns throughput model itself validated against our own published measurement (model 5.3 ms vs measured 6.08 ms), so only the unit was wrong.

    index_pages is likewise scoped to the pages a probed scan touches. The arithmetic lives in two pure, unit-tested functions because it is not observable through EXPLAIN: the index-scan node’s cost also carries PostgreSQL’s heap-fetch and qual costs, which swamp it.

Documentation

  • docs/FILTERING.md — the allowlist crossover is now actionable. The measured table (2.6–14.7× faster below ~7 % selectivity, 2.6× slower at 100 %) was only useful if you knew your filter’s selectivity. Adds a copy-pasteable way to read PostgreSQL’s own estimate, a fraction→technique table, and two caveats: the crossover percentage is host-dependent (the shape transfers, not the number), and the estimate is only as good as your statistics.
  • Phase Z3 rescoped to a non-gap, with the reasoning recorded. “Automatically turn a WHERE into a kernel mask” is not implementable by a PostgreSQL AM: a scan key is index_key operator constant over an index column, so a qual on any other column becomes an executor Filter and never reaches the AM. amgetbitmap is not a route either — it returns an unordered bitmap, and ordering is the whole value of an ANN scan. The useful behaviour shipped in v1.8.0 as iterative scan, which is demand-driven and needs no view of the filter.
  • Phase Z5 scoped, with both routes shown blocked. A bounded mutable delta needs a wire-format change (both scan paths assert directory.total_vectors() == n_live and centroids.len() == lists * dim; a synthetic “delta cell” would need a centroid it does not have). Cheap in-place cell reassignment needs a decode/reconstruct that turbovec does not expose (pack.rs has only repack/unblock). The reporting half shipped instead.

Tests

435 → 437 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.9.0'; — creates the new function. No REINDEX.

[2.8.4] — 2026-09-21

Code-only release. Wire format unchanged from 2.8.3 (MetaPageData::version stays 8); no REINDEX needed and no SQL surface change.

Fixed

  • Phase Z1 — an IVF index that takes writes now REPORTS its degradation (ordinary TurboQuant path). An aminsert cannot place a row in its cell without an O(n) reshuffle, so the deferred-commit flush appends and the index falls back to a flat scan. The 1-bit BQ path always preserved lists and stamped ivf_degraded so turbovec.index_is_degraded() and the throttled ambeginscan WARNING fired; the TurboQuant path blanked lists, so index_was_ivf() went false and the latency cliff was silent — an operator got a quietly slower index with no signal and nothing to act on. Root cause: reconcile_and_write_flush planned its meta page via plan_with_blocked, which hardcodes lists: 0. It now captures the on-disk lists under the already-held exclusive rewrite lock and stamps both fields, leaving the coarse/cell-directory offsets at zero so the readers return empty and the scan takes the flat fallback deterministically rather than by a length coincidence. Both fields are existing v4 meta scalars — no chain is added or moved, so the chain-offset running-sum class is not implicated.

  • An INSERT into an index built WITH (assign_dups > 1) no longer claims the index is corrupt. IVF-4a soft assignment stores a boundary row in several cells on purpose, so its external id appears in several slots and slot_to_id is deliberately not a bijection — while the insert path loads the index into a flat IdMapIndex, which requires one. The rejection reported corrupt relfile pages: duplicate ids with a REINDEX hint; both halves were wrong (turbovec_check verifies such an index clean, and a rebuild reproduces the same by-design duplicates), sending operators hunting for corruption that does not exist. It now reports FEATURE_NOT_SUPPORTED, names assign_dups, states the index is effectively READ-ONLY, and the HINT says how to confirm it is healthy. Behaviour is otherwise unchanged: the INSERT still fails, SELECT still works. Zero cost on the healthy path — from_id_map_parts has exactly one failure mode, so reaching the handler already identifies the cause.

Documentation

  • docs/PARITY_GAPS.md gains a zvec source review (local alibaba/zvec checkout d88357b vs deac2d9), annotated throughout as un-benchmarked: four real gaps in priority order, the one lesson worth stealing (bounded mutable delta + explicit consolidation), explicit non-gaps (hybrid fusion, ColBERT, WAL, durability, scalar filtering — PostgreSQL’s or already ours), and deliberately deferred items (RaBitQ/PQ, DiskANN). Adds phases Z1–Z5; Z3 (automatic predicate→ANN handoff) is gated on Z4 (costing), because pushing filters while index_selectivity is hardcoded to 0.0 would only make bad plans confident.
  • assign_dups > 1 is now documented as making the index read-only, in both docs/BENCHMARKS.md (which recommended it for recall without saying so) and the reloption reference in src/index/options.rs.
  • AGENTS.md gains four steering rules: degradation must be observable; competitor comparisons stay source reviews until measured; the BQ and TurboQuant insert paths write at different times (BQ synchronously in aminsert, TurboQuant deferred to PreCommit, which a #[pg_test] never reaches — so a plain INSERT in a test exercises nothing on that path); and assign_dups > 1 indexes are read-only.

Tests

429 → 430 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane. New: ivf_flush_degradation_is_reportable, ivf_soft_assign_index_rejects_insert_and_is_not_corrupt, ivf_soft_assign_insert_error_names_assign_dups.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.8.4'; — nothing else. No REINDEX.

[2.8.3] — 2026-09-11

bit_width = 4 + IVF measured at 1M — flat wins at every target. This answers the production user’s question with data rather than the forecast v2.8.2 shipped. Documentation only; the binary is byte-identical to 2.8.2. No wire change (v8), no SQL surface change, no REINDEX.

The measurement

1M × 1024-d real Cohere corpus, AVX-512, 48 configs, every one confirmed to run a real Index Scan. Artefacts: benches/results/bq_1m_bw4ivf_20260911/.

target bw4 flat bw4 lists = 1024
R@10 ≥ 0.90 6.08 ms (w=32) 16.04 ms (p=64) flat by 62 %
R@10 ≥ 0.95 6.08 ms (w=32) 16.13 ms (p=128) flat by 62 %
R@10 ≥ 0.98 6.08 ms (w=32) unreachable flat only
R@10 ≥ 0.99 6.08 ms (w=32) unreachable flat only

The write-up leads with the recall ceiling, not the latencies, because the conclusion does not depend on any timing. At probes = 128, widening the rerank window from 32 to 2000 leaves recall at exactly 0.959 across all 8 windows — it does not move by a single query, because the true neighbours are not in the probed cells. Recall is CPU-independent, so discard every latency number and 4-bit IVF still loses at the top two targets. Flat’s cheapest config is also its most accurate (R@10 = 1.000 at window 32), so IVF never gets an opening. The 62 % gap is corroboration.

Published for completeness rather than left for a reader to find: the one sub-flat p50 in 48 rows is at window 2000, IVF 116.77 ms versus flat 122.19 ms (0.96×). It is not a win — 0.959 recall against 1.000 — so it fails a matched-recall comparison. The tightest honest framing: IVF’s fastest configuration anywhere is 15.76 ms at R@10 = 0.875, against flat’s 6.08 ms at R@10 = 1.000.

The OOM worry is retired for this configuration

Possibly more actionable for an operator than the latency result. At 1M × 1024-d with maintenance_work_mem = '4GB' and 4 parallel maintenance workers:

arm bytes/vec index build peak build anon-RSS
bw4 flat 559.95 534 MB 14.89 s 2.641 GiB
bw4 lists = 1024 564.17 538 MB 101.22 s (6.8×) 2.154 GiB

Real peaks, not lower bounds — 0.25 s sampling, 466 in-build samples, from a pure-bash sampler that left idle loadavg at 0.00–0.01. IVF’s peak is below flat’s, and 2.6 GiB is nowhere near the 20.3 GB an unbounded 3GB setting reached on the earlier 250k build. The hazard is leaving maintenance_work_mem unbounded, not 4-bit IVF. The genuine cost is the 6.8× build time.

Caveats, recorded rather than glossed

  • All 48 latency rows are contention-flagged, and the unflagged-row filter is unavailable here: 0 of 8 flat and 0 of 40 IVF rows survive it — exactly the trap § 0.6e documents. The bias runs against IVF (flat ran at 22.0 % mean cpu_busy versus IVF’s 6.1 %, a 3.64× difference; mean loadavg 1.87×) and flat still won by 62 %, so a quiet re-time could only widen flat’s margin.
  • A second corpus-identity surprise. This arm’s resolvability spread was 18.6–69.1 % against § 0.6e’s 154–203 % on nominally the same corpus and shards. Both clear the ~10 % unusable floor so each run’s internal comparison stands, but the discrepancy is unexplained and the two runs must not be treated as same-corpus. The first such surprise forced the v2.8.1 corrections.
  • Ground truth took 438.4 s against § 0.6f’s ~275 s on the same parallel-CTAS path. Unexplained, flagged so a future diff does not misread it as a harness regression.
  • Scope: a ≥ 0.98 user is answered at any scale, since the per-probe ceiling is structural rather than a size effect. A ≲ 0.90 user well above 1M is not answered by this arm.

Also closed a gap the previous 1M run left open: held-out queries are value-disjoint as well as id-disjoint (value_overlap = 0 via md5 hash join).

Operational: AWS burner rules

AGENTS.md had no AWS section, which is what let a mid-run account expiry stand a benchmark instance up with no way to terminate it (account bene expired at 12:03 UTC, ~48 min after launch; InvalidClientTokenId on a previously-working key means the account went away, not that you broke something). Added the current burner (lava) and the rules that made that incident cost zero data — chiefly pull artefacts as you go.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.8.3'; — no REINDEX.

[2.8.2] — 2026-09-10

4-bit IVF is supported — the “1-bit-only” result was about benefit, not support. Documentation only; the binary is byte-identical to 2.8.1. No wire change (v8), no SQL surface change, no REINDEX.

The report, and the answer

A production user running bit_width = 4 read v2.8.0’s “the 1M IVF+BQ crossover is 1-bit-only” as meaning IVF cannot be combined with 4-bit, and asked whether support could be added.

It already exists and always has. WITH (lists = N) composes with every bit_width — 4-bit IVF is the original IVF path, out-of-core end-to-end since v1.13.0, and 2-bit and 3-bit work too. The only bit_width/kind combination the code rejects is bit_width = 1 with graph = true (src/index/options.rs). Verified by grepping every rejection site: there is no bit_width gate on IVF anywhere in the tree. Nothing to enable, nothing to wait for.

The misreading is this project’s fault, not the user’s — the guidance paragraph sits inside the README’s 1-bit section, so a 4-bit reader lands on it naturally. Corrected in the README, and § 0.6e’s heading is restated as “the crossover EXISTS, and the benefit is 1-bit-only” with a callout naming the misreading so the next reader does not repeat it.

New § 0.6g — a 4-bit user’s three questions, answered separately

They were being conflated:

  1. Is it supported? Yes.
  2. Will it help? Probably not — and this combination was never measured, stated plainly rather than implied. bit_width = 4 + lists = N does not appear in any artefact at either scale; bw4 is only ever swept flat. The mechanism predicts no win: the crossover needs both an expensive O(n) scan and a quantizer lossy enough to demand a wide rerank window, and 4-bit needs only a 32-wide window versus 800 for 1-bit, so its scan is already cheap and cell-restriction mostly adds overhead. 2-bit is the direct evidence — a clean loss at 1M for exactly that reason.
  3. What could it cost? Two things. The per-probe recall ceiling (0.986 at probes = 128; R@10 ≥ 0.99 unreachable at any setting, and widening the rerank window does not recover it, because the true neighbours are not in the probed cells). And build memory — bw4 + lists = 512 at 1024-d reached 20.3 GB anon-RSS and was OOM-killed at maintenance_work_mem = 3GB.

New: “Should you enable IVF?” in docs/PRODUCTION.md

A decision table plus a 20-minute experiment an operator can run on their own data: baseline at their recall target, build an IVF copy with maintenance_work_mem bounded, sweep probes, compare at matched recall (not matched settings — an iso-knob comparison flatters whichever index gets a wider effective window), and verify with EXPLAIN that it really is an Index Scan rather than a sequential-scan fallback dressed up as a latency result.

The reasoning behind pushing users toward their own measurement: since v2.8.1 this project’s own 1M and 250k figures come from different corpora (the older Cohere dataset became gated), so published cross-scale deltas are suggestive rather than measured. A user’s corpus is the only authority for their workload, and a negative result from their data is worth more than a positive one from ours.

A measurement of bw4 + lists = 1024 at 1M is in flight and will replace § 0.6g’s prediction with a number when it lands.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.8.2'; — no REINDEX.

[2.8.1] — 2026-09-10

Corrections to v2.8.0’s benchmark write-up. Documentation only — the binary is byte-identical to 2.8.0. No wire change (v8), no SQL surface change, no REINDEX. All three items came from the 1M benchmark agent’s final report and were verified before being accepted.

The 1M corpus is not the same corpus as the 250k runs

v2.8.0 presented 1M-versus-250k deltas — IVF’s storage overhead “halving”, the per-probe ceilings “landing within a hair” of the 250k figures — as though both scales shared a corpus. They do not. Cohere/wikipedia-22-12-en-embeddings is now gated (confirmed HTTP 401), so the 1M arm used CohereLabs/wikipedia-2023-11-embed-multilingual-v3: same publisher and dimensionality, but a different model and snapshot. The 250k artefacts carry no corpus label at all, corroborating that they came from a different pre-existing table.

Every cross-scale delta is now labelled suggestive, not measured. The within-run flat-versus-IVF comparisons are unaffected — both arms of each run share one corpus, and those are what the conclusions rest on.

A trap in the method used to validate v2.8.0’s headline

The 47 % / 38 % IVF wins were validated by re-checking on contention-unflagged rows only. That is unsound whenever the baseline does not survive the filter — and on the lists = 4096 arm it does not: all 8 bw1 flat rows are flagged (they ran first, while loadavg was still decaying from the k-means build) and zero survive, so a filtered comparison there would “prove” IVF wins against an empty set.

Re-verified: the lists = 1024 arm used for the headline keeps all 8 flat rows unflagged, so that check was valid — but partly by luck of execution order. The rule is now in docs/TESTING.md beside the existing control-arm rule: assert the filtered baseline is non-empty before trusting a filtered comparison.

Restored: the ground-truth-fix documentation (now § 0.6f)

An earlier rewrite of § 0.6b had overwritten it, and a later edit’s assertion masked the loss — found by grepping for the measured numbers and getting nothing. The code fix was never affected. The restored section also records an accidental validation at 1M scale: the run’s two arms straddled the fix, giving 3755.7 s (pre-fix INSERT path) versus 275.4 s (CTAS path) — 13.6×, matching the 12× predicted from the plan shape — and the 16 flat configs shared by both arms reproduce bit-identically across the two GT implementations, independent confirmation the fix changes no measured number.

Also

  • 1M build-memory figures relabelled as lower bounds, not peaks: the sampler polled every 2 s and was stopped before the second arm, so the lists = 4096 build has no RSS measurement at all.
  • The contention cause is named: loadavg decaying after each parallel k-means build (a monotonic decline across consecutive rows) plus an RSS sampler forking a Python interpreter every 2 s. cpu_busy_pct of only 3.1–3.2 % with cpu_steal ≤ 0.01 confirms it was never saturation.
  • All 160 configs across both 1M arms ran a real Index Scan — zero masked sequential scans.
  • Recorded honestly: held-out queries are id-disjoint from the corpus (join count 0), but value-disjointness was not proven — the 100 × 1M text comparison was abandoned as too slow. A duplicate would require the dataset itself to contain duplicate embeddings.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.8.1'; — no REINDEX.

[2.8.0] — 2026-09-10

The 1M IVF+BQ crossover, measured on a real corpus — plus a parallelised ground-truth path and the v2.7.4 latency caveat resolved with data. Minor rather than patch because the harness’s GT path changed shape (same output, different plan) and the operator guidance for WITH (lists = N, bit_width = 1) is now materially different. No index wire-format change (stays v8), no SQL surface change, no REINDEX.

The crossover exists at 1M, and it is 1-bit-only

Corpus: CohereLabs/wikipedia-2023-11-embed-multilingual-v3 (en), 1 000 000 × 1024-d — real, not synthetic — with 100 held-out queries, on an AVX-512 host. The § 0.6d resolvability gate passed at 154–203 % nn1→nn100 spread (the discarded synthetic corpus was 6.6–10.4 %).

target bw1 flat bw1 IVF (lists=1024) verdict
R@10 ≥ 0.90 25.6 ms 13.7 ms (p=64) IVF wins 47 %
R@10 ≥ 0.95 33.2 ms 20.7 ms (p=128) IVF wins 38 %
R@10 ≥ 0.98 39.3 ms 44.2 ms IVF loses 12 %
R@10 ≥ 0.99 56.8 ms unreachable flat only

Re-computed using only contention-unflagged rows: identical 47 % / 38 %, so this is not a load artefact. For bit_width ≥ 2 it is a clean no — flat is 4.8 ms at every target and IVF never gets under 16 ms.

Mechanically the crossover needs both conditions: flat’s O(n) scan grown expensive and a quantizer lossy enough to need a wide rerank window. 1-bit at 1M needs w=100–256 over 1M rows, so restricting to 64–128 cells of ~1000 rows is a real saving; 2-bit needs only w=32, so its scan is already cheap.

Two-axis guidance, not one: - bit_width = 1, n ≳ 1M, target ≲ 0.95 → lists = N - bit_width = 1, target ≳ 0.98 → flat (IVF cannot reach it) - bit_width ≥ 2 → flat, at least to 1M

lists = 4096 is worse than lists = 1024

A useful negative: quadrupling the list count made every axis worse — 8.7 % more storage, 11× the build (1067 s vs 94 s), and ~50 % higher latency at matched recall (21.5 vs 13.7 ms at R@10 ≥ 0.90). Each cell holds 4× fewer rows, so a given recall needs ~4× the probes (p=256 where 1024 lists needed p=64). The lists ≈ sqrt(n) rule is load-bearing; exceeding it is a pure loss for BQ.

Also confirmed: the per-probe recall ceiling is structural, not a small-corpus artefact (0.846/0.904/0.944/0.971/0.986 at probes 8/16/32/64/128 at 1M, within a hair of the 250k figures), and IVF’s storage overhead halves at 1M (+3.1 % vs +6.0 %) as the fixed centroid/cell-directory cost amortises.

Ground truth parallelised — in two steps, the second correcting the first

Both verified to leave GT row-for-row identical (old EXCEPT new = 0, new EXCEPT old = 0 on (qid, hit_id, rk), membership ignoring rank = 0), so no published recall figure moves:

  1. The correlated LATERAL + row_number() OVER (ORDER BY dist) forced WindowAgg → Sort → Gather: every row shipped to the leader for a single-threaded sort. At 250k × 20 queries, Gather (actual rows=250000, loops=20) with a 13.9 MB leader quicksort, 148.2 s. Replaced with a per-query InitPlan constant so each worker top-N sorts its own share: 76.8 s, 1.93×.
  2. My first version of that fix wrapped the SELECT in INSERT INTO ... SELECT, which silently threw the parallelism away again. PostgreSQL generates no parallel plan for a data-writing statement — only CREATE TABLE AS / SELECT INTO / CREATE MATERIALIZED VIEW are exempt. Serial Seq Scan at 7.49 s versus 3.99 s for CTAS; ~36 s vs ~2.9 s per query at 1M, a 12× gap. Now CTASes each query’s top-k into a temp table and inserts those few rows.

The second defect was caught by the 1M benchmark run’s pg_stat_activity (one backend at 99.9 % CPU, zero parallel workers), not by me. Root cause worth naming: I benchmarked a bare SELECT, measured 1.93×, then shipped an INSERT — a different statement with a different plan — and never re-timed. Benchmark the statement you are going to ship, not a proxy for it.

The v2.7.4 contended-latency caveat is resolved with data

The bench host’s load problem was fixed, so the 250k sweep was re-run on the same corpus: 18 of 24 rows now clean (was 0 of 24), recall reproduced exactly, the contended p50s were uniformly 14–16 % pessimistic, and the published matched-recall ratios held to two decimal places (2.70× / 6.13× → 2.75× / 6.09×). The absolute milliseconds in the § 0 tables are left as published — ~15 % conservative — with the correction pointing at the clean artefact. An honest record beats a tidy one.

Also

ivf_streaming_build_temp_file_cleanup asserted on a cluster-wide pg_ls_tmpdir() delta, so a concurrent #[pg_test]’s transient spill file failed it — observed on a docs-only commit, which is definitionally not a regression. It now re-samples and fails only if the growth persists. Same shared-global-state class as the bench-harness collisions fixed in v2.7.4.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.8.0'; — no REINDEX.

[2.7.6] — 2026-09-09

Documentation consistency pass, plus a narrowed root cause and salvaged numbers from the discarded 1M arm. No shippable code change — binary byte-identical to 2.7.5. No wire change (v8), no SQL surface change, no REINDEX.

The BQ docs contradicted themselves

Across the 2.7.3–2.7.5 runs the docs accumulated claims the measurements had already falsified. docs/BQ_RECALL_BENCH.md still opened by stating it “contains no measurements” and pointing at a README row that “currently and correctly reads not yet published” — both untrue since 2.7.3. docs/ONEBIT_BQ.md carried an orphaned fragment, stranded mid-paragraph by an earlier merge, asserting the harness “has NOT been run — no numbers exist yet”. Every doc was swept for claims the measurements contradict; all are now consistent. The runbook sections are retained and relabelled as how to reproduce rather than not yet done.

The discarded 1M root cause, narrowed — and a competing diagnosis ruled out

A plausible alternative was raised: that the generator’s uncorrelated ARRAY(SELECT randn() ...) subquery had been hoisted, making all 200 cluster centres identical — the same v1.24.0 test-harness bug class AGENTS.md warns about. It was tested rather than accepted or dismissed, and both halves matter:

  • The hazard is real. The subquery never references the outer c. Reproduced minimally: uncorrelated gives 1 distinct vector across 5 rows; adding a correlating WHERE c = c gives 5. Worth recording, because the behaviour is plan-dependent.
  • It did not happen here. Measured on the loaded corpus: 200 distinct centres, 5000 distinct vectors per cid, same-cluster distance 0.108 vs cross-cluster 0.990. randn()’s volatility forced per-row evaluation.

So the earlier “distance concentration” diagnosis was right in substance but imprecise about where. Narrowed by measurement: the clustering worked; the tie is inside each cluster. Within one cluster a member’s 100 nearest neighbours span just 7.22 %, because 5000 iid Gaussian points at d = 768 concentrate — pairwise separation norm has relative spread ≈ 1/sqrt(2d) ≈ 2.6 %. Cluster separation is not sufficient for rankability, which is the non-obvious part and the reason the original prompt’s “make it clustered” instruction was not enough.

Salvaged from that arm: valid 1M storage/build numbers

Corpus geometry doesn’t affect these — per-vector codes are dim/8 * bit_width and the flat build is a fixed O(n·dim) pass:

1M × 768-d build bytes/vec peak build RssAnon
bw1 167.3 s 104.87 8.98 GiB
bw2 152.0 s 209.72 5.95 GiB
bw4 161.0 s 404.76 6.18 GiB

bw2/bw1 = 2.000 exactly, confirming the dim/8 sign-code stride holds at 1M rows. Note bw1 has the highest peak build RSS despite the smallest output: BQ reads the corpus back resident to compute the corpus mean before it can take signs, so its build memory tracks n · dim · 4 regardless of bit width. Worth knowing before building a 1-bit index on a RAM-constrained host.

Harness defect: the ground-truth query serialises on the leader

The GT query puts row_number() OVER (ORDER BY <distance>) in a Sort above the Gather, so parallel workers ship raw vectors to the leader and the leader recomputes every distance single-threaded. That is why ground truth took 23 084 s (6.4 h) for 100 queries at 1M rows — a plan pathology in the harness, not a turbovec cost, and the main thing making a real 1M arm expensive.

New: mandatory pre-flight for synthetic corpora

docs/BQ_RECALL_BENCH.md § 0.6d. Before trusting any recall number from generated data, run the resolvability probe and require the nn1→nn100 spread to be comparable to a real corpus:

corpus spread verdict
real Cohere-wiki 1024-d 37–268 % rankable
the discarded synthetic 768-d 6.6–10.4 % unusable

With an explicit warning that is_degenerate() and rankability are different checks — a corpus can pass the first and still be statistically unrankable.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.6'; — no REINDEX.

[2.7.5] — 2026-09-09

The 1-bit dimension sweep is measured, and it produces the sharpest practical guidance the BQ work has yielded: 1-bit is a high-dimension technique. Also corrects a hi_dim_rerank documentation error. No shippable code change — binary byte-identical to 2.7.4. No wire change (v8), no SQL surface change, no REINDEX.

The penalty collapses as dimension rises

256/512/1024-d, 250 000 rows and 100 held-out queries per dim, matching § 0’s setup. Artefacts: benches/results/bq_dimsweep_20260909/.

1-bit R@10 at a fixed rerank window rises with dim at all seven swept windows, no exception:

window 256-d 512-d 1024-d
32 0.394 0.581 0.744
256 0.729 0.888 0.967
800 0.866 0.966 0.994
2000 0.933 0.990 1.000

At matched recall the window penalty versus 2-bit collapses:

dim 1-bit window for R@10 ≥ 0.95 2-bit penalty
256 4000 100 125×
512 800 32 25×
1024 256 32 8×

At 256-d, reaching R@10 ≥ 0.99 requires a window of 16 000 — reranking 6.4 % of the whole corpus. 1-bit is effectively unusable at 256-d and below; prefer it at 768-d and up. Storage moves the same way (1.902× → 1.946× → 1.971× versus 2-bit) because 1-bit’s fixed per-index overhead amortises away as dim grows: 30 % overhead over the dim/8 ideal at 256-d, 11 % at 1024-d. Both axes favour 1-bit more strongly at higher dimension.

Depth confirms it: R@100 at the auto default is 0.455 (256-d), 0.737 (512-d), 0.948 (1024-d) — at 256-d the default loses more than half the true top-100.

Validation. The 1024-d arm was re-run from a fresh database on a re-sliced corpus and reproduced § 0’s published recall bit-identically at all seven windows, p50s within 1–3 %. Independently re-verified against the published artefact. That validates the published figures and the rebuilt, isolated harness.

Caveat that bounds this. The 256-d and 512-d corpora are prefix slices of the 1024-d Cohere-wiki embedding, not natively-trained embeddings of those dimensions. A native 256-d model concentrates its information into 256 coordinates; a truncated 1024-d vector keeps only the first quarter of a representation spread across all of them. That likely makes the sliced low dims look worse than a native model would, so the trend is an upper bound on dim-sensitivity — directionally sound, magnitude not transferable. Nothing was padded to fabricate a higher dim.

Same contention caveat as § 0: recall and storage stand; absolute p50s are indicative.

Correction: the 1-bit hi_dim_rerank special case is a no-op at dim ≥ 256

docs/ONEBIT_BQ.md and docs/BQ_RECALL_BENCH.md said hi_dim_rerank treats a 1-bit index as high-dim “at any dim”, implying it widens BQ’s rerank window generally. That is literally true of the code and misleading about the effect. The auto window is clamp(effective_dim, 256..=1024), and since effective_dim = max(dim, 256) for 1-bit but dim for 2/¾-bit, the two are identical for every dim >= 256. The special case only widens the window below 256-d.

This strengthens the published results rather than weakening them: a 1-bit-vs-2-bit comparison at dim >= 256 with default settings compares equal windows, so § 0 and § 0.6c are quantizer-vs-quantizer, not knob-vs-knob. Both docs now carry a table showing exactly where the special case bites.

Found by re-deriving the clamp independently against the Rust source while chasing down why the sweep’s pre-registered prediction P-D3 had named the wrong mechanism.

On pre-registration

Three of the sweep’s four predictions were wrong in some respect — the storage direction was backwards, the hi_dim_rerank mechanism was misnamed (which is what surfaced the doc error above), and the latency-vs-dim shape was non-linear. Only the central hypothesis was confirmed, and it was understated. Recording predictions before the run is what made all of that visible instead of being quietly rationalised afterwards.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.5'; — no REINDEX.

[2.7.4] — 2026-09-09

Documentation accuracy and benchmark-harness isolation. No shippable code change — the binary is byte-identical to 2.7.3. No wire change (v8), no SQL surface change, no REINDEX.

Correction to v2.7.3’s latency figures

v2.7.3 published the 1-bit BQ p50s without disclosing that the harness had flagged every one of them. All 24 rows carry latency.contention.contended_flag = true — 1-minute loadavg 3.16–4.64 against the harness’s gate of 1.5. The artefact recorded this correctly; the write-up did not surface it. That was my error and this is the correction.

The gate is unreachable on that host and not because of the benchmark: a stuck systemd --user spinning at 77–86 % CPU for 31 days pins its idle loadavg near 2.0, so any run there is flagged. Mitigating measurement, taken rather than assumed: cpu_busy_pct on the pinned cores was only 16–26 % — the load is runnable-elsewhere processes, not saturation of the bench cores. So the ratios and the shape of the window-vs-recall curve stand, and the absolute milliseconds are indicative. Recall, storage, bytes/vector and build time are CPU-independent and unaffected.

README.md, docs/RECALL.md, docs/ONEBIT_BQ.md, docs/UPGRADING.md and docs/BQ_RECALL_BENCH.md now all carry the caveat. Found by a sub-agent checking its own contention flags; I had not been checking mine.

Measured: IVF + 1-bit BQ

docs/BQ_RECALL_BENCH.md § 0.6a, artefact benches/results/bq_ivf_20260909/. WITH (lists = N, bit_width = 1) builds and scans correctly. Storage overhead over flat BQ is +6.0 % — a fixed ~8.5 B/vector of coarse centroids and cell directory, so proportionally worst for the smallest codes.

The finding: IVF imposes a per-probe-count recall ceiling that a wider rerank window cannot break. At probes = 8, 1-bit saturates at R@10 = 0.846 and stays there from window 256 through 2000 — the true neighbours are not in the probed cells, and exact re-ranking cannot invent them. Flat BQ reaches 0.994 with no probe tuning.

This is the mirror image of Gap-B (v1.25.0), and the distinction is the useful part: there, high-dim recall loss was not retrieval-bound (cell recall 0.98–0.996) and a wider window fixed it; here it is retrieval-bound and the window is irrelevant. Same symptom, opposite cause — diagnose which one you have before reaching for a knob.

At 250k, flat BQ dominates IVF+BQ. That is a scale-dependent boundary, deliberately not a verdict: IVF’s whole value is bounding scan cost as n grows, and 250k is below the crossover. Recording it as a boundary is what the graph kind’s early iso-beam numbers should have done.

A 1M-scale arm was discarded, not published

A synthetic 1 M × 768-d run produced R@10 = 0.031 at window 32, which looks like catastrophic scale collapse. It isn’t — the generated corpus was statistically unrankable. Resolvability probe: 1st vs 100th nearest neighbour differed by only 6.6–10.4 % in cosine distance, versus 37–268 % on the real Cohere-wiki corpus. When the top-100 is effectively a tie, no quantizer can rank it and “recall” measures the tie-break order. Cause, narrowed on 2026-09-09 with direct measurement: the clustering worked (200 distinct centres; same-cluster distance 0.108 vs cross-cluster 0.990, a 9× separation) — the tie is inside each cluster. Within one cluster, a member’s 100 nearest neighbours span only 7.22 %, because 5000 iid Gaussian points at d = 768 concentrate: the pairwise separation norm has a relative spread of ≈ 1/sqrt(2d) ≈ 2.6 %. A plausible competing diagnosis — that the generator’s uncorrelated ARRAY(SELECT randn() ...) subquery had been hoisted, making all 200 centres identical (the v1.24.0 harness-bug class) — was tested and ruled out: the SQL genuinely invites that hoist (reproduced minimally: 1 distinct vector across 5 rows uncorrelated, 5 when correlated), but randn()’s volatility forced per-row evaluation here and the loaded centres are distinct.

Discarded with a full post-mortem, including the probe to run before trusting any generated corpus, in benches/results/bq_scale_20260909/DISCARDED.md.

Also surfaced by that arm: the harness’s ground-truth query wraps row_number() OVER (ORDER BY <distance>) around the per-query scan, forcing a WindowAgg → Sort → Gather shape where workers ship rows to the leader and the leader sorts the whole corpus (confirmed by EXPLAIN ANALYZE: Gather (actual rows=200000, loops=5) with a 4.3 MB external disk merge, versus per-worker top-N heapsort in 31 kB with a constant query vector). A real defect — but not the reason GT took 23 084 s, which the v2.7.6 notes over-credited it for. Restructuring measures 1.00× narrow / 1.25× on 3 KB-wide rows, and 6.4 h over 100M comparisons is 231 µs each against a 154 GFLOP job: the cost is per-row overhead across five full sorts of a 1 M-row wide table on a pre-AVX2 host, not the plan. Left unapplied pending validation at 1 M scale on an AVX2 host. Storage did confirm the dim/8 stride and an exact 2.00× 1-bit:2-bit ratio at 1 M rows. Note a synthetic corpus can pass is_degenerate() and still be unrankable — different checks, and only the second predicts whether recall means anything.

Harness: --run-id so concurrent arms cannot corrupt each other

Two arms running in one database silently corrupted each other three ways: a shared bq_query_set (one arm’s 256-d query set replaced another’s 1024-d one mid-sweep), DROP TABLE IF EXISTS bq_gt destroying 860 s of ground truth, and a hardcoded bqbench_ index prefix — index names are schema-scoped, so one arm’s CREATE INDEX failed against a sibling’s index on a different table, silently costing it an entire bit_width leg. All three are now namespaced by --run-id / BQ_RUN_ID, validated by the existing identifier guard. Default behaviour without --run-id is byte-identical.

Stale documentation corrected

  • Test counts: 334/334 (AGENTS.md) and 341/346 in six rows of PG_VERSION_SUPPORT.md → the actual 427 passed / 8 ignored, uniform across all seven CI legs pg13–19.
  • AGENTS.md: migration matrix was stuck at v1.27.1 and the wire version said 7 when it has been 8 since v2.0.0. Both corrected, and ~160 lines of v1.x release prose that duplicated CHANGELOG.md replaced with current state (kinds, which to recommend, the corruption-history warning for anyone touching persist/scan code, and the upstream ctid bug).
  • docs/PRODUCTION.md: new section on bounding maintenance_work_mem for high-dim IVF builds. Measured: bit_width = 4, lists = 512 at 1024-d reached 20.3 GB anon-RSS and was OOM-killed at maintenance_work_mem = 3GB, then completed in 50.6 s at 1 GB with two parallel workers. The effective ceiling is maintenance_work_mem × (1 + max_parallel_maintenance_workers), and an operator sees this as an unexplained backend termination.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.4'; — no REINDEX.

[2.7.3] — 2026-09-08

The 1-bit sign-BQ frontier is measured and published, and a real insert-path bug found while running it is fixed. No wire-format change (stays v8), no SQL surface change, no REINDEX.

Fixed: a 1-bit index built on an EMPTY table rejected its first insert

CREATE TABLE → CREATE INDEX ... WITH (bit_width = 1) (while empty) → INSERT failed with dim mismatch — index expects 0, row has 1024. The empty build stamps dim = 0 (no rows, no reloption-pinned dim) and insert_bq_row read the dim from the meta page. The flat path never had this bug because it takes the dim from the incoming row; the BQ path now does the same, seeding the mean as zeros for the 0-row case (a one-row corpus mean is that row, so every centred coordinate is 0 — exactly what a rebuild would produce).

Found on a real 1024-d corpus host while setting up the benchmark, not by a test: every in-tree BQ fixture happened to index an already-populated table, so all of them missed it. The regression test covers the exact failing order and also asserts a wrong-dim row is still rejected once the dim is pinned, so the fix does not paper over dim checking.

Measured: the 1-bit recall / storage / latency frontier

Closes the last “not yet published” claim in the README. arnold (i9-12900H, AVX2 — latency is only publishable on an AVX2 host per AGENTS.md), PostgreSQL 17.9, pg_turbovec 2.7.2, 250 000 × 1024-d Cohere-wiki as a native turbovec.vector column, 100 held-out queries (zero corpus overlap, verified), exact top-100 in-DB ground truth using the same operator the index serves, postmaster and driver pinned to P-cores 2-5, and all 24 configs confirmed to run a real Index Scan.

bit_width bytes/vector vs 4-bit R@10 ≥ 0.99 needs p50 there
1 142.3 3.98× smaller window 800 36.7 ms
2 280.5 2.02× smaller window 32 6.0 ms
4 565.7 1.00× window 32 9.2 ms

1-bit trades latency for storage, steeply: twice the storage saving of 2-bit, for 2.7–6.1× the latency at matched recall and a 25× wider exact-rerank window to reach R@10 ≥ 0.99. That is a storage-constrained-workload option, not a default — which is how the reloption was already documented, now with numbers behind it.

All four predictions registered before the run held. P2’s falsification condition was “1-bit crosses at a comparable window, which would mean the hi_dim_rerank 1-bit special-case is unnecessary” — it did not, so the data justifies that special-case rather than merely tolerating it.

Two further findings:

  • 1-bit degrades faster at depth than at k=10. R@100 tops out at 0.981 (window 2000) and is 0.948 at the auto default, while 2-bit reaches 1.000 by window 800. Budget a wider window if you paginate or re-rank past the top 10.
  • Harness cost trap, documented. --vec-expr accepts any SQL expression, and a cast like (emb::real[]::turbovec.vector) is evaluated per row per query — 200 M casts for a 1M × 200-query ground-truth build, ~53 s per query. Materialising a native turbovec.vector column first took one query from 53 s → 2.7 s (20×). The 1M arm was abandoned for this reason, not for anything about turbovec.

Caveats are stated in docs/BQ_RECALL_BENCH.md § 0.5 and not glossed: one corpus, one dim, one shared host (load ≈ 2 at start, so treat absolute milliseconds as indicative and the ratios as the result), and no IVF arm. This corpus also cannot separate 2-bit from 4-bit on recall — both sit at ≈1.000 nearly everywhere; it separates 1-bit from both, which is what it was run for.

Artefact: benches/results/bq_frontier_20260908/ (24 configs + sweep log).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.3'; — no REINDEX.

[2.7.2] — 2026-09-08

BUG#6 is now reported upstream. Documentation-only: the binary is byte-identical to 2.7.1 (the sole src/ edit is a doc comment). No wire change, no SQL surface change, no REINDEX.

The root-cause analysis and one-line core fix that v2.7.1 verified have been filed on pgsql-hackers:

ExecForceStoreHeapTuple() loses tts_tid, so ORDER BY-op index scans project an invalid ctid — 2026-09-08, with the patch attached as v1-0001-....

The filed patch’s execTuples.c hunk is identical to the one verified here by A/B build. The filed version additionally adds a core regression test to src/test/regress/{sql,expected}/gist.*, and that test was itself checked against both builds: it reports ctid_matches = 1 on unpatched 18.4 and 5 on the patched build, so it genuinely gates the fix rather than passing vacuously.

docs/FILTERING.md and the knn_scan_ctid_projection_upstream_limitation test now carry the thread link. That test still asserts the current (broken) behaviour, so it will fail once a fixed PostgreSQL reaches CI — which is the intended signal to flip the assertion (gated on the fixing version) and relax the docs, not a regression in this AM.

Until a fixed PostgreSQL ships nothing changes for users, and the workaround remains a proven necessity rather than a preference: MATERIALIZED CTEs, text casts inside a subquery, and extra subquery nesting were all tested and all still yield the sentinel. Chain on your own key column, or use turbovec.knn().

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.2'; — no REINDEX.

[2.7.1] — 2026-09-08

BUG#6 root cause proven against stock PostgreSQL, and the one-line core fix verified. Documentation + upstream-patch release: the binary is byte-identical to 2.7.0 (no wire change, no SQL surface change, no REINDEX). The only src/ edit is an expanded doc comment on the existing tripwire test.

BUG#6 is the one where SELECT ctid ... ORDER BY emb <=> q LIMIT k projects the invalid-item-pointer sentinel (4294967295,0). It was previously argued to be a core bug; it is now demonstrated:

  • Reproduced on stock PostgreSQL 18.4 using only core GiST, with zero turbovec loaded — thin diagonal polygons, where the bounding-box distance under-estimates the true polygon distance, so was_exact comes out false and tuples are routed through nodeIndexscan.c’s reorder queue.
  • Traced line by line in stock source. indexam.c:983 sets xs_recheckorderby for any AM that asks; nodeIndexscan.c:290 queues the tuple; reorderqueue_pop calls ExecForceStoreHeapTuple, whose TTS_IS_BUFFERTUPLE branch calls ExecClearTuple (which does ItemPointerSetInvalid(&slot->tts_tid)) and never restores tts_tid from tuple->t_self; tuptable.h:420’s slot_getsysattr returns &slot->tts_tid for the ctid system column. The sibling tts_heap_store_tuple does set tts_tid, so it is an asymmetry, not a design choice.
  • The fix is proven, not proposed. PostgreSQL built both ways on one machine, one script:

    unpatched patched
    ctid self-join, expect 5 1 5
    UPDATE ... WHERE ctid, expect 5 rows 1 5
    sentinel ctids at LIMIT 50 49/50 0/50

    The UPDATE line is the dangerous one: no error, it just affects one row instead of five.

  • Byte-identical across branches: the affected function body hashes the same in the 13.23, 14.22, 15.17, 16.14, 17.9 and 18.3 trees, so the one-line fix applies unchanged to every supported major.
  • No query-level workaround exists (verified, not assumed): WITH ... AS MATERIALIZED, casting to text inside a subquery, and extra subquery nesting all still return the sentinel, because it is already in the slot before any of them run. Only abandoning the index scan avoids it. So the documented guidance — chain on your own key column, or use turbovec.knn() — is the only real answer, and now it is known to be rather than assumed to be.

New finding. turbovec sees this on 100% of rows while core GiST loses only some. IndexNextWithReorder sets was_exact = (cmp == 0), so a tuple skips the queue when the AM’s advertised ORDER BY value compares exactly equal to the recomputed one. turbovec advertises f64::NEG_INFINITY, which never compares equal to a real distance, so every one of its tuples is queued. That is a difference in exposure, not in cause — advertising a real lower bound would only make the fault intermittent, which is harder to diagnose, not safer, so the -inf choice stands.

Ships docs/upstream/bug6-pgsql-hackers-DRAFT.md, a submission prepared for review and not sent, alongside the existing patch file (now carrying the A/B table).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.1'; — no REINDEX. Nothing in the extension’s behaviour changes; the fix this release documents is a PostgreSQL core patch, not something pg_turbovec can apply.

[2.7.0] — 2026-09-08

IVF + 1-bit sign-BQ compose, the Hamming kernel gets ~4.4× faster, and two corruption-class bugs shipped by v2.6.0 are fixed. No wire-format change (stays v8), no REINDEX, no SQL surface change.

Two bugs in v2.6.0’s BQ code — fix these by upgrading

Found by an audit while composing IVF with BQ, not by a field report:

  • MetaPageData::set_ivf_chains omitted bq_mean_count from its running chain-offset sum. An IVF+BQ build would have written the coarse-centroid chain on top of the corpus mean — destroying the centring vector that every sign code and every query depends on. This is the fourth occurrence of this bug class (v1.24.0 omitted graph_count; v2.6.0 found and fixed three sites). set_graph_chain had the identical omission (unreachable today, since graph + 1-bit is rejected) and is fixed too. Now guarded by an exhaustive pairwise no-overlap test, verified to fail when the omission is reintroduced.
  • Flat-BQ aminsert did not re-persist the tombstone bitmap, so every INSERT after a VACUUM silently resurrected every deleted row — the same M2 bug the graph kind fixed in v2.1.0, reintroduced because the BQ write path had no tombstone parameter. It also appended unconditionally, so re-inserting an existing heap TID added a second slot for the same row (unbounded growth under upserts, plus a duplicate id — the shape the flat bijection guard calls corruption). Both fixed; the write path now carries tombstones through a single meta write.

Only bit_width = 1 indexes were affected, and only on insert-after-VACUUM (resurrection) or re-insert (duplicate slots). 2/¾-bit indexes were never affected. There is no on-disk format change, so upgrading is sufficient — but an existing 1-bit index that has taken inserts after a VACUUM should be REINDEXed, since it may hold resurrected or duplicated slots.

WITH (lists = N, bit_width = 1)

Previously rejected with a clear ERROR. Now builds and scans: sign codes are stored cell-contiguous (the permutation IVF already applies) alongside the coarse centroids and cell directory, so a scan probes only the nearest turbovec.probes cells and runs Hamming within them — the dim/8 storage win combined with the probe-a-fraction scan win. Wire version stays 8: the shape is kind = KIND_BQ plus the existing v4 IVF chain fields.

BQ cells live in the raw L2-normalised space, not the rotated space TurboQuant IVF uses. TurboQuant trains cells in the rotated space because that is where its fine quantizer encodes; a BQ code is the sign of the centred raw coordinate, so there is nothing to align to. Getting this asymmetric is the sharpest available silent failure — a rotated query against un-rotated centroids probes the wrong cells and collapses recall with no error — so the scan skips the rotation explicitly and the tests assert per-id self-neighbour recovery.

An IVF+BQ aminsert degrades the index to a flat BQ scan (it appends and drops cell metadata). This is observable: lists is preserved and ivf_degraded is stamped, so turbovec.index_is_degraded() reports it — strictly better than the TurboQuant IVF insert path, which blanks lists and cannot be reported. Rebuild with REINDEX to restore cell-scoped scanning.

Hamming kernel: ~4.4× faster, and AVX2 measured then declined

The kernel now counts 8 bytes at a time instead of one. Measured 4.4–4.8× at embedding dims (100k×768d: 6.13 ms → 1.28 ms; 1M×768d: 59.0 ms → 13.7 ms; independently reproduced at 4.5–5.4× on a second machine). Stable from L3-resident to firmly-DRAM sizes, so the scan is not bandwidth-bound at these sizes.

No unsafe, no runtime CPU dispatch — every machine and architecture runs the identical instruction sequence. An AVX2 intrinsics kernel was written and proven bit-identical, then declined on measurements: it buys only ~1.8× more above dim = 512 and is a net loss below it (the 32-byte loop never runs at dim ≤ 256), while costing a second unsafe block on the scan hot path, a dim-dependent dispatch threshold, and a code path CI cannot exercise both sides of — the exact blind spot that let the v1.7.3 wrong-results bug ship. u64x4 was also measured and rejected (LLVM already extracts the ILP). Full numbers in docs/ONEBIT_BQ.md §7.

Bit-identity is proven in-tree, not assumed: 5720 random code pairs across 143 dims against a bit-by-bit reference sharing no code with the kernel, 825 top-k cases asserting the full sequence including tie order, a tie-density guard so that is not vacuous, and mutation testing (dropped tail, byte-swapped operand, AND for XOR, skipped last word, count_zeros) where each mutation fails 3–6 tests.

Benchmark harness for the unpublished BQ frontier

benches/scripts/bq/ plus docs/BQ_RECALL_BENCH.md: a sweep over bit_width ∈ {1, 2, 4} measuring R@10 against exact in-DB ground truth, storage, build time and — only on an AVX2 host — latency. It has not been run; no numbers exist yet. The AVX2 gate is structural: on a scalar host the driver does not time queries at all, rather than emitting a number that reads as fast. Ground truth uses the same operator the index serves (avoiding the cosine-vs-L2 trap), the re-rank window is recorded per row against a mirror of the Rust clamp, and iso-recall rows emit null when a bit_width never clears the target — that absence being the result. Predictions are written down in advance so a real run can falsify them.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.7.0'; — no REINDEX for 2/¾-bit indexes. A 1-bit index that has taken inserts after a VACUUM should be REINDEXed (see the bug notes above).

[2.6.0] — 2026-09-07

1-bit sign binary quantization (WITH (bit_width = 1)) works end to end — build, scan, aminsert, VACUUM. No wire-format change (stays v8), no REINDEX, no SQL surface change. Previously the reloption was accepted for forward compatibility but the build raised a clear “not yet implemented” ERROR.

1-bit is not TurboQuant-at-1-bit: the turbovec crate hard-rejects bit_width < 2 in both constructors, so this is a distinct scheme — sign binary quantization, the DiskANN/pgvector/Qdrant coarse code. Per-vector storage is dim/8 with no per-vector scale (half the 2-bit stride, and it also drops the codebook, the rotation/TQ+ chain and the blocked chain), scored by Hamming (popcount(XOR)) and then exactly reranked by the AM’s existing xs_recheckorderby machinery.

Why there is no wire bump

A new kind byte (KIND_BQ = 3), not a version bump. Every existing index keeps kind = SINGLE/COLBERT/GRAPH and decodes byte-identically — the same additive per-kind path v4→v5→v6 used. The three bq_mean_* meta fields occupy page offset 316, reserved-and-zero on every prior version, so an old meta page reads them as “absent”. (docs/ONEBIT_BQ.md §4 originally specified a 7→8 bump; that was written when v7 was current.)

Mean-centering is load-bearing

The naive sign-at-zero rule sets every bit on dense-positive data (measured R@10 = 0.0 on GIST), so the per-dim corpus mean is subtracted before taking signs, persisted, and applied to queries too. A corpus still collapsed after centering (constant / near-constant) is rejected at build rather than shipping a signal-free index. Scanning an index whose mean is missing ERRORs with a REINDEX hint rather than returning garbage.

Three instances of the v1.24.0 corruption class, found and fixed

write_tombstones_and_meta, the tombstone placement inside the rewrite path, and MetaPageData::total_blocks() each summed chain page counts without bq_mean_count. On a BQ index that would have placed the tombstone chain on top of the mean vector and under-sized the relation — the identical shape of the v1.24.0 graph bug (which omitted graph_count). Found by auditing every chain-offset sum in the tree, not only the path being added.

Notes

  • aminsert does not recompute the mean: that would invalidate every sign code already packed against the old mean, silently degrading the whole index’s ranking from one insert. Build-time mean is fixed; drift is a REINDEX concern. The insert is one exclusive lock across read-modify-write — an unlocked RMW is exactly what silently lost graph rows before v2.1.0.
  • VACUUM is tombstone-only, sharing the graph kind’s path (compacting would mean renumbering every slot).
  • turbovec_check reports kind = 'bq' and skips the v2.2.2 scales validation for that kind only — BQ has no scales chain, so validating one would report every BQ index corrupt.
  • The Hamming kernel is deliberately scalar (count_ones lowers to POPCNT); any hand-vectorised version must first be proven bit-identical against it. This is the v1.7.3 lesson — a mis-specialised kernel returned wrong ANN results on pre-AVX2 CPUs — applied pre-emptively.

Not yet supported

bit_width = 1 with lists > 0 (IVF) is rejected with a clear ERROR; the cell-contiguous sign layout and per-cell Hamming scan are unwired. bit_width = 1 with graph = true stays rejected in the reloption validator (and the graph kind is deprecated as of 2.5.0). BQ’s recall/latency frontier on real corpora still needs an AVX2-host run; the tests prove correctness, storage and scan behaviour, not a published frontier.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.6.0'; — no REINDEX.

[2.5.0] — 2026-09-07

Two changes: the graph kind is deprecated, and Phase S-1 partition pruning lands (additive SQL). No index wire-format change (stays v8), no REINDEX.

WITH (graph = true) is deprecated

It now emits a deprecation WARNING; the build path is scheduled for removal, with decode retained one further release so a stale graph index fails loudly with a REINDEX hint rather than silently.

The kind was added in v1.23.0 to chase HNSW’s query latency while keeping TurboQuant’s storage compression. Measured at matched recall it never delivers. Its apparent sublinearity holds only at iso-beam (p50 1.11× for a 10× corpus — but recall falls 0.605 → 0.472); once recall is held equal the curves diverge and never cross:

corpus / target flat IVF graph
SIFT-1M/128d, R@10 ≥0.95 0.98 ms, qps@8 1380 1.8 ms, qps@8 2039 26.2 ms, qps@8 299
GIST-1M/960d, R@10 ≥0.95 5.88 ms, qps@8 279 11.3 ms, qps@8 480 unreachable
GIST-10M/960d, R@10 ≥0.98 34.2 ms, qps@8 31 28.4 ms, qps@8 161 unreachable

It also loses on build time (57–90×), storage, and has no out-of-core path. It is IVF, not the graph, that beats flat’s O(n) wall. Use the default flat index below ~1M vectors and WITH (lists = N) at scale.

Retained deliberately: turbovec.graph_ef, pack::repack, and the coarse_graph work (Phase G-1) — that one navigates centroids, and IVF’s win partly rests on it. Deprecating the graph kind is not deprecating graph techniques.

Also recorded: the “60× parallel build speedup” was an artefact. graph_build_partitions_decide coupled shard count to thread count, and shards cost recall (GIST-1M R@10 0.920 at P=4 → 0.605 at P=83); threads at a recall-preserving P buy <5×. The partitioned-build parity test is annotated with the blind spot that hid this (2.5k rows/shard at dim 64, versus ~12k at 960d in reality) rather than re-tuned, since the kind is on its way out.

Phase S-1: partition-level coarse quantizer

At 1T scale the design is a partitioned parent with one turbovec index per partition and native Merge Append doing scatter → gather. That is correct for any N but the per-query fan-out is O(N) in the partition count: at ~100k partitions, opening an Index Scan per partition dominates. S-1 lifts IVF one tier up — each partition gets a summary (its mean in the original, un-rotated space, so summaries are comparable across partitions; each partition trains its own rotation, so the persisted coarse centroids are not) — and nearest_partitions() returns the Kp nearest so the caller fans out to Kp instead of N.

New SQL (additive): table turbovec.partition_summary, plus turbovec.refresh_partition_summary(parent, vec_col) and turbovec.nearest_partitions(parent, query, k_partitions, metric).

S-1 was first attempted in v1.29.0 and reverted when its #[pg_test] failed CI. The assertion was wrong, not the code: it demanded the pruned top-k be identical to the full-fan-out top-k, but pruning is an approximation by construction — scoring a query against partition means cannot be equivalent to scanning every partition, so on unclustered data a Kp < N probe can legitimately miss a true top-k member. The revived test asserts what pruning actually guarantees: exactly Kp partitions returned, the query’s own cluster ranked first on content-clustered data, and a recall floor against full fan-out.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.5.0'; — no REINDEX. Existing graph indexes keep working and will warn on rebuild.

[2.4.0] — 2026-09-07

WAL amplification follow-up: stable chain starts. No wire-format change (stays v8), no SQL surface change, no REINDEX. Minor rather than patch because the on-disk allocation layout of new/rewritten indexes changes (chain contents and decoding do not).

v2.3.0 stopped WAL-logging pages whose contents hadn’t changed, which cut the reported ~500 MB-per-commit to ~28 MB. This release addresses what was left. Chain starts were packed back-to-back — scales_first = codes_first + codes_count, ids_first = scales_first + scales_count — so any growth in the codes chain moved every scales and ids page to a new block number, and a relocated page is genuinely different, so it had to be rewritten. On the reporter’s 768d/4-bit index a codes page holds only 21 rows, so essentially every flush crossed an allocation boundary and paid ~27.6 MiB to move the scales+ids chains.

  • The three growing chains' allocations are now rounded up to a multiple of MetaPageData::PAD_PAGES (256 pages), so a shift happens once per ~5400 rows instead of once per 21. Modelled on that index, WAL per row falls from ~55 → ~5.7 KiB at batch=512, and ~1375 → ~37 KiB at batch=1, on top of v2.3.0’s own ~33×.
  • Cost is bounded slack: at most 3 * PAD_PAGES pages (6 MiB) per index, independent of index size. Chains smaller than one padding unit are not padded, so small indexes keep the previous layout byte-for-byte — without that gate a 1000-row 768d index would have gone 0.4 → 6.0 MiB (15×) to buy an amortisation it never needs.
  • padded_pages_needed() is the single definition, used by the planner and the incremental grow/shrink paths in relfile.rs. If one site padded and another didn’t they would disagree about where the following chains start — precisely the class of block-offset bug that silently corrupted a graph index in v1.24.0.

Why this is safe without a REINDEX: *_count keeps its existing meaning of “blocks allocated to this chain”, which is what every consumer already uses it for (sizing the relation, and locating the next chain via running sums). A chain’s contents are located by *_first plus n_vectors/rows_per_page and never by *_count (see read_chain), so a padded index decodes identically. Existing unpadded indexes keep working as-is; their chains relocate on the first full rewrite, which is safe by the v1.29.4 invariant (all chains are written before the meta page, so an interrupted rewrite leaves the old meta pointing at the old, intact chains).

Four new layout tests cover the invariants directly: chain starts don’t move when up to 10k rows are added; padding never changes content addressing; small and empty indexes aren’t padded; and across a dim × bit_width × n sweep every padded extent is large enough for its rows with overhead bounded by one padding unit.

docs/TESTING.md gains a section on measuring global counters, distilled from the three self-inflicted CI failures in v2.3.0 — concurrent #[pg_test]s share one cluster, so pg_current_wal_lsn() deltas capture other tests' WAL (the same operation measured 32 KB / 1097 KB / 1425 KB across runs and once ranked the arms inverted). Count a backend-local value you control, reset each arm to an identical starting state, and assert the control arm is non-trivial so “0 vs 0” can’t pass.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.4.0'; — no REINDEX.

[2.3.0] — 2026-09-07

WAL amplification fix: a flush now only WAL-logs the index pages that actually changed. No wire-format change (stays v8), no SQL surface change, no REINDEX. Minor rather than patch because insert-path write behaviour changes materially.

A field report (2026-09-08, pg.ddx.io) measured pg_turbovec inserts accounting for ~100 % of all WAL on the host: 1322 MB / 25 s with an embedding backfill running versus 31 kB / 25 s with that one unit stopped — a ~42,000× difference from a single writer doing ~100–1400 rows/minute. The decisive detail: pg_stat_statements attributed only 434 kB of a ~3 GB/60 s window to SQL, so >99.98 % of the WAL was generated outside any statement — index maintenance, not the INSERT.

WAL per commit was roughly constant and close to the 882 MB index size (~500 MB at 16 rows/txn, ~754 MB at 128, ~670 MB at 512) and WAL per row a clean 1/batch curve with no knee — the signature of the whole relfile being rewritten and fully WAL-logged on every flush. A scratch-index A/B isolated it to commit boundaries alone: 32 × 16 rows = 643 MB WAL versus 1 × 512 rows = 2.6 MB (245×), both after wal_compression=zstd. Cost: ~4.3 TB WAL/day, archived off-host so paid for twice, and the dominant consumer of the host NVMe’s endurance (49 % used, 172 TB written in 5306 power-on hours). It also kept num_requested checkpoints at ~2× num_timed with a 220–390 % FPI ratio, which sent the operator chasing a checkpoint misconfiguration.

Cause: write_chain_at registered every page of every chain with GENERIC_XLOG_FULL_IMAGE. But reconcile_flush_image appends new slots at the end and updates touched slots in place, so every other page was already byte-identical on disk — we were paying a full-page image to rewrite pages with their own contents.

  • Each full page is now compared against the bytes about to be written (under a shared buffer lock) and skipped when identical, so WAL scales with bytes changed rather than index size. Only pages already carrying our own no-hole header are eligible, so a page whose header GenericXLogFinish would treat as a hole is never inherited; a skipped page is by definition already correct, so crash recovery is unaffected.
  • The GenericXLog state is started lazily, so a batch in which every page is skipped emits no WAL record at all instead of an empty one.
  • Removed the dead, never-called write_chain() helper. It wrote pages with MarkBufferDirty and no WAL — a latent footgun that would have made a skipped page’s missing WAL permanent. Every page mutation now provably goes through GenericXLog (zero live MarkBufferDirty sites).
  • Test insert_wal_scales_with_change_not_index_size drives a real flush (via the existing flush_to_relfile_for_test hook — a #[pg_test]’s outer transaction always rolls back before PreCommit fires, so the aminsert path can’t be observed in-band) on a ~109-chain-page index and asserts a no-change flush WAL-logs <1/10th the pages of a full rewrite. Fail-before/pass-after via SKIP_UNCHANGED_PAGES.

    It counts pages registered for WAL rather than diffing pg_current_wal_lsn(), which is load-bearing: #[pg_test]s run concurrently against one cluster, so a global LSN delta also captures every other test’s WAL. Three CI runs reported the same operation as 32 KB, then 1097 KB, then 1425 KB, and once ranked the two arms inverted — the LSN measurement was never valid in either direction. Registered pages are per-backend, deterministic, and are the direct driver of WAL volume (one full-page image each).

    Note the residual cost on an append (as opposed to a no-change flush): chain starts are packed back-to-back (scales_first = codes_first + codes_count), so when the codes chain grows by one page every scales and ids page lands at a new block number and genuinely must be rewritten. Measured 392 KB vs 1416 KB always-rewriting on a 20k-vector/64d index. Making chain starts stable across growth would shrink that further and is left as follow-up.

Batching still matters and the docs now say so: WAL scales with commits, so a flush’s remaining cost (tail pages, meta page, IVF cell directory) is paid per transaction. docs/PRODUCTION.md gains a “WAL cost of inserts — batch your writes” section including how to measure it with LSN deltas, since pg_stat_statements will not show it.

Also documents the partial-index verification gotcha (Ask 3 of the 2026-09-05 report): a KNN query that omits a partial index’s predicate silently gets a sequential scan, which masks index faults as latency problems — always confirm EXPLAIN shows Index Scan using <name>.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.3.0'; — no REINDEX.

[2.2.2] — 2026-09-05

turbovec_check() no longer has a scan-fatal blind spot. Code-only — no wire-format change (stays v8), no SQL surface change (the reason column already exists as of 2.1.0), no REINDEX required to upgrade.

A field report (2026-09-05, pg.ddx.io) found an IVF index — ~2.18 M vectors, 768 d, bit_width = 4, lists = 1400 — on which every KNN scan failed instantly with turbovec’s InvalidScaleValue { slot: 1, value: -2.559434e22 }, while turbovec_check() reported is_corrupt = false, count_matches = t, duplicate_id = NULL. Because the operator’s automated self-heal trusts that check, the index degraded silently until a manual REINDEX (~12 min) cleared it. The reporter correctly identified this as the more important of the two issues they filed: a checker that cannot detect an index failing 100 % of scans is a false-confidence generator.

Root cause of the blind spot (a deliberate, documented decision that this incident falsifies): the check read only the meta page and the ids chain, skipping “the much larger codes/scales chains”. But the scales chain is the cheap one — one f32 per vector, 8.7 MB at 2.18 M vectors, versus the 17.4 MB ids chain it already read, versus 0.84 GB for the codes chain it rightly still skips — and it is load-bearing for TurboQuantIndex::from_parts, i.e. exactly the gate a scan hits.

  • turbovec_check() now validates every persisted scale (finite, non-negative, sane magnitude — the same invariants from_parts enforces) for every index kind, inside the same shared rewrite-lock bracket as the meta/ids read so all three are one consistent observation, and names the failing slot in reason. Cost is +50 % on an already-cheap query; no opt-in deep flag, because a default that cannot see a scan-fatal fault is the bug.
  • The scan-path rejection is now a proper PostgreSQL ERROR naming the fault and the REINDEX INDEX recovery, instead of the bare Rust .expect("...from_parts rejected raw parts") string an operator used to get with no hint that REINDEX fixes it.
  • Regression test turbovec_check_detects_invalid_scale writes the reported value (-2.559434e22) into the persisted scales chain of an IVF index via a page-straddle-safe test-only helper and asserts the checker flags it; fail-before/pass-after.

Note this is the same failure shape as the graph lost-update called out in the 2.1.0 notes: internal self-consistency of meta+ids is not evidence of scannability. The originating corruption itself (a damaged scales chain on an index that had survived several pg_cancel_backend/pg_terminate_backend events, a kill -QUIT immediate shutdown and a restart that aborted an in-flight backfill) is not yet reproduced and remains under investigation; this release makes it detectable and actionable rather than silent.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.2.2'; — no REINDEX to upgrade. (An index already damaged needs a one-time REINDEX INDEX <name>;, which this release will now actually tell you about.)

[2.2.1] — 2026-09-05

Safety patch: the PARALLEL graph build was effectively uncancellable. Code-only — no wire-format change (stays v8), no SQL surface change, no REINDEX.

v2.1.0 made the graph build interruptible, but its hook lives in thread-local storage and is therefore invisible to rayon workers by design (a worker must never longjmp out of the pool). That left a hole nobody had measured: during the partitioned build the driver parks in a futex inside rayon’s join for the entire parallel phase, so it cannot reach a poll either. Measured on a 10M-node build: pg_cancel_backend() and a direct SIGINT were both ignored for over 13 minutes while 32 threads ran at 100% CPU — and because no backend code was executing, pg_stat_activity reported wait_event = NULL, so the runaway build looked idle. An operator had no way to stop it.

Workers now consult a cheap, thread-safe abort predicate at their safe points and stop producing work; the driver raises PostgreSQL’s real cancel/terminate error at its own poll once the phase collapses, so all error raising still happens only on the backend thread. The predicate is injected by the PG-side caller, so src/index/graph.rs stays Postgres-free and the mechanism is unit-testable without a server (partitioned_build_bails_out_when_abort_is_requested).

Found while measuring whether the graph kind could be made production-grade at scale (see the 2.2.1 note in docs/UPGRADING.md and the deprecation discussion for v2.3.0).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.2.1'; — no REINDEX.

[2.2.0] — 2026-09-04

Graph-kind scan-beam retune. The graph kind’s scan-time beam width is now its own knob (turbovec.graph_ef, default auto = 512) instead of a side effect of the candidate count, and the auto default is set from a measured 1M-scale recall-vs-latency frontier on SIFT-128 and GIST-960. No wire-format change (stays v8), NO REINDEX — but a new GUC is a SQL-surface addition, so this is a MINOR.

Fixed — the graph “high-dim recall cliff” was a BEAM bug, not a scorer bug

The tracked “GraphScorer diverges from the kernel at dim >= 512” is disproved. An exact-f32-oracle control gives recall identical to the LUT scorer (within 0.015), and recall recovers monotonically with the BEAM on the same index and the same scorer (search_k 32 → 128 → 512 gives R@10 0.750 → 0.990 → 1.000). The real mechanism was the scan-time beam width, and — worse — which knob was setting it:

  • Through v2.1.0, graph_search computed ef = (k * 4).max(64) where k was the candidate count scan.rs had already widened. Since turbovec.hi_dim_rerank = auto raises that count to clamp(dim, 256..=1024) for dim >= 256 — a knob whose actual job is the flat/IVF exact-rerank window over a cell scan’s quantized ranking — the graph’s beam was being set as a side effect of an unrelated feature.
  • The visible symptom was an inverted cliff. With hi_dim_rerank = auto, 128d was the WORST (R@10 0.720: 128d is below the rerank threshold, so the beam stayed at the bare 64 floor) while 384/512d looked perfect (0.99/1.00: the inflated candidate count multiplied the beam up). With hi_dim_rerank = off recall collapsed at EVERY dim (~0.40/0.38). Recall was never dim-dependent; it was beam-dependent, and the beam was accidental.

Fixed by splitting graph_search into graph_search_with_ef(.., ef, ..) (the live scan path, beam resolved from the new GUC) and the original graph_search (the pure-k fallback for callers with no live GUC — unit tests, the aminsert findability probe). scan.rs’s graph arm now passes the user’s OWN candidate count (search_k * oversample), not hi_dim_rerank’s floor. hi_dim_rerank keeps doing its real job for flat/IVF, unchanged.

Verified on 42 paired (k, ef) configs across SIFT-200k, SIFT-1M and GIST-1M: off and auto now give recall within 0.0000 of each other at every beam, at 128d AND at 960d.

Added

  • turbovec.graph_ef (int, Userset, default 0 = auto = 512, range 0..=1000000) — the graph kind’s recall/latency dial, the direct analogue of hnsw.ef_search (and of turbovec.probes for IVF). 0 = auto; a positive N pins the beam. Always clamped up to the query’s candidate count (a beam narrower than it cannot fill it) and down to the live corpus size. Pure scan-time knob: no wire change, no REINDEX, honoured immediately by any graph index built by any 2.x binary.

Changed — defaults

  • Graph scan beam default 64 → 512 (GRAPH_SCAN_EF_DEFAULT), and it is now ABSOLUTE rather than (k * 4).max(64). Measured frontier (32 vCPU AVX-512, 1M rows, R@10 vs exact cosine GT, warm p50, search_k pinned to max(32, k)):

    corpus beam 64 beam 512 beam 2048
    SIFT-1M (128d) R@10 0.943 / 1.25 ms R@10 0.990 / 5.19 ms R@10 0.991 / 14.13 ms
    GIST-1M (960d) R@10 0.760 / 8.65 ms R@10 0.920 / 34.55 ms R@10 0.966 / 93.44 ms

    At 128d 512 is the unambiguous knee: it is within 0.001 R@10 of the 2048-beam ceiling at 0.37× its p50, and doubling past it buys +0.001 for 1.7× the latency. At 960d the curve has NOT plateaued by 2048, so 512 is a deliberate latency cap — the widest beam that keeps 960d p50 under ~35 ms — not a knee; SET turbovec.graph_ef = 2048 reaches 0.966 for ~95 ms if that is the trade you want.

    ⚠️ One configuration REGRESSES on recall, deliberately. A 960d graph index queried with hi_dim_rerank = auto (the default) goes R@10 0.971 → 0.920 and R@100 0.958 → 0.826, in exchange for p50 194 ms → 34 ms (5.6×) and 196 ms → 42 ms (4.7×). The old numbers came from the side-effect beam of 3840 that auto was silently buying at 960d — never a documented beam, and it evaporated the moment a user set hi_dim_rerank = off (0.840, the “collapse”). A ~200 ms p50 is not a defensible default, and the old recall is now REACHABLE and documented for the first time: SET turbovec.graph_ef = 3840 reproduces the pre-patch auto beam exactly. 128d indexes improve in BOTH modes (R@10 0.976 → 0.990 at 1M) and lose nothing — 128d is below hi_dim_rerank’s dim >= 256 threshold, so 128d auto was never above off.

Known limitation — the graph kind is NOT the fastest kind at 1M

Measured on the same box, same corpora, same GT, at k=10: the flat kind beats the graph on BOTH recall and latency on BOTH corpora. SIFT-1M flat R@10 0.993 / 0.96 ms vs the graph’s best-defensible 0.990 / 5.19 ms (5.4× lower latency, higher recall). GIST-1M flat R@10 0.997 / 5.91 ms vs the graph’s 0.920 / 34.55 ms (5.8× lower latency, +0.077 recall) — and even beam 2048 (0.966 / 93 ms) does not catch flat. At 1M rows on a 32-core AVX-512 host, turbovec’s SIMD full scan (one linear sweep of a compact quantized buffer at memory bandwidth) is simply faster than navigating a graph (~ef scattered gathers with a serial dependency between hops). The graph’s advantage is asymptotic and 1M is below the crossover on this hardware. Guidance: use flat (or IVF for a smaller RAM footprint) at this scale. The graph’s one measured win is CONCURRENCY scaling — the flat scan saturates memory bandwidth and goes 713 → 1313 qps from 1 to 8 clients (1.8×) while the graph goes 688 → 4906 (7.1×) — so it is the better throughput engine under load even where it loses isolated p50. Full curves and the reasoning in docs/GRAPH_EF_BENCH.md.

Docs

  • docs/GRAPH_EF_BENCH.md — the full measured frontier (both corpora, both hi_dim_rerank settings, the whole ef grid, at k = 10 and 100), the flat/IVF comparators at the same k, the before/after A/B on byte-identical index files, and the reasoning behind 512.
  • docs/PRODUCTION.md — turbovec.graph_ef reference section.
  • README.md — GUC table + count (18 → 19).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '2.2.0'; and restart the backend (the new GUC is registered in _PG_init). No REINDEX: the wire format is unchanged (v8) and the beam is resolved at scan time, so existing graph indexes get the new default immediately. Flat and IVF indexes are unaffected — the beam does not exist on those paths. To restore a pre-v2.2.0 effective beam exactly: SET turbovec.graph_ef = 128 for what hi_dim_rerank = off gave at any dim (and what auto gave below 256d), or = 3840 for what auto gave at 960d.

[2.1.0] — 2026-09-04

Graph-kind correctness release. The WITH (graph = true) kind had a critical concurrent-INSERT data-loss bug, could return fewer than k rows, and its build could not be cancelled. All three are fixed here, plus graph health monitoring and an upstream-bug writeup. No wire-format change (stays v8), NO REINDEX — but turbovec_check()’s return type changes, so this is a MINOR (see the upgrade note).

Fixed — graph kind

  • C1 (CRITICAL): concurrent INSERT into a graph index lost data. insert_graph_row did an unlocked whole-index read, then a blind whole-relfile rewrite. Two concurrent graph inserters produced either (a) a torn adjacency chain — corrupt graph adjacency chain: graph offsets[n]=54272 != neighbors.len()=54240, a loud ERROR — or, worse, (b) a silent lost update: the losing inserter’s rows were simply gone (heap rows committed with no index entry) while turbovec_check() still reported is_corrupt = false, because the surviving relfile was internally self-consistent. Measured fail-before: 8 of 100 concurrent inserts lost, 8 heap rows unindexed. Fixed by holding ONE lock_relfile_write across the entire read-modify-write, re-reading the meta page under it, and adding guards that ABORT rather than reconcile onto torn state. Also folded in the tombstone two-write gap: the bitmap is now planned into the single meta write (previously a second write re-attached it, and a crash in that window permanently un-tombstoned every VACUUM delete). Validated: 320/320 concurrent inserts clean, and a sustained 240 s run of 6 inserters against a concurrent DELETE/VACUUM loop stayed is_corrupt = false with zero SIGABRT.
  • BUG#2: graph_search could return fewer than k rows. Two real causes: tombstoned nodes are never routing hops, so a large dead set disconnects the live remnant (n=2000 at 99% dead returned 2 rows for k=10); and a degenerate corpus of tied distances collapses RobustPrune’s out-lists (500 identical rows returned 8 for k=10). Fixed with a bounded backfill from unvisited live slots, gated behind the existing out.len() >= k guard so the normal path is byte-identical.
  • FINDING#2: a graph CREATE INDEX could not be cancelled and hung pg_ctl stop -m fast (the rayon build had zero interrupt polling). Added driver-thread stage polls mirroring ivf_build_and_write, plus a TLS-scoped hook the Postgres-free graph module calls at its own stage boundaries and every 256 nodes — a rayon worker cannot see the TLS slot, so it can never longjmp out of the pool.

Documented — not a pg_turbovec bug

  • BUG#6: a kNN index scan projects a sentinel ctid (4294967295,0) instead of the real heap tid. Root-caused to PostgreSQL core, not this extension: ExecForceStoreHeapTuple()’s buffer-slot branch calls ExecClearTuple() (which invalidates slot->tts_tid) and never restores tts_tid from tuple->t_self, while its sibling ExecStoreHeapTuple() does — a plain asymmetry. slot_getsysattr() reads the projected ctid straight out of tts_tid. Any AM that sets xs_recheckorderby = true is affected; core GiST reproduces it with no turbovec loaded, and a one-line core patch fixes both. Row data is unaffected (xs_heaptid is correct), and the documented turbovec.allowlist-from-ctid recipe still works (it harvests from a filter scan). Only chaining off a kNN scan’s ctid breaks. The proposed core patch is in docs/upstream/, the limitation and the supported alternatives are in docs/FILTERING.md, and a tripwire test fails the moment upstream fixes core so this note gets retired.

Fixed — BUG#5 (monitoring, details)

  • BUG#5: turbovec_check() was graph-adjacency-blind. It validated the flat/IVF ids bijection + tombstones but never looked at a graph index’s CSR adjacency chain, so a corrupt adjacency (a torn write, or the concurrent-insert corruption) reported is_corrupt = false — operators had no health signal for the graph kind at all. For a kind = graph index the check now decodes the adjacency chain and validates: offsets monotonically non-decreasing and consistent with the neighbor-array length, every neighbor id < n_vectors, the entry point in range AND having at least one out-neighbor, no self-loops, strictly-ascending (deduplicated) neighbor lists, the adjacency’s node count matching meta.n_vectors, and at least one edge when n_vectors > 1 (not trivially disconnected). Cost is O(n + edges) over the adjacency chain only — no vector data is read, so it stays cheap enough to poll. Flat/IVF/ColBERT behaviour is byte-identical (the new code is gated on meta.is_graph()).
  • The chain decode used by the monitoring path is now non-fatal (relfile::try_read_graph_adjacency), so a malformed chain is REPORTED rather than ERRORing out of the operator’s health query. It also bounds the chain against the physical relation length before reading, so a truncated relfile is reported instead of raising read_chain’s hard past-EOF ERROR. The scan / insert / VACUUM paths keep the ERROR policy (they cannot proceed on a broken graph) and share the same decode implementation.

Changed (SQL surface — additive column, but see the upgrade note)

  • turbovec.turbovec_check(regclass) gains a trailing reason text column: NULL when healthy, otherwise a human-readable description of the first problem found (duplicate id, count drift, or the specific graph-adjacency invariant violated). Adding an OUT column changes the function’s return type, which CREATE OR REPLACE FUNCTION cannot do — the release that ships this needs a DROP FUNCTION turbovec_check(oid); CREATE FUNCTION ... pair in its sql/pg_turbovec--<prev>--<this>.sql upgrade script, and it is a MINOR bump at minimum (never a patch). No wire-format change (still v8); no REINDEX.

[2.0.0] — 2026-08-26

MAJOR: pg_turbovec now runs on upstream turbovec 1.0.0. This adopts the stable turbovec 1.0.0 crate (Ryan Codrai’s TurboQuant implementation, first stable release) in place of the long-lived 0.9.0 fork. On-disk wire format v7 → v8 — a breaking change, with a documented REINDEX-from-heap migration (below).

What changed

  • turbovec 0.9.0 → 1.0.0. The fork’s hand-rolled centroids/boundaries codebook + QR rotation are replaced by turbovec 1.0.0’s TQ+ per-coordinate calibration and v5 block-Hadamard rotation (which also drops the OpenBLAS build dependency). Every encoded byte differs from v7, hence the wire bump. pg_turbovec carries two small additive fork patches on top of stock 1.0.0 (pub fn repack and a re-exposed IdMapIndex parts API used by the buffer-manager cache-fill path), both tracked to be offered upstream.
  • Wire format v7 → v8. MetaPageData::version = 8, EXPECTED_WIRE_FORMAT_VERSION = 8.
  • Materially faster, same storage. On EC2 i4i.8xlarge (AVX-512), head-to-head vs v1.29.7: SIFT-1M flat p50 12.9 ms → 1.3 ms (9.6×); GIST-1M IVF p50 @ R@10 0.95 512 ms → 92 ms (5.6×); cold-scan (SIFT IVF) 3139 ms → 399 ms (7.9×); build 1.2–2.1× faster. Index size byte-for-byte unchanged (77 MB SIFT-1M, ~4.96 GB @ 10M). Recall matched-or-better at matched config. Determinism intact (the build-parallelism byte-identity gates all pass; v5 rotation removed the old OpenBLAS nondeterminism). Two minor warm-flat regressions (~11–15% at already-saturated recall), no correctness impact.
  • All v1.29.x corruption fixes (meta-LAST torn-write ordering, reconcile-on-flush, pre-flush validate-all, IVF lists==0 dup-gate, VACUUM shrink guard, read_chain bounds, SubXact rollback) are re-proven to fire on the v8 persist path. The 90-minute no-VACUUM upsert+writer-restart corruption A/B ran clean on v8 (92 checks is_corrupt=f, 0 dup-id, 0 SIGABRT, 26 restarts).

Migration — REQUIRES a one-time REINDEX per index (wire v7 → v8)

ALTER EXTENSION pg_turbovec UPDATE TO '2.0.0'; + restart the backend, then REINDEX INDEX <name>; once per turbovec index. A pre-v8 index opened under 2.0.0 is detected (MetaPageData::is_legacy_v7()) and ambeginscan ERRORs at first scan with a REINDEX INDEX <name>; hint — never a silent misread (validated end-to-end on EC2: build v7 → open on 2.0.0 → ERROR+hint → REINDEX → serves correctly on v8, same neighbors).

An in-place page converter was investigated and rejected: measured recall loss of −20.7 pp @ R@10 (SIFT-1M 4-bit) and catastrophic at 2-bit, from double quantization at re-encode. REINDEX-from-heap re-encodes the heap’s source vectors directly (full recall) — the heap is the corpus, so this is not a rebuild-from-external-corpus. See docs/UPGRADING.md.

[1.29.7] — 2026-08-25

Numerical-robustness patch — normalise_into (run on every indexed row via normalize_on_insert) computed the reciprocal norm as (1.0_f64 / norm) as f32, which overflows to +inf when norm is a tiny-but-nonzero f64 (a vector whose elements are near f32 underflow, e.g. a single ~2e-39 coordinate). The +inf reciprocal then poisoned every element (x * inf = inf), so the “normalised” vector had inf coordinates and infinite norm — feeding garbage into the quantizer for that row. Now divides per-element in f64 and casts each result to f32 ((f64::from(x) / norm) as f32), which stays finite because |x/norm| <= |x| for a real vector. Patch bump — no wire change (v7), no SQL surface change, no REINDEX.

Found during the pg_turbovec 2.0.0 (turbovec 1.0.0) port’s full test re-run: the pre-existing Hegel property test prop_normalise_is_unit_norm_and_idempotent flaked (~1 in N seeds) on the norm inf assertion. Added a deterministic regression test normalise_tiny_norm_stays_finite (fail-before/pass-after proven).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.29.7'; — no REINDEX. A row inserted under an older binary whose vector hit this edge would have stored a garbage (inf) code; such a row (if any) is corrected by re-inserting it. In practice the trigger requires near-underflow input magnitudes, which real embeddings do not produce.

[1.29.6] — 2026-08-15

Dependency-hygiene patch — clears all outstanding RustSec advisories in the transitive tree. Patch bump, Cargo.lock-only (no source change, no wire change (stays v7), no SQL surface change, no REINDEX, byte-identical build output — the gemm, turbovec, and pgrx = 0.19.1 pins are unchanged, so IVF build determinism is preserved).

cargo update pulled semver-compatible upstream fixes:

  • crossbeam-epoch 0.9.18 → 0.9.20 (RUSTSEC-2026-0204, invalid pointer dereference in the fmt::Pointer impl) — reached via rayon, which pg_turbovec uses for build/scan parallelism. This is the one advisory in a shipped runtime path; pg_turbovec never formats those pointers, so exposure was nil, but the dependency is now patched.
  • tokio-postgres 0.7.17 → ≥ 0.7.18 (RUSTSEC-2026-0178), postgres-protocol 0.6.11 → ≥ 0.6.12 (RUSTSEC-2026-0179 / 0180) — these reach the tree ONLY through pgrx-tests, a dev-dependency (the test harness’s PostgreSQL client). They are not present in the shipped pg_turbovec.so and were never a production exposure; patched for a clean cargo audit.

Remaining cargo audit output is two “unmaintained” warnings (serde_cbor via pgrx, paste via turbovec/statrs and hegeltest) — not vulnerabilities, and both are upstream-controlled (pgrx’s PostgresType serialization and a transitive math dep), not fixable from this crate.

Upstream turbovec: intentionally NOT advanced. Our pinned rev (befc4cbf, the tip of the feat/unblock-inverse-repack branch, 2026-07-06) carries the pub from_parts + pack::unblock surface pg_turbovec depends on and is NEWER than upstream main’s last release (0.6.0, 2026-05-25). main diverged in a slim-down direction that removes that surface, so advancing the pin would break the build; rebasing our branch onto main’s encode-vectorization / codebook-caching improvements is a deliberate, separately-validated kernel change (wire / determinism-sensitive), not a routine bump.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.29.6'; — no REINDEX, no behavior change (dependency hygiene only).

[1.29.5] — 2026-08-15

Production-hardening patch from a deep code re-audit + an at-scale feature stress test. Patch bump — no wire change (stays v7), no SQL surface change, no REINDEX (ALTER EXTENSION pg_turbovec UPDATE is sufficient). Every fix is a guard or ordering correction; none change results on valid input.

  • IVF incremental INSERT regression fix (introduced in v1.29.4). v1.29.4 added an on-disk duplicate-id guard to the deferred-flush reconcile path, but it ran ungated — and an IVF index legitimately stores the same external id in multiple cells (soft-assignment to the 2nd..Mth nearest cell). So the first INSERT into any soft-assigned IVF index tripped refusing to reconcile onto a corrupt .tvim id table (id N appears in more than one slot on disk) and aborted — breaking IVF incremental insert. The guard is now gated to the bijective flat/single kind (lists == 0), matching the insert/read paths. Regression test ivf_insert_after_soft_assign_does_not_falsely_abort.
  • Multi-index partial-flush corruption-spreader (C-NEW-1). On a table with 2+ turbovec indexes, if the PreCommit flush of the second index tripped a persist guard (ERROR/longjmp), the first index’s relfile pages were already physically written (GenericXLog is not rolled back on abort) — leaving one index with phantom CTIDs and a sibling missing rows. The PreCommit flush now does a pure validate-all pass (no I/O) before writing any index, so a guard trip aborts with zero physical writes.
  • Corrupt/torn-meta unbounded read + palloc guard (H-NEW-2/3). read_chain now bounds the chain against the physical relation length (RelationGetNumberOfBlocksInFork) and rejects rows_per_page == 0 before allocating or walking blocks — a bit-flipped/truncated meta no longer over-reads past EOF or allocates a Vec sized by a corrupt n_vectors; it raises ERRCODE_DATA_CORRUPTED + a REINDEX hint.
  • SAVEPOINT / subtransaction rollback (M-NEW-4). Registered a SubXactCallback; a ROLLBACK TO SAVEPOINT now conservatively invalidates the dirty cache set, so a rolled-back insert is no longer persisted onto disk at top-level commit (was index bloat + silent recall loss, masked by recheck).
  • VACUUM shrink guard gap (M-NEW-5). The flat swap-remove VACUUM path — the one whole-relfile mutation that bypassed write_full_inner’s id-0/dup guard — now re-checks the surviving ids' bijection before committing the shrink (flat kind only).
  • Empty-index KNN query fix (NEW BUG #1). ORDER BY emb <-> q LIMIT k on a freshly-created, not-yet-populated index returned 0 rows instead of ERROR: query dim N != index dim 0 (the dim check now runs after the empty-index early return) — fixes the “create index, backfill later” pattern.
  • Overflow hardening (L-NEW-7): MetaPageData::total_blocks() uses saturating_add so a corrupt meta can’t wrap a large block count to a small total.

The re-audit also VERIFIED (independently, at production scale on EC2): v1.29.4’s field-report torn-write fix holds clean for 95 min / 85 mid-flush writer restarts (v1.29.3 corrupts after 2), and the whole flat/IVF feature matrix + the sparsevec OOM guard are solid. Known issues still tracked for a dedicated release (graph-kind only): graph concurrent-insert lost-update, graph wrong-results at dim>=512, graph under-return at 128d, graph build un-cancellable, turbovec_check graph-adjacency-blind, and the scan sentinel-ctid projection. The graph kind (WITH (graph = true)) should be treated as experimental.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.29.5'; — no REINDEX, no downtime beyond the .so swap + reconnect.

[1.29.4] — 2026-08-14

Data-corruption fix: VACUUM-INDEPENDENT torn-write of the .tvim relfile on writer interrupt/restart. Patch bump — no wire change (stays v7), no SQL surface change, no REINDEX to upgrade (existing indexes are read + written correctly in place by the new binary). A running index that already corrupted needs a one-time REINDEX to clear the bad state, but the upgrade itself is ALTER EXTENSION + restart.

v1.29.2/1.29.3 fixed the concurrent-VACUUM lost-update races. A deployment (pg.ddx.io/agora) that runs no VACUUM at all (autovacuum_count = 0, no DELETE) still corrupted continuously — signature always id 0 in more than one slot, ~every 30 min, on a 1.84M-row 768d bit_width=4 partial flat index driven by a continuous INSERT … ON CONFLICT DO UPDATE upsert writer that the orchestrator restarts every ~15 min.

Root cause (reproduced + dumped on EC2, unpatched corrupts in ~40 s): the deferred aminsert PreCommit flush rewrites the whole relfile via write_full_inner, which polls check_for_interrupts!() inside write_chain_at and commits each GenericXLog batch as a physical, WAL-logged page change that is NOT rolled back on transaction abort. The writer-restart is a pg_terminate_backend (SIGTERM → ProcDiePending); when it lands mid-flush it longjmps (FATAL) out of the chain write. With the pre-fix ordering — meta page written FIRST, then the codes/scales/ids chains — an interrupt after the meta commit but before the ids chain finished left meta.n_vectors = N pointing at an ids chain whose newly-appended (or in-place) region was still the zero-initialised page bytes. On reload that reads back as a contiguous run of id 0 slots — duplicate ids … id 0 appears in more than one slot (XX001) + SIGABRT — with count_matches still TRUE (the chain is physically N slots long). On disk we dumped the exact fingerprint: a contiguous zeroed run equal to one batch’s append count (16 slots at n=1,840,016; 730 slots == the tail of a single ids page at n=1,840,149), never scattered garbage. VACUUM is not involved.

Fixed by making write_full_inner write the row CHAINS FIRST and the META PAGE LAST (the same crash-safety invariant the build path’s write_blocked_phase_and_meta already documents). An interrupt during the chain writes now leaves block 0 (the meta) UNCHANGED, so readers and the next reconcile-flush observe the previous, fully-valid n_vectors + chains; any pages written past the old count are simply unreferenced until the next full rewrite / RelationTruncate (which stays after the meta write, so a shrink never leaves the meta referencing a page past EOF). The meta write is a single GenericXLog page op with no interrupt poll, so it cannot tear.

Also added a belt-and-suspenders guard on the deferred-flush /reconcile path (reconcile_and_write_flush): before splicing this transaction’s upserts onto the current on-disk ids and re-persisting, it runs first_duplicate_id over the ids it just re-read and, if the on-disk state is ALREADY corrupt (e.g. a hole left by an older binary), ABORTS the transaction with a REINDEX INDEX hint instead of entrenching it — a retryable abort beats propagating on-disk corruption.

A NON-restarting, long-lived single writer (no statement_timeout, no cancels, never pg_terminate_backend’d) never hits the interrupt window and does not trigger this bug — a valid immediate operational workaround.

Validation

  • Fail-before/pass-after unit test torn_flush_on_interrupt_meta_last_stays_clean: injects the exact interrupt window and shows meta-FIRST + tear → id-0/duplicate on reload, meta-LAST (fixed) + tear → clean bijection.
  • Sustained no-VACUUM A/B on the reproduction (1.84M-row 768d bit_width=4 partial flat index, upsert writer, writer restarted via pg_terminate_backend): unpatched corrupts in ~40 s at the first restart; patched runs 2 h+ across dozens of restart cycles with turbovec_check is_corrupt = false throughout AND real forward progress (n_vectors climbs cleanly).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.29.4'; — no REINDEX, no downtime beyond the .so swap + reconnect. An index already corrupted by a pre-1.29.4 binary needs a one-time REINDEX INDEX <name>;.

[1.29.3] — 2026-08-14

Runtime-hardening patch from a full code audit. Patch bump — no wire change (stays v7), no SQL surface change, no REINDEX (ALTER EXTENSION pg_turbovec UPDATE is sufficient; existing indexes are read and written unchanged). All fixes are defensive guards that turn a crash/OOM into a clean ERROR; none change results on valid input.

  • Adversarial-input OOM guard on the sparsevec densify paths. sparsevec permits a dimension up to 1e9, but the ::vector cast (sparsevec_to_vector) and sum(sparsevec) (SparsevecAccum:: ensure_dim) densified into vec![0.0; dim] (up to 4–8 GB) BEFORE the 16000-element vector cap was checked — so one unprivileged SELECT '{}/1000000000'::sparsevec::vector (or sum() over such a row) could OOM a backend. Both paths now reject dim > 16000 before allocating. Test: sparsevec_oversized_dim_densify_errors_not_ooms.
  • Clean ERROR (not a Rust panic across the FFI boundary) on a torn/corrupt scan slot lookup. ReadOnlyIndex::search / search_masked / id_at_slot and the out-of-core IVF id remap indexed slot_to_id[slot] directly; a corrupt/torn read yielding a slot past the id table panicked (a backend-abort risk under load). Now id_at_slot_checked / id_at_global_checked raise a PG ERROR with a REINDEX INDEX hint.
  • VACUUM stays cancellable. The flat swap-remove loop (vacuum.rs) holds the exclusive relfile-rewrite lock across every dead slot; added check_for_interrupts!() at the loop top so a VACUUM deleting millions of rows can be cancelled (and doesn’t hold the lock uncancellably against readers) between slots.

The audit also confirmed the v1.29.2 flat/IVF corruption fixes, determinism gates, dup-id guards, GUC safety, and 1-bit build fencing are solid. It surfaced one CRITICAL follow-up NOT fixed here: the graph-kind incremental INSERT still has the pre-v1.29.2 lost-update window (a non-atomic read-then-rewrite) that reconcile-on- flush closed for the flat/IVF kinds — tracked for a dedicated fix release with a concurrent-inserter reproduction test + sustained-load validation per the corruption HARD MANDATE. Graph inserts should stay serial/one-writer until then (the documented build-then-serve model).

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.29.3'; — no REINDEX, no downtime beyond the .so swap + reconnect.

[1.29.2] — 2026-08-13

Data-corruption fix: concurrent VACUUM + deferred-flush lost-update in the flat/single relfile. Patch bump — no wire change (stays v7), no SQL surface change, no REINDEX to upgrade (existing indexes are read + written correctly in place by the new binary). A running index that already corrupted needs a one-time REINDEX to clear the bad state, but the upgrade itself is ALTER EXTENSION + restart.

Continuous concurrent inserts + autovacuum into a large flat turbovec.vector index corrupted the .tvim id table (“duplicate ids … id 0 appears in more than one slot”, XX001) and SIGABRT-crashed backends. Root cause was a set of concurrency races between the whole-relfile operations, NOT on-disk-format corruption:

  1. Deferred-flush lost-update (the load-bearing bug). Each writer loads the whole index into a per-backend in-memory snapshot on its first mutation and, at PreCommit, rewrote the ENTIRE relfile from that snapshot. A concurrent VACUUM that shrank the relfile (deleted dead rows) in between was clobbered — the committing writer resurrected the deleted rows from its stale pre-VACUUM snapshot, reintroducing dead/duplicate ids. Fixed with reconcile-on-flush: under the exclusive rewrite lock at PreCommit, re-read the CURRENT on-disk (codes, scales, ids) and splice ONLY this transaction’s upserted ids onto it (relfile::reconcile_flush_image / reconcile_and_write_flush), never a blind whole-snapshot overwrite. VACUUM’s deletes and other backends' committed inserts survive.
  2. VACUUM stale-snapshot write. ambulkdelete read the ids chain, computed dead slots, and only THEN took the write lock for the swap-remove — using the pre-lock meta/dead_slots. A flush completing in that window left VACUUM swap-removing against a just-rewritten chain. Fixed by taking the exclusive rewrite lock for the WHOLE bulkdelete (ids read → dead-slot compute → swap → shrink → truncate) on a snapshot re-read under the lock.
  3. read_rotation unlocked window. install_whole_index / install_graph_index read the chains under the shared lock, then read the rotation chain UNLOCKED; a concurrent rewrite could move or truncate it. Fixed by bracketing the whole install (chains + rotation + adjacency + tombstones) under one outer shared lock.
  4. Stale-meta torn reads in read_ids_only (used by turbovec.turbovec_check and the scan visibility path): it read the ids chain with the caller’s pre-lock meta. Fixed by re-reading meta under the shared lock, atomically with the ids.
  5. Hardening (last line of defense): the persist path now refuses to write an id-0 or a duplicate id for the bijective flat/single kind, aborting the transaction rather than persisting the corruption signature.

Proven on an i4i.8xlarge (PG18) against a 1.76M-row 768d flat index: a comprehensive sustained-load run (VACUUM every 3s + 4 writers doing batch-16/128 ON CONFLICT DO UPDATE with fresh-backend-first-insert churn + 3 scanners) that corrupts the pre-fix binary in 60 s (id-0 duplicate, 35 SIGABRT/recovery events) ran 70.5 minutes on the fixed binary with ZERO corruption: 46/46 in-flight turbovec_check reads clean, 0 dup-id, 0 SIGABRT, 0 recovery, ~19,470 write transactions concurrent with VACUUM. Insert throughput regresses ~13% (15→13 rows/s on a 1M-row flat index, single writer) — the reconcile adds a full ids+scales re-read at flush; the dominant per-flush whole-relfile rewrite is unchanged. Warm kNN and build time unaffected (1M build 72 s, warm kNN ~112 ms/query server-side). Repro harness in benches/corruption-repro/.

Migration

ALTER EXTENSION pg_turbovec UPDATE + restart the backend to pick up the new .so. No wire change, no REINDEX for the upgrade. An index that has already corrupted under the old binary should be REINDEXed once to clear the bad on-disk state.

[1.29.1] — 2026-08-11

Packaging fix: ship in-place ALTER EXTENSION UPDATE scripts. Patch bump — no wire change (stays v7), no code change, no REINDEX.

v1.28.4 added turbovec.turbovec_check() to the full-install schema but not to any upgrade path, so operators who ran ALTER EXTENSION pg_turbovec UPDATE TO '1.28.4' in place never got the function the changelog advertised (agora report 2026-08-11). Root cause: the repo shipped only full-install pg_turbovec--<version>.sql and no pg_turbovec--<from>--<to>.sql upgrade scripts, so PostgreSQL’s ALTER EXTENSION UPDATE had no SQL delta to apply — only the .so changed.

  • Adds the runnable upgrade scripts under sql/: the 1.28.3->1.28.4 edge CREATEs turbovec_check; forward edges (1.28.4->1.29.0, 1.29.0->1.29.1) are documented no-ops so the update chain resolves.
  • New scripts/drift-check.sh gate (#12) fails the build if a release lacks its sql/pg_turbovec--<prev>--<this>.sql upgrade edge, so this can never silently regress again.
  • Also codifies the 2026-08-11 hard mandate in AGENTS.md: no corruption ever; non-major upgrades must be zero-format-change or online-upgradable in place (no REINDEX-from-corpus for a minor); even major format breaks must offer an offline in-place converter tool.

[1.29.0] — 2026-08-07

Partitioned-scale support (toward 1T+ vectors) + the 1-bit quantization foundation. Minor bump — purely ADDITIVE SQL surface; no wire-format change (MetaPageData::version stays 7), no REINDEX. ALTER EXTENSION pg_turbovec UPDATE TO '1.29.0'; is sufficient.

Partitioned scale (Phase S-0)

PostgreSQL caps a single heap at 32 TB (~11B rows @768d), so 1T+ vectors must live in a partitioned table. pg_turbovec now documents this directly, and the surprising-but-verified core finding is that PostgreSQL’s native Merge Append over a hash-partitioned parent already does a correct, LAZY global top-k across per-partition turbovec indexes with zero new code — the scatter→gather→merge is byte-identical to a single-table exact top-k (PoC: 10/10 overlap), and the merge pulls ~k + probed rows, not k·N.

  • docs/PARTITIONED_SCALE.md: a cookbook for scaling to 1–10B vectors today with no AM changes — partition sizing (10M–50M/partition), embarrassingly-parallel per-partition build orchestration (the key build-time lever: ~2.2 days at 800-way vs ~47 days naive for 1T), native insert routing, per-partition VACUUM/REINDEX CONCURRENTLY, and the query patterns (native parent ORDER BY emb <=> q LIMIT k, plus the explicit UNION-ALL fallback). PoC at benches/poc/scatter_gather_partitioned_topk.sql.
  • Partition pruning for very large N (Phase S-1, the partition-level coarse quantizer) is designed (doc §6) but deferred to a later release — its first cut didn’t pass its own correctness #[pg_test] when run, and this project does not ship unproven partition-selection code (a silent recall bug is worse than a deferral). It lands alongside the 1-bit completion, real-PG validated. The free scatter-gather above covers 1–10B today.

1-bit (sign binary quantization) foundation

WITH (bit_width = 1) is now accepted by the reloption (range 1..=4; 1 = sign-BQ, 2/¾ = TurboQuant; the GUC default stays 2..=4, so 1-bit is opt-in and never a default). The turbovec kernel hard-rejects bit_width < 2, so 1-bit is a distinct sign-BQ scheme (per-coord sign bit + Hamming coarse + exact heap rerank), matching pgvector/DiskANN/Qdrant. The exact-heap-rerank floor auto-widens for a 1-bit index (reusing hi_dim_rerank/search_k/oversample — no new mechanism). The sign-BQ core is implemented + unit-tested: mean-centering (the mandatory footgun fix — raw sign-at-zero collapses to R@10≈0 on non-zero-centered data), degenerate all-same-sign detection, and half-of-2-bit storage (dim/8 codes/vec, no scale).

Not yet functional end-to-end: a CREATE INDEX ... WITH (bit_width = 1) currently ERRORs clearly (“not yet implemented; use bit_width = 2, 3, or 4”) — it never half-builds or ships a silent all-ones landmine. The end-to-end 1-bit encode/scan path requires a wire bump to v8 and lands in a later release after full real-PG validation of the new scan kernel + wire format (deliberately not rushed into this release). The reloption is accepted now for forward compatibility.

[1.28.4] — 2026-08-07

Corruption fix: eliminate the dual row-counter drift that could persist a .tvim id table claiming more rows than it holds, plus a new turbovec.turbovec_check(regclass) integrity function. Patch bump — no wire-format change (stays v7), no REINDEX required by the upgrade itself (a currently-corrupt index still needs REINDEX or DROP + CREATE; see Migration).

Reported 2026-08-06 (agora / pg.ddx.io) on a 1.74M-row PG18.4 index: the .tvim id table developed “duplicate ids (id 0 appears in more than one slot),” write-blocking the whole indexed table, and REINDEX didn’t durably repair it (only DROP + CREATE held). This release fixes the root cause on the write path and adds an operator-facing way to detect the corruption without attempting a write.

B (root cause) — single source of truth for the persisted row count

The deferred aminsert flush (xact::flush_to_relfile) persisted PersistState.n_vectors — a SEPARATELY-incremented i64 counter — as the on-disk row count, passed to relfile::write_full_with_prepared as an INDEPENDENT argument alongside idx.slot_to_id(). Those two counters can drift (an IdAlreadyPresent remove+re-add, a mark_dirty closure that didn’t run, a future insert-path refactor). If n_vectors ever exceeded slot_to_id.len(), the meta page claimed more rows than the ids chain held, and reload (read_full) over-read the ids chain into zeroed trailing slots — surfacing as “id 0 appears in more than one slot” (a real CTID never encodes to 0, so an id-0 slot is always zeroed bytes).

  • The flush now DERIVES the persisted count from idx.slot_to_id().len() (the authoritative array whose bytes actually land in the ids chain) via the new pure, unit-tested xact::reconciled_row_count, making the drift class structurally impossible on the write path.
  • A hard runtime guard aborts the transaction (the matching Abort callback then evicts the dirty cache entry, so the next access reloads clean committed state) if the mirror ever drifts — belt-and-suspenders with the pre-existing assert_eq! in relfile::write_full_inner (now documented as a hard release-build persist-site guard that must NOT be downgraded to debug_assert!).
  • PersistState.n_vectors is retained only as a mirror (and the scan-visibility snapshot); the on-disk truth is slot_to_id.len().

A (recovery) — REINDEX now durably repairs

With B fixed, a REINDEX (which allocates a fresh relfilenode; the per-backend cache drops the stale entry on the relfilenode mismatch in am_lookup_for_mutation) produces a clean id bijection that the write path can no longer re-corrupt. The pre-1.28.4 “REINDEX reports success but the corruption returns minutes later” behavior was the ongoing insert workload re-corrupting the freshly rebuilt index via the drift above (or a crash — see C); B removes the write-path source.

C (crash-safety) — KNOWN REMAINING GAP, scoped honestly

This release does NOT make the .tvim id table fully WAL-crash-safe. An unclean shutdown / pg_resetwal can still discard WAL that extended the id chain, leaving two slots claiming id 0. Full WAL-crash-safety of the id table is a larger durability change, deliberately not attempted inside a corruption patch (shipping a half-done risky durability change would be worse than the documented gap). Mitigations in place: the insert AND read paths already ERROR loudly on such a relfile with HINT: REINDEX INDEX <name>; (v1.28.2), and turbovec_check() (below) now makes it detectable without a write. Recovery: REINDEX INDEX <name>; (now durable per A), or DROP + CREATE for an index corrupted before 1.28.4.

D (integrity check) — turbovec.turbovec_check(regclass)

New read-only, ownership-checked function that reads the meta + ids chain and reports enough for an operator to detect the corruption WITHOUT attempting a write (the only signal pre-1.28.4 detection gave, which blocks the whole table):

SELECT * FROM turbovec.turbovec_check('my_idx'::regclass);
--  wire_version | kind   | n_vectors | slot_count | count_matches
--  duplicate_id | is_corrupt | tombstone_density

is_corrupt (true when a flat index has a duplicate id OR the counts disagree) is the column monitoring should alert on. Reuses scan::first_duplicate_id. Takes only AccessShareLock, so it never blocks writers; non-owners get a permission-denied ERROR (portable across PG13-19 via pg_class_ownercheck / object_ownercheck).

Tests

  • xact::reconcile_tests — load-independent pure-Rust drift-guard unit tests (run under cargo test --lib, no cluster needed): the regression gate for B.
  • persist_large_batch_no_duplicate_id_after_reload (#[pg_test]) — builds a populated flat index, adds a 128-row batch of new ids via the same path aminsert uses, flushes, and re-reads the persisted relfile asserting no duplicate id and meta.n_vectors == slot_to_id.len().
  • persist_row_count_drift_aborts (#[pg_test]) — a deliberately drifted PersistState must abort the flush, never persist a corrupt relfile.
  • turbovec_check_reports_healthy_flat_index (#[pg_test]) — D.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.28.4'; is sufficient and cannot fail on existing indexes. No wire change (stays v7), no REINDEX required by the upgrade. A currently-corrupt index still needs recovery: REINDEX INDEX <name>; (now durable) or DROP + CREATE. The fix STOPS new write-path corruption; it does not repair an index already corrupted by a prior crash.

[1.28.3] — 2026-07-31

Managed-PostgreSQL readiness (audit follow-up): interrupt handling + portability gates + deployment docs. Patch bump — no wire-format change (stays v7), no SQL-surface change, no REINDEX.

Addresses a managed/hosted-PostgreSQL readiness audit of v1.28.2.

  • P0 — interrupt handling. The extension had zero CHECK_FOR_INTERRUPTS call sites, so long scans/builds/flushes were uninterruptible and statement_timeout / pg_cancel_backend() (and the platform admin signals delivered by the same mechanism) were ignored for seconds. Added polling at every point pg_turbovec controls where no buffer content lock is held:
    • each amgettuple entry (makes the iterative-scan refill loop promptly cancellable),
    • each CREATE INDEX build-stage boundary (k-means / assign-sweep / quantize-encode / prepare-and-persist),
    • the top of each GenericXLog batch in the relfile write path (large aminsert PreCommit flushes). Known limitation: a single large flat-kind search() call is still internally uninterruptible (one call into the Postgres-free kernel crate). Finer-grained polling there needs a kernel block-scoring API and is tracked follow-up; prefer the IVF/graph kinds for large corpora meanwhile.
  • P2 — portability drift-check gates. scripts/drift-check.sh now fails the build if (11a) any raw-WAL / raw-storage primitive appears at a call site (log_newpage, XLogInsert, smgrwrite, smgrextend, PageSetLSN, FlushRelationBuffers, …) — durability must stay 100% GenericXLog; or (11b) any GUC is registered with a non-Userset context. Both invariants hold today; the gate stops a future commit from silently reintroducing a managed-adoption blocker.
  • P2 — docs. New docs/DEPLOYING_ON_MANAGED_POSTGRES.md consolidating durability/replication, restricted-superuser compatibility, read-replica behavior, cancellation, the memory / O(n)-per-transaction insert model, build cost, and the upgrade / wire-format policy.

Deferred (tracked, larger scope): making the in-place index rewrite atomic (write the meta page last, as the single commit point; today it self-heals via the xs_recheckorderby heap backstop); a full out-of-process TAP suite (crash recovery / replication / failover / cancellation-latency / codec fuzz); a kernel block-scoring API for fine-grained flat-scan cancellation. See the audit for detail.

[1.28.2] — 2026-07-31

Bug fix: detect duplicate-id corrupt .tvim relfiles and fail loudly with a REINDEX hint, instead of silently mis-serving reads while failing every write. Patch bump — no wire-format change (stays v7), no SQL-surface change, no REINDEX required by the upgrade itself (a corrupt index does need one; see below).

Reported 2026-07-30 (agora / pg.ddx.io): a 1.6M-row index on PG18.4 rejected 100% of INSERTs for ~1.5 days (~47k logged failures) with an opaque aminsert: corrupt relfile pages: duplicate ids in .tvim file, while pg_index.indisvalid stayed true and scans kept answering — so the corruption was invisible to health checks and a backfill retry-loop turned it into an availability incident. The likely origin was several crash-recovery / pg_resetwal cycles that left the on-disk id table with duplicate ids.

Root cause was an asymmetry: the write path validated the id table is a bijection (turbovec’s from_id_map_parts rejects duplicates) and hard-errored, but the read path (ReadOnlyIndex::from_parts, which only builds slot_to_id) never checked — so a corrupt index failed writes yet still served (duplicated / mis-ranked) scan results and looked valid.

Fix: - The read/open path (scan::assert_ids_unique_or_reindex, called in the whole-index, graph, and out-of-core cache-install paths) now detects duplicate ids once per backend at cache-install and ERRORs (ERRCODE_DATA_CORRUPTED) with HINT: REINDEX INDEX <name>; — the same actionable shape as the pre-v7 legacy gate. A corrupt index now fails loudly on reads too, rather than silently mis-serving. - The insert-path error carries the same REINDEX hint, so a backfill loop gets a clear “rebuild me” signal instead of retrying an opaque error forever. - Recovery: REINDEX INDEX <name>; (or ... CONCURRENTLY) rebuilds a clean id bijection from the heap. There is no lighter in-place dedup — a rebuild is the supported repair.

Gate: scan::duplicate_id_tests::detects_duplicate_ids (pure-Rust) covers the bijection predicate incl. the exact reported id.

Still open (tracked, larger scope): the .tvim id table is not fully crash-safe against pg_resetwal / unclean shutdown — WAL-logging it well enough to recover cleanly (or refusing to come up valid when it can’t) is the deeper durability fix and is deferred to a future release. v1.28.2 makes the corruption detectable and actionable; it does not yet prevent it from being introduced.

[1.28.1] — 2026-07-28

Nix flake packaging. Patch bump — packaging only; no code change, no wire change (stays v7), no SQL-surface change, no REINDEX.

A customer installing on PostgreSQL 18 via Nix found the v1.28.0 tag had no flake.nix (“no Nix output”). This release adds one:

  • nix build github:gburd/pg_turbovec#pg_turbovec_NN for NN in 13–19 (default = PG18), built with nixpkgs' buildPgrxExtension + a flake-local cargo-pgrx 0.19.1 (nixpkgs' pinned set stops at 0.18.x). The turbovec git dependency’s vendor hash is pinned via cargoLock.outputHashes.
  • Output layout: lib/pg_turbovec.so + share/postgresql/extension/pg_turbovec--<ver>.sql + .control. On PG18+ a stock server can use it in place via extension_control_path / dynamic_library_path (validated: CREATE EXTENSION + distance fns + a turbovec index scan against nixpkgs' postgresql_18, served straight from the store path).
  • nix develop: dev shell with rust, cargo-pgrx 0.19.1, clang/bindgen, openblas, and the PG-from-source deps (bison/flex/readline/zlib/icu).
  • The #[pg_test] suite stays CI’s job (needs a live cluster); the flake package build itself runs no tests (doCheck = false).

[1.28.0] — 2026-07-28

PostgreSQL 19 (beta1) support via the pgrx 0.17 → 0.19.1 upgrade. Minor bump — framework/platform change; no wire-format change (stays v7), no SQL-surface change, no REINDEX. Existing indexes and queries are unaffected; ALTER EXTENSION pg_turbovec UPDATE suffices.

  • New: pg19 Cargo feature + CI matrix leg. PG19 is upstream beta (19beta1); support is experimental until PG19 RC/GA, at which point it will be re-validated.
  • pgrx 0.19.1 (from 0.17.0, a two-major jump):
    • The pgrx_embed second-compilation-pass model is gone in pgrx 0.18+ — deleted src/bin/pgrx_embed.rs and the [[bin]] stanza; SQL entity metadata now lives in the .so’s own linker section.
    • Rust edition 2024, MSRV 1.96 (was 2021 / 1.85). The edition-2024 unsafe_op_in_unsafe_fn lint is allowed crate-wide: the index-AM callbacks are unsafe extern FFI boundaries whose entire bodies are unsafe by construction; the per-fn # Safety contracts remain the audit surface.
    • cargo-pgrx 0.19.1 required (cargo install --locked cargo-pgrx --version 0.19.1).
  • PG19 C-API deltas handled (all version-gated; pg13–18 builds byte-identically unaffected):
    • LockBuffer’s mode parameter changed from int (i32) to BufferLockMode::Type (u32) — new relfile::lock_buffer_mode shim used at all 5 call sites.
    • relfilenode_from_relation gained the pg19 arm (same rd_locator.relNumber layout as pg16+; the cfg just didn’t cover pg19, leaving the fn body empty on pg19).
  • Verified: all 7 versions (pg13–pg19) compile with 0 errors and 0 warnings; pure-Rust determinism gates (reservoir, IVF k-means, build-pool invariance) green on pg13 and pg19; the full #[pg_test] suite gate runs in CI across the 7-leg matrix.

[1.27.3] — 2026-07-12

Phase Q-4c: clear the IVF build cliff — batch the k-means reservoir rotation into one parallel GEMM. Patch bump — build-SPEED change only. The persisted IVF centroids + codes are byte-identical to v1.27.2 for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, no REINDEX.

Profiling a real 1M×1024/lists=4096 build on a 32-vCPU AVX-512 host showed the v1.27.2 lazy-per-row-rotation was still ~85% of a ~14.5-min build: the scalar O(dim²) rotate_unit ran ~600k times (reservoir fill + replacements) single-threaded, and at dim=1024 each rotation is ~1M FLOPs. Reducing the call count (v1.27.2) wasn’t enough — the per-row scalar rotation was the wrong primitive.

The fix: the reservoir now stores normalised-but-unrotated rows and rotates the whole training sample once at drain via ivf::rotate_corpus_into — a parallel BLAS GEMM that already exists and is already gated byte-identical (rotate_corpus_bit_identical_across_pool_sizes). The per-row hot path is now just the O(dim) normalise; the O(dim²) work is one GEMM over the ≤256·lists kept rows across all cores. Byte-identical output: same rows selected (RNG/seen/cap unchanged, value-independent), same rotation math (GEMM corpus @ R^T == scalar unit @ R^T).

Measured (real corpus, 32-vCPU AVX-512, PG17.10): - 1M×1024, lists=1024: ~870s → 138s (~6.3×), identical 540 MB index. - 10M×1024, lists=4096: 1562s (~26 min), 5343 MB — previously did NOT complete (DNF, cancelled at 33.7 min / 16%). The build cliff is cleared; the build now scales roughly linearly (per-stage trace shows no single dominant stage). - Query correctness verified at both scales (self-top-1 returns the query id).

Also in this release (verified on the same host): the v1.27.2 G3 concurrency fix holds — the 1M in-RAM kNN throughput sweep rises monotonically and plateaus at ~40 TPS through 64 connections (no collapse), where the pre-fix per-query-pool-churn build collapsed from 46 TPS @16 conns to 1.7 TPS @32 conns. Latency saturates gracefully (797 ms @32 → 1590 ms @64) instead of thrashing.

Gate: reservoir_tests::deferred_batch_rotation_is_byte_identical_to_eager (pure-Rust) proves the store-unrotated-then-batch-rotate sample matches the original rotate-every-row path; the IVF determinism + pool- invariance gates are unchanged and green.

[1.27.2] — 2026-07-11

Phase Q-4b: kill the IVF build cliff’s real bottleneck — the per-row rotation in k-means reservoir sampling. Patch bump — build-SPEED change only. The persisted IVF centroids + codes are byte-identical to v1.27.1 for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, no REINDEX.

Profiling the real 10M×1024/lists=4096 build (v1.27.1 still did not complete in budget despite Q-4a’s parallel k-means) found the dominant cost was not k-means (~4.4%, flat in n) but the single-threaded scalar rotate_unit inside BuildState::ivf_reservoir_push — ~92% of the projected build. It ran an O(dim²) rotation on all N accepted rows during the heap scan, even though the reservoir only ever keeps cap = 256·lists samples. At 10M/lists=4096 that is ~9M O(dim²) rotations computed and immediately discarded.

The fix defers the rotation: the reservoir selection uses only the RNG stream + seen-count + cap (none depend on the rotated values), so rotate_unit now runs only for the ≤ cap rows that actually land in the sample. Rotation cost drops from O(N·dim²) to O(cap·dim²), removing the n-scaling that was the cliff — while producing the identical k-means training sample (hence identical centroids, identical codes). The RNG draw stays unconditional in the replacement branch, exactly as before, so the same source rows land in the same reservoir slots and receive the same rotation.

Verified: index::build::reservoir_tests:: lazy_rotation_is_byte_identical_to_eager reproduces both the old (eager-rotate-always) and new (lazy-rotate-on-keep) reservoir algorithms with a shared ChaCha8 seed across 4 corpus sizes × 3 seeds (exercising both the fill and the replacement paths) and asserts the sample buffers are bit-identical (f32::to_bits). Pure-Rust, load-independent, green. The full cargo pgrx test pg16 gate (including ivf_streaming_build_determinism_byte_identical) must be re-run on a quiet host before tagging — the local box was under load ~73 at authoring time.

[1.27.1] — 2026-07-11

Phase Q-4a: parallelize the IVF k-means build. Patch bump — build-SPEED change only. The persisted IVF centroids + assignment are byte-identical to v1.27.0 for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, no REINDEX. This only makes NEW WITH (lists = N) builds faster.

Q-1 exposed the IVF build cliff as the blocker for scale: a real 10M×1024 IVF build did not complete in a 90-min budget. Two remaining serial hot loops in the k-means build are now parallelized bit-identically (the rest was already parallel from v1.20.0/v1.22.1): - gemm_lloyd_assign’s per-row argmin + top-2 tie-break — run every Lloyd iteration over the whole sample — is now a data-parallel par_iter_mut().enumerate() writing each assign[i] to a fixed index, so the result is byte-identical regardless of thread count. - rotate_corpus_into’s GEMM — Parallelism::None → Rayon(0), on the same “gemm tiling never reduces across threads” guarantee v1.22.1 established for the Lloyd cross-term GEMM. Left serial deliberately: the k-means++ next-seed CDF pick (sequential by construction) and the empty-cell reseed scan (data-dependent mutation, not a parallel-safe map).

Measured ~1.91× build speedup (lists=4096, n=131072, dim=256, 8-core, release: 418.2s → 219.4s) with bit-identical output. Sub-linear because Lloyd iterations are sequentially dependent (only within-iteration work parallelizes) — the honest ceiling for k-means, unlike the embarrassingly-parallel graph build (v1.26.0). Verified: ivf_coarse_model_bit_identical_across_pool_sizes + rotate_corpus_bit_identical_across_pool_sizes assert byte-identical centroids + assignment across pool sizes {1, 2, auto}; kmeans_deterministic_across_pool_sizes stays green; full suite 338/338.

This is the first step toward the scale-and-heavy-load build goals; the empty-cell/reseed serial remainder and a larger-scale (10M→100M) build validation are the follow-ups.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.27.1'; — a no-op upgrade for existing indexes; new IVF builds are just faster.

[1.27.0] — 2026-07-10

Phase Q-0: de-duplicate the on-disk quantized-codes storage, roughly halving the per-vector index footprint. Minor bump (wire-format change — REINDEX required; additive capability, no SQL-surface removal). This clears the storage blocker for large single-node indexes.

The problem. Every prior version persisted each vector’s quantized codes TWICE: the row-major bit-plane packed_codes chain AND the SIMD-blocked chain (the output of pack::repack(packed_codes, …)). That doubled the dominant O(n) storage term. At 100M vectors that is the difference between ~78 GB and ~39.6 GB for 768d/4-bit (double- storage vs single).

The fix (Option A — persist only the packed codes). The blocked layout is a PURE FUNCTION of the packed codes, so v7 drops the blocked chain entirely from disk and recomputes it once per backend at index-open via pack::repack. This is the same O(n) one-time compute a pre-Phase-P index already paid lazily on first scan; it’s paid once per (backend, am_version) at cache-install and cached in the per-backend ReadOnlyIndex, so warm per-query latency is unchanged and scan results (recall, ordering) are bit-identical to before (the recomputed blocked layout equals the layout that used to be persisted). Option A was chosen over Option B (persist blocked, recompute packed via the new pack::unblock) because repack is the forward, already-used direction and the OOC path never touches the blocked chain — A keeps every hot path simplest.

Measured storage reduction (phase_q0_storage_is_deduplicated

[pg_test], 512×128d 4-bit): the dropped blocked chain is ≥ the packed

codes chain, i.e. persisting it doubled the code-storage term. Per-vector on-disk code bytes: dim/8 * bit_width stored ONCE (was twice). 100M projections: 768d/2-bit 19.8 GB (was 39.6), 768d/4-bit 39.6 GB (was 78), 1536d/2-bit 39.6 GB (was 78).

Wire format (v6 → v7), NOT additive. Unlike the additive v4→v5→v6 per-kind bumps, dropping a persisted chain is a real break for EVERY kind: single-vector, ColBERT, IVF, and graph indexes all now emit wire version 7 (the kind byte still discriminates). A pre-v7 index (v1..v6) is detected by the new MetaPageData::is_legacy_v6() predicate (version < 7); ambeginscan raises a clear ERROR that NAMES the index with a HINT: REINDEX INDEX <name>; at the first scan — never silent corruption. Verified by ambeginscan_errors_on_legacy_v6_meta (and the retained v1/v2 forgeries, which now also hit the unified v7 gate).

No SQL surface change (no operators/types/functions/GUCs/opclasses added or removed). All index kinds (flat, IVF, graph, ColBERT) continue to work.

Migration: 1. ALTER EXTENSION pg_turbovec UPDATE TO '1.27.0'; 2. REINDEX INDEX <name>; — once per turbovec index, ANY kind.

Until an index is REINDEXed under v7, scans against it ERROR with the REINDEX hint (they do NOT return wrong results). See docs/UPGRADING.md.

[1.26.0] — 2026-07-10

Phase G-2d(a): a partitioned/merge parallel build for the graph index kind, so it scales past the single-pass build’s serial ceiling. Minor bump (new GUC turbovec.graph_build_partitions, new build path, additive). NO wire-format change (stays v6, byte-identical on-disk CSR), no new operators/types/functions, no REINDEX.

The single-pass Vamana build is serial by necessity (each insertion navigates the graph every prior insertion left; G-2c showed thread-parallelizing it doesn’t amortize) and did not complete at 5M rows (>2h26m). This adds a structurally-parallel build: partition the corpus into P shards (contiguous ranges of the deterministic shuffled insertion order — each shard a uniform random sample), build each shard’s sub-graph in parallel across the bounded rayon pool, then stitch via a parallel cross-shard refinement pass (greedy-search the merged graph from a global medoid entry + RobustPrune per node) and a deterministic parallel reverse-edge pass. No explicit persisted bridge edges are needed — the refinement pass creates cross-shard navigability, so the frozen v6 CSR is sufficient.

Measured (verified, not just claimed): - Recall parity — partitioned MATCHES or BEATS single-pass. In a findable high-recall regime (queries drawn from the corpus, 20k×64): single-pass R@10=0.958, partitioned (P=8) R@10=0.996 (+0.038). In a cross-cluster regime: 0.630 → 0.754 (+0.124). The refinement pass’s greedy search over the merged graph surfaces better cross-shard neighbours than incremental insertion, so the partitioned graph is HIGHER quality, not a speed-for-recall trade. - ~8× parallel build speedup (relative serial-vs-partitioned wall-clock, 8-core AVX2 box, 200k×64): P=16 ≈ 7.99×, P=8 ≈ 6.45×. Both the per-shard builds and the refinement/reverse passes parallelize — unlike the single-pass build. - Deterministic: bit-identical for a fixed (corpus, seed, P) AND across rayon pool sizes {1, 2, auto} (partition is a pure function of the seed; every parallel stage uses an index-ordered collect / staged fixed-id writes).

turbovec.graph_build_partitions (int, default auto): auto derives P from corpus size + build-pool budget (single-pass below a size threshold); 0/1 forces the single-pass reference build; N forces N shards. The single-pass build_vamana is retained intact (all its tests pass) and is what P<=1 runs — verified identical.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.26.0';. No REINDEX — existing v6 graph indexes are unaffected; the parallel build only changes how NEW WITH (graph = true) indexes are constructed, to the identical on-disk shape. The 5M/10M a cloud VM gate re-run (now that the build completes at scale) is the follow-up this unblocks.

[1.25.1] — 2026-07-09

Release-tooling + docs/benchmark patch. No shippable code change — the compiled binary is byte-identical to v1.25.0 (no src/ runtime change; wire format stays v6; no SQL-surface change; no REINDEX). This is the release that first exercises the new automated publish pipeline.

  • Automated release publishing (.forgejo/workflows/release.yml, scripts/make-dist.sh, META.json.in, ci/announce.sh): a vX.Y.Z tag on Codeberg now gates (compile + drift-check), builds a PGXN source distribution (renders META.json from Cargo.toml’s version, runs cargo pgrx schema for the install SQL, zips a PGXN-layout archive), attaches it to a Codeberg release, uploads to PGXN, and submits a postgresql.org news announcement (feeds pgsql-announce). Runs on a self-hosted Forgejo runner (Codeberg’s hosted 10-min cap can’t fit a pgrx build); each publish step no-ops cleanly if its secrets are unset. The PGXN dist is a source archive built with cargo pgrx install, not a pgxn install-able package — published for discoverability + version-pinning. See RELEASING.md for the one-time runner + secrets setup.
  • Qdrant + ANN-Benchmarks-protocol competitive benchmark (benches/results/qdrant_annbench_20260709/): validated v1.25.0’s turbovec.hi_dim_rerank at real scale — on GIST-960-1M, auto lifts R@10 from 0.876 (the off ceiling) to 0.953, crossing the ≥0.90 and ≥0.95 bands the pre-fix engine never reached. vs Qdrant (in-RAM, a benchmark host, 1M): Qdrant wins raw latency 3–18×, pg_turbovec wins storage (SIFT 141 MB = 7–8× smaller; GIST 983 MB = 5.4× smaller than Qdrant’s 5.3 GB). At 10M×960 only Qdrant built in a 90-min budget (turbovec’s single-threaded k-means build cliff).
  • Docs correction: the 2026-07-08 competitive doc’s pg_turbovec SIFT R@10=0.99 latency (16.84 ms) was the out-of-core path; the in-RAM number is 2.6 ms (verified R@10=0.9915 @ 2.61 ms). Footnoted in place.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.25.1'; — a no-op upgrade (nothing in the SQL surface or on-disk format changed).

[1.25.0] — 2026-07-09

Gap-B fix: turbovec.hi_dim_rerank, a dimension-aware exact-L2 rerank-window widening that recovers high-dimensional recall. Minor bump (one new GUC, additive). No wire-format change (stays v6, byte-identical to v1.24.0), no new operators/types/functions, no REINDEX.

An offline investigation (using FAISS’s trusted quantizers as measurement vehicles) established that the high-dim recall gap — e.g. GIST-1M/960d capping ~0.86 where pgvector HNSW and VectorChord reach 0.95-0.98 — is NOT retrieval- bound. The true nearest neighbours DO land in the probed IVF cells (measured cell recall 0.978 at probes=64, 0.996 at probes=128). The real bottleneck is in-cell quantized ranking: at high dim the lossy 4-bit score is noisy enough that a true neighbour often sits at rank ~200-800 within the probed cells, below a small search_k, so it never enters the exact-L2 reorder recheck. This corrects the prior “retrieval-recall ceiling” framing (corrections noted in-place).

The cure is scan-side only: fetch a wider candidate set so the always-on exact-L2 reorder queue (xs_recheckorderby) re-ranks enough survivors to recover the true top-k. Measured: an SQ4 analog of TurboQuant’s per-coordinate scalar quant lifts R@10 from 0.666 to 0.978 at 960-dim by reranking ~800 candidates instead of ~64.

This must not blanket-widen: at low dim recall already plateaus by search_k≈25, so widening there is pure latency tax. Hence a dimension-aware floor. turbovec.hi_dim_rerank: - auto (default): apply a candidate floor of clamp(dim, 256..=1024) only for indexes with dim >= 256, and only ever RAISE the count — an explicit search_k/oversample override past the floor always wins. SIFT-128 is untouched (zero latency cost); GIST-960, OpenAI-1536, and the 512-768d embedding families get the wider window out of the box. - on: apply the floor regardless of dim. - off: honour search_k/oversample exactly (pre-1.25.0 behaviour).

The result set is identical to setting search_k/oversample by hand to the same candidate count — this is a smarter default, not a new mechanism. Verified end-to-end by the hi_dim_rerank_raises_high_dim_recall #[pg_test] (384-dim corpus: auto measurably beats off, never regresses any query) plus the hi_dim_rerank_tests unit tests on the dim-scaling decision function.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.25.0';. No REINDEX. The new auto default improves high-dim recall out of the box at a small high-dim-only latency cost; SET turbovec.hi_dim_rerank = off restores exact pre-1.25.0 scan-candidate behaviour.

[1.24.0] — 2026-07-08

Phase G-2b: VACUUM + incremental INSERT for the graph index kind. Both operations previously raised a clear ERROR against a WITH (graph = true) index (v1.23.0 was build+scan only, correctness- first); this release turns them into working functionality. No wire- format change — wire format stays v6, existing v4/v5/v6 indexes decode byte-identical, no REINDEX. Minor bump because a previously-ERRORing operation becomes real capability (same reasoning v1.23.0 used to justify G-2a as a minor).

VACUUM (ambulkdelete) now uses the same per-slot tombstone bitmap mechanism IVF already uses (the generic relfile::read_tombstones / write_tombstones_and_meta path, confirmed to have zero IVF-specific assumptions). Tombstoned nodes never enter scan results and their out-edges are never followed; a fully-tombstoned corpus returns empty cleanly. graph_search gained a tombstones parameter threaded through the scan path.

INSERT (aminsert) does a whole-relfile rewrite per insert (read every chain back, quantize+append the new vector, run a Vamana insertion via insert_one_node_via_oracle, persist). This is a deliberate O(n)-per-insert cost — explicitly NOT the deferred/batched path other kinds get — appropriate for the build-then-serve model the graph kind targets; heavy incremental churn should still REINDEX. The insertion distance oracle uses the exact quantized-code scan kernel for dist(new, existing) and a triangle-inequality lower bound for the dist(existing, existing) RobustPrune diversity check (documented safe-direction approximation: can only make pruning less aggressive, never drops a genuinely diverse edge, still hard-capped at degree R).

Two real bugs found and fixed during G-2b’s own test-writing:

  1. Relfile corruption on insert-after-VACUUM. write_tombstones_and_meta’s block-offset formula for placing a new tombstone chain omitted + graph_count, so a graph index’s tombstone chain (once VACUUM wrote one) computed an offset that collided with the already-persisted graph adjacency chain, corrupting it on the next incremental insert (corrupt graph adjacency chain: graph offsets[n]=0 != neighbors.len()=...). Root-caused via instrumentation showing graph_first == tombstone_first. graph_count is 0 for every non-graph kind, so the fix is a no-op for flat/IVF/ColBERT. A second, related fix re-persists a pre-existing tombstone bitmap after the main graph write (which plans a fresh meta from scratch and would otherwise silently drop it).

  2. VACUUM entry-point dead-end. The entry-point fallback fired only when the entry point itself was tombstoned, never when the entry point SURVIVED but every one of its out-neighbors was tombstoned — a low-degree entry point whose only edge points at a now-dead slot is a genuine dead-end (first beam-search hop expands to zero live candidates). Both the VACUUM-side fallback selection and the scan-side entry pick now treat “no live out-neighbor” as equally disqualifying as “dead” and prefer a fallback that itself has a live neighbor.

Also corrected a test-harness bug (not a shipped-code bug): the graph #[pg_test] corpora were generated with an uncorrelated random() subquery that PostgreSQL hoisted and evaluated once, making every test row identical (n_distinct = 1) and producing spurious “recall collapse” numbers that had nothing to do with the feature under test. Correlating the inner generate_series to the outer row restored genuine per-row randomness; the insert/vacuum/quantization paths were correct all along.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.24.0'; only. No REINDEX, no wire change, no SQL-surface change. A v1.23.0 graph index that was built and never mutated is unaffected; graph indexes can now be VACUUMed and incrementally inserted into in place.

Still deferred (unchanged from v1.23.0): G-2c (SIMD traversal + build parallelism), G-2d (the 5M-scale AVX2 HNSW-latency gate).

[1.23.0] — 2026-07-06

Phase G-2a: WITH (graph = true), a new Vamana-style navigable- graph index kind — the first step toward matching HNSW’s query latency while keeping TurboQuant’s storage compression, per the long-standing roadmap.

Builds a real single-pass Vamana graph (DiskANN’s algorithm: greedy search + RobustPrune per node, in a deterministic randomized insertion order) over the full corpus. A graph node’s vector storage is identical to a flat index’s (same TurboQuant encode path); the new adjacency chain (CSR: offsets + flat neighbor ids) and an entry-point slot id are the only new on-disk structures. Scan is a greedy beam search from the entry point, feeding results through the existing xs_recheckorderby machinery every other kind already uses — no new correctness surface there.

Wire format v6, additive (same pattern as v5’s KIND_COLBERT): existing v4 (single-vector) and v5 (ColBERT) indexes decode byte-identical under the v6 binary — verified by dedicated tests. No REINDEX for any existing index. A graph index is a brand-new shape only a v6 binary produces.

Determinism: relaxed for this kind only, per the plan doc’s explicit “three fundamental tensions” framing — deterministic for a fixed seed on one machine/thread-count (required for the test suite and for REINDEX reproducibility on a given host), but NOT byte-identical across machines/ISAs the way flat/IVF/ColBERT indexes are. WAL/streaming replication are unaffected either way (replicas replay the primary’s actual page bytes, never rebuild independently).

Scope of this release (G-2a, correctness-first — see for the full sub-phase breakdown): build + scan work end to end with real, verified recall against an exact linear scan on test corpora. Explicitly NOT yet done, each tracked as a numbered follow-up: - G-2b: VACUUM/tombstone integration. ambulkdelete and aminsert against a graph index currently raise a clear ERROR rather than silently corrupting the graph — rebuild the index after bulk loading/deleting, the same operational model IVF’s original build-then-query-only phase used before tombstones landed. - G-2c: SIMD-optimized traversal + build parallelism. This release’s distance computation is plain scalar Rust; the build is correctness-first, not speed-optimized. - G-2d: the real 5M-scale, AVX2-hardware HNSW-latency gate measurement this whole feature is ultimately judged against (the gate: p50 ≤ 1.3× HNSW AND storage ≥ 6× smaller AND recall ≥ IVF at matched budget, at 5M rows/ R@0.96 on AVX2). Not run in this release — no latency or recall-vs-HNSW claim is made here; that measurement is the honest next step before this feature can be called production-ready, and the plan doc’s own framing (“the bet failed, keep IVF” is a real possible outcome) still stands.

New reloption WITH (graph = true), mutually exclusive with WITH (lists = N) on the same index (enforced in amoptions). 24 new tests (287 total, up from 263), drift-check/compile-matrix/fmt clean, CI green.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.23.0'; is sufficient. No REINDEX for existing indexes.

[1.22.2] — 2026-07-06

Raises turbovec.probes’s default from 8 to 16 — the out-of-the- box recall floor was unreasonably low. Scan-side default change only, no wire change, no SQL surface change, no REINDEX.

The v1.22.1 a cloud VM competitive re-benchmark measured pg_turbovec’s shipped defaults (probes=8, search_k=32) capping at R@10=0.796 on SIFT-1M and R@10=0.407 on GIST-1M — both far below any reasonable recall SLO, and a real footgun for anyone who runs CREATE INDEX ... USING turbovec without reading the tuning docs. probes=16 (same search_k=32) measured:

Corpus probes=8 (old default) probes=16 (new default)
SIFT-1M R@10=0.796, p50=3.0ms R@10=0.918, p50=4.8ms
GIST-1M R@10=0.407, p50=18.5ms R@10=0.557, p50=20.4ms

Roughly 1.5-1.6× the latency for +12-15 recall points on both corpora — the better point on the curve than probes=32 (which roughly triples latency for a similar recall gain). Existing sessions/deployments that explicitly SET turbovec.probes are unaffected; this only changes the compiled-in default.

New regression test index_am_probes_defaults_to_16 guards the default against silent drift (matches the index_am_iterative_scan_ defaults_to_off precedent from v1.20.1).

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.22.2'; is sufficient. No REINDEX.

[1.22.1] — 2026-07-05

Closes a real fraction of the IVF build-cliff gap — scan/build-path only, no wire change, no SQL surface change, no REINDEX.

v1.20.0/v1.21.0’s parallel-build work row-blocked the normalize/rotate/assign-sweep stages of ambuild, measuring only a modest ~1.27× speedup on a 64-core box. A FLOPs analysis (triggered by the v1.22.0 GUC audit) found the real dominant cost was never row-blocked: gemm_lloyd_assign’s cross-term GEMM runs over the whole k-means training sample (n_sample = lists × 256, which equals the full corpus size at high lists) once per Lloyd iteration, up to 25 times, single-threaded (Parallelism::None, kept that way for on-disk determinism). At GIST-1M/960d/lists=4096 scale this GEMM is ~26-112× more FLOPs than the already-parallelized stages — explaining why the earlier fix barely moved the needle.

The fix: gemm 0.18’s own internal Parallelism::Rayon(n) tiling produces bit-identical output to Parallelism::None for every shape/seed/thread-count tested — a GEMM’s output tiles are independent dot-product reductions over the shared contraction dimension, so (unlike a cross-thread SUM, which does need the fixed-partition-order bookkeeping k-means' centroid-update step already has) thread count can never perturb a GEMM’s per-element result. One line changed: Parallelism::None → Parallelism:: Rayon(0). Via rayon::current_num_threads(), this automatically and correctly respects turbovec.build_parallelism’s bounded pool with zero extra plumbing (train_kmeans already runs inside build_pool::install(pool, ..)).

Measured on real hardware, real scale (16-core AVX-512 a cloud VM instance, GIST-1M corpus shape: n_sample=1,048,576, dim=960, lists=4096, full 25-iteration k-means training):

Variant Wall clock vs today
Parallelism::None (v1.22.0, shipped) 2686.6s (~44.8 min) baseline
Parallelism::Rayon(0) (this release) 768.4s (~12.8 min) 3.50×

Centroids confirmed bit-identical between the two runs.

The investigation took two wrong turns before this number, both worth recording rather than hiding: (1) an early microbenchmark of gemm at specific shapes segfaulted with a real gdb backtrace into gemm-common internals, looking exactly like a crate memory-safety bug — root cause was the test harness’s own transposed-read stride bug (a genuine ~16M-element out-of-bounds read), not a gemm bug; replicating the real call site’s exact strides showed no issue. (2) A later a cloud VM timing harness’s naive “serial baseline” — an explicit 1-thread rayon::ThreadPoolBuilder wrapped around the whole training call, meant to isolate the GEMM’s own parallelism — also accidentally forced the unrelated, already-parallel k-means++ seeding phase down to 1 thread, making the “today” baseline look far slower than v1.22.0 actually behaves in production (where seeding always runs on the real build_parallelism pool regardless of the GEMM’s parallelism setting). Both were caught and retracted before being reported as findings; the final harness varies only pool size as an independent axis and compares GEMM modes strictly within each pool size, matching what the real code actually does.

New regression test kmeans_deterministic_across_pool_sizes (sized to exceed gemm’s DEFAULT_THREADING_THRESHOLD so it genuinely exercises multi-threading, not a no-op) asserts byte-identical CoarseModel.centroids across pool sizes {1,2,3,4,8}. 263/263 tests (1 ignored), drift-check clean, compile-matrix clean all 6 PG versions.

Migration: ALTER EXTENSION pg_turbovec UPDATE TO '1.22.1'; is sufficient. No REINDEX — this changes build wall clock only, not the on-disk bytes (centroids/assignment/everything downstream of train_kmeans is byte-identical to v1.22.0 for the same input).

[1.22.0] — 2026-07-04

Repo cleanup, no functional change. Prompted by an audit for “silent GUC” traps, unfinished/debugging artifacts, and general repo hygiene before a release. No wire-format change (MetaPageData::version stays 5), no REINDEX.

Removed

  • turbovec.mmap_static_blocked — a deprecated no-op GUC since v1.19.0 (it toggled a relfile-mmap fast path that v1.19.0 deleted). Removed after a three-minor deprecation window (v1.19.0 warn → v1.20.0/v1.21.0 still-warning → v1.22.0 remove), per AGENTS.md’s SQL-surface-removal policy (two-release minimum). SET turbovec.mmap_static_blocked = ... now errors like any other unknown GUC instead of silently no-op'ing.
  • .woodpecker/ci.yaml — an orphaned CI config from before the project moved to Forgejo Actions (last touched at v0.3.0, never referenced by any current doc or actually run). It was also the only place cargo fmt --all -- --check was ever wired up, which is how 244 formatting violations accumulated across the tree without CI ever catching them (see below).
  • relfile_mmap_static_round_trip_matches_buffer_manager (test) — compared the mmap read path against the buffer-manager fallback; meaningless now that there is only one read path. Also incidentally wrote a stray debug file (/tmp/pg_turbovec_phase_r3_smoke.txt) on every run.

Fixed

  • cargo fmt’d the whole tree (244 pre-existing violations, purely mechanical/cosmetic — no behavior change). fmt-check is now wired into .github/workflows/test.yml and .githooks/pre-push so this can’t silently reaccumulate.
  • Literal \uXXXX escape-sequence artifacts (e.g. \u2014 instead of an actual em dash) in, , src/extras.rs, src/index/cost.rs, and the now-removed .woodpecker/ci.yaml — cosmetic (doc comments, not code behavior), but a real artifact of a write-tool double-escaping bug worth stamping out repo-wide rather than file-by-file.
  • A stale dead-code compiler warning: highdim_oversample_recovers_ recall’s unused lists: i64 = 141 local (the CREATE INDEX right below it hardcoded the literal 141 instead of interpolating the variable).
  • src/guc.rs’s own module-doc GUC table was missing turbovec.search_k and turbovec.probes — two of the most-used GUCs in the extension — and had the wrong range for turbovec.cache_size_mb (documented as 1..=65536; the actual registered range is 0..=65536, and 0 has real, documented meaning: it disables caching). Fixed in src/guc.rs and propagated to docs/ARCHITECTURE.md §9’s GUC table, which had drifted to only 6 of the 17 real GUCs.
  • README.md’s “Operations note: shared_buffers” section still described the v1.5.0–v1.18.x mmap-era guidance (“1.5× the index size is no longer required”, “shared_buffers size no longer bounds warm-scan latency”) as current fact. It’s the opposite of current reality since v1.19.0 removed mmap: shared_buffers sizing matters again, and pg_turbovec’s 7–15× compression is what makes fitting the hot index in shared_buffers achievable. Rewritten to describe the actual current (buffer-cache-only) read path.
  • docs/BUFFER_CACHE_ONLY_DESIGN.md was never actually committed to git despite describing a change that shipped in v1.19.0 (it sat untracked in the working tree for multiple sessions). Committed now with its status header corrected from “DESIGN” (proposal) to “IMPLEMENTED” (it already is, and has been since v1.19.0).

Documentation-only, for context

  • Added an explicit warning to turbovec.bit_width_default’s GUC description: the name is turbovec.bit_width_default, not turbovec.bit_width — PostgreSQL silently accepts SET turbovec.<anything> as a no-op placeholder custom GUC when the name doesn’t match one this extension actually registered (a generic PostgreSQL behavior, not a pg_turbovec bug), so a typo’d SET turbovec.bit_width = N neither errors nor does anything. A benchmark driver script hit exactly this during the v1.21.0 Phase G-1 validation (see that release’s CHANGELOG entry) — every “bw=2” row in the original Phase G-0 results was silently built at bw=4. Use the bit_width index reloption (WITH (bit_width = N)) to set it at CREATE INDEX time.

Migration

No REINDEX. Wire stays v5; no SQL surface change besides the removed deprecated GUC. ALTER EXTENSION pg_turbovec UPDATE TO '1.22.0'; is sufficient.

[1.21.0] — 2026-07-03

Phase G-1: centroid graph for sublinear IVF coarse-cell selection. (Gated in by the finding that the IVF-vs-HNSW latency gap at SIFT-1M didn’t clear the bar for a full corpus graph, so this release attacks the coarse-probe cost instead.) In-memory / scan-path only; no wire change (MetaPageData::version = 5), one new GUC, no REINDEX.

Added

  • Centroid graph coarse-cell selection. For an out-of-core IVF index (turbovec.out_of_core cell-scoped path) with lists >= 4096, coarse_probe can now navigate a small fixed-out-degree (16) undirected graph over the coarse centroids — a Vamana/ HNSW-lite greedy beam search — instead of scoring every centroid. The graph is built once per backend, in-memory, from the already-persisted coarse centroids (ivf::build_centroid_graph); nothing new is persisted, so existing IVF indexes get it for free on the next scan, no REINDEX.
    • Undirected by construction. A pure directed k-NN graph (each centroid’s own nearest-16 others) can strand a cell that’s someone else’s close neighbour but has no close neighbours of its own pointing back — a real navigability gap for greedy search that a randomized-corpus test caught during development. build_centroid_graph symmetrizes every edge (adds the reverse of each directed edge) before search ever runs, which is what makes the recall-preservation guarantee below actually hold.
    • Byte-deterministic. The directed pass is per-row independent (parallel-safe); the symmetrization pass is a fixed sort+dedup over the full edge list. Same centroids ⇒ byte-identical graph. Verified by centroid_graph_build_deterministic (unit) and ivf_coarse_graph_build_is_deterministic_across_cache_rebuilds (#[pg_test], end-to-end through the relfile + cache).
    • Recall-preserving. graph_probe’s beam width (ef = max(nprobe*4, 32), the classic HNSW-style slack budget) is sized so the graph-navigated result SET matches the exact linear scan’s nprobe-nearest cells exactly at the tested scales (verified by graph_probe_matches_linear_scan_exactly, 150 random queries across 5 nprobe values, and the end-to-end ivf_coarse_graph_matches_linear_scan #[pg_test]). Existing recall-floor tests (index_am_recall_floor_{2,3,4}bit) still pass unmodified.
  • turbovec.coarse_graph (GUC, enum, default auto): auto builds/uses the graph only when lists >= 4096 (below that the plain linear scan is already cheap and a graph’s build + per-query overhead isn’t worth paying — see ivf::GRAPH_MIN_LISTS’s doc); on forces it regardless of lists; off always uses the exact linear scan. ivf_coarse_graph_auto_falls_back_below_threshold proves the small-lists fallback is correctness-neutral (matches off exactly) and that forcing on below the threshold still matches too.

Honest notes

  • G-1 is scoped to the out-of-core (OocIvfIndex) scan path only. The whole-load path (ivf_setup_and_search in src/index/scan.rs) re-reads centroids fresh from the relfile on every scan-open (there is no per-backend cache of that struct today), so building an O(lists²) graph there per scan would be pure overhead, not the “build once per backend” amortised cost the plan requires. That path is also gated to comfortably-RAM-resident indexes, i.e. the small-lists regime where the linear scan is already cheap — the OOC path is also where lists >= 4096 (the scale G-1 targets) actually shows up in practice.
  • Correcting a docs-drift bug found while implementing this release: the v1.20.0 CHANGELOG entry below claims a “sublinear two-level coarse quantizer” (O(lists)→O(√lists)) shipped in that release. That was never implemented. v1.20.0’s actual diff (verified against git show) only parallelized k-means++ seeding and the build-time assign-sweep, and added turbovec.scan_parallelism for the fine-scan. coarse_probe remained the plain O(lists·dim) linear scan through v1.20.1. There is no TwoLevelCoarse type or equivalent anywhere in the v1.20.0–v1.20.1 source tree. v1.21.0 (this release) is the first to actually ship sublinear coarse-cell selection. The v1.20.0 CHANGELOG entry and docs/UPGRADING.md’s corresponding row are left as historical record (not rewritten) but this release’s docs/UPGRADING.md row calls out the correction explicitly.
  • Measured effect is a rough local sanity check (small-corpus pgrx test host, not a benchmark AVX-512/AVX2 latency benchmark): correctness and recall-preservation are verified by the tests above; a proper before/after p50 comparison at lists in the low-to-high thousands (where auto actually engages) on an AVX2+ host is follow-up bench work, not part of this patch.

Migration

No REINDEX. Wire stays v5; existing v4/v5 IVF indexes benefit from the centroid graph (when lists >= 4096) with no rebuild. ALTER EXTENSION pg_turbovec UPDATE TO '1.21.0'; is sufficient.

[1.20.1] — 2026-07-03

CRITICAL PERF FIX: turbovec.iterative_scan default flipped relaxed_order → off. Wire format unchanged (MetaPageData::version = 5); no SQL surface change; no REINDEX needed — ALTER EXTENSION pg_turbovec UPDATE is sufficient and the new default takes effect on the next backend.

The bug

PostgreSQL’s reorder queue (IndexNextWithReorder in nodeIndexscan.c) can only return a candidate tuple early when the index AM’s advertised ORDER BY value for that tuple is exact. pg_turbovec always advertises f64::NEG_INFINITY for every tuple (deliberately opclass-agnostic — correct across L2/cosine/inner- product without per-opclass bounds logic), so that exactness condition can never be satisfied. Under the old default (relaxed_order), this forced the executor to drive the AM’s own iterative-refill schedule (probe-widening, search_k doubling, up to turbovec.max_scan_tuples = 20,000 and turbovec.max_probes = 64) all the way to completion on every ORDER BY dist LIMIT n query — no matter how small n was — before the executor’s reorder queue could signal it was safe to return even the first row.

Measured on an AVX-512 a cloud VM host (SIFT-1M/128d IVF, probes=8, otherwise-default GUCs): ~2 ms with turbovec.iterative_scan = off vs ~900 ms with the old default relaxed_order — a 450x latency tax paid by every default-configuration KNN query since relaxed_order first shipped as the default in v1.8.0. Every benchmark and load-test in this repository explicitly set turbovec.iterative_scan = off, which is why this went undetected for eleven releases: the bug only manifests when a caller does NOT override the GUC, and no internal benchmark left it at its default. It was caught while measuring the Phase G-0 IVF-vs-HNSW frontier, which (deliberately) exercised the untouched defaults for the first time.

Changed

  • turbovec.iterative_scan now defaults to off (was relaxed_order). off matches pgvector’s own hnsw.iterative_scan default and only under-returns on a selective WHERE filter combined with ORDER BY ... LIMIT — a much rarer shape than the plain unfiltered KNN query this bug taxed. Opt back into relaxed_order (SET turbovec.iterative_scan = relaxed_order;) if your workload relies on the under-return-avoidance guarantee for selective filters; see docs/FILTERING.md and docs/PRODUCTION.md.
  • Added index_am_iterative_scan_defaults_to_off regression test (src/lib.rs) that asserts the compiled-in default — without an explicit SET — caps at search_k rather than draining to max_scan_tuples, so this can’t silently regress back.
  • Fixed one pre-existing test (ivf_lists_scan_matches_flat) that was implicitly relying on the old relaxed_order default to guarantee finding an exact self-match under quantization noise; it now opts into relaxed_order explicitly, since that’s a correctness anchor for k-widening behaviour, not a test of the compiled-in default.
  • Corrected the documented default in src/guc.rs’s module-level GUC table, docs/PRODUCTION.md, docs/FILTERING.md, docs/MIGRATING_FROM_PGVECTOR.md, and docs/PARITY_GAPS.md.

Migration

ALTER EXTENSION pg_turbovec UPDATE TO '1.20.1'; (empty migration file, migrations/032_pg_turbovec_v1.20.1.sql). No REINDEX. The new GUC default applies to new backends/sessions; a long-lived backend that already read the old compiled-in default at connection start keeps using it until it reconnects (ordinary PostgreSQL GUC semantics — not specific to this fix).

[1.20.0] — 2026-07-02

IVF scaling — parallel build + parallel scan + sublinear coarse quantizer. Surfaced by the benchmark A/B/C benchmark (the benchmark host, AVX-512). Scan-path / build-path / in-memory only; no wire change (MetaPageData::version = 5), one new GUC, no REINDEX — existing indexes benefit with no rebuild.

Added

  • Sublinear two-level coarse quantizer (the key scaling enabler). For an IVF index with lists > 4096, cell selection is now O(√lists) instead of O(lists). The two-level structure is computed in-memory at index-open from the already-persisted coarse centroids (a deterministic function of them) — nothing new is persisted, so existing IVF indexes get it for free on the next scan, no REINDEX. Measured 39–364× fewer centroid distance computations per query at recall 1.0, which breaks the coarse-probe wall: lists can grow large (tiny cells) without the coarse step becoming the bottleneck — the enabler for IVF at 10M+.
  • turbovec.scan_parallelism (GUC, int, default 0 = auto = min(cores, 4); 1 = serial). Parallelizes the per-query IVF fine-scan across probed cells (out-of-core path), cutting single-query latency at high dimension. Conservative default to protect aggregate QPS under concurrency. Results identical to the serial scan (same top-k; verified).

Changed

  • Parallel IVF build. k-means training + the assign-sweep now use the bounded build pool (turbovec.build_parallelism) across cores, memory-bounded and byte-identical (the reduction order is a fixed function of the input, independent of thread count — so the relfile is reproducible across machines with different core counts).

Honest notes

  • The parallel build measured only ~1.27× on 64-core GIST-960d: the single-threaded GEMM (Parallelism::None, for bit-exactness) remains the dominant term, and row-blocking around it can’t parallelize the GEMM itself. This release makes the parallel build safe and correct (no OOM, deterministic) but it is not the full build-cliff fix — a deterministic parallel GEMM (or an IVF bit-exactness policy change enabling BLAS threads) is scoped follow-up work.
  • The high-dim recall ceiling was investigated and found retrieval-bound (addressed by the sublinear coarse + more probes), not quantization-bound (widening the reorder-rescore pool recovered ~0 recall).
  • Query-latency-at-scale and 10M-build validation on a cloud VM are pending (the 10M run OOM’d before the memory fix in this release; re-run needed to confirm).

Migration

No REINDEX. Wire stays v5; existing v4/v5 indexes benefit from the sublinear coarse + parallel scan with no rebuild. ALTER EXTENSION pg_turbovec UPDATE TO '1.20.0'; is sufficient. Tests: 241 → 249.

[1.19.0] — 2026-06-18

All index reads through the PostgreSQL buffer manager. Read-path architecture change; no wire change (MetaPageData::version unchanged), no SQL-surface change, no REINDEX. Required for managed/sandboxed Postgres and any environment that restricts direct file access.

Changed

  • Removed the direct relfile mmap. Every byte of index data is now read through PostgreSQL’s shared-buffer cache (ReadBufferExtended) — there is no mmap/pread of the relfile. The buffer manager is the single source of truth for page access (consistent pinning/locking; clean crash + streaming-replication semantics). src/index/mmap_static.rs is deleted and the memmap2 dependency dropped (net −700 lines). The buffer-manager readers this routes through already existed (they were the mmap fallback), so this is mostly a deletion. See docs/BUFFER_CACHE_ONLY_DESIGN.md.
  • Out-of-core (>RAM) IVF serving is preserved without mmap. The cell-scoped gather (OocIvfIndex::search_ooc → relfile::gather_codes_ranges) reads only the probed cells' pages through the buffer manager, so the per-backend resident set stays O(probes * cell_size), not O(n). Out-of-core serving needs cell-contiguous layout + range-scoped reads — not mmap.

Deprecated

  • turbovec.mmap_static_blocked is now a no-op (ignored). It is retained for one minor release so an existing SET does not error, and will be removed in a future minor.

Performance notes

  • Warm queries: unchanged (the prepared index is cached per-backend; warm scans never touch the buffer manager).
  • Cold cache-fill on a >shared_buffers index: slower than the old mmap path (per-page lookup/pin/lock + a buffer-manager copy). Mitigation: size shared_buffers to hold the hot index — pg_turbovec’s 7–15× compression is what makes “the index fits shared_buffers” practical where fp32 HNSW could not. For a >RAM index, use IVF + turbovec.out_of_core so only probed cells' pages are read.

Migration

No REINDEX. Read-path only; wire format unchanged. ALTER EXTENSION pg_turbovec UPDATE TO '1.19.0'; is sufficient. Tests: 241 (unchanged) on pg16; all 6 PG versions compile; drift-check clean.

[1.18.0] — 2026-06-18

Tier-1 IVF latency optimizations (scan-path). No SQL-surface change, no wire change (MetaPageData::version = 5; single-vector stays v4), no REINDEX. Closes the Tier-1 backlog — narrowing the IVF-vs-HNSW latency gap by attacking the actual per-query floor, with evidence rather than speculation.

Changed

  • Default turbovec.search_k lowered 100 → 32 (#1a). The dominant per-query cost is the executor’s reorder-recheck of every returned candidate (a heap-tuple fetch + an exact full-precision distance recompute each) — not the vector scan. The new searchk_recall_frontier test shows recall@10 plateaus by search_k≈25 (25/50/100/200 identical), so the old default of 100 over-provisioned the recheck ~3× for zero recall gain. The real-corpus recall_floor_{2,3,4}bit tests pass at the new default, confirming recall safety. Raise it for LIMIT > ~20 or a hard corpus; lower it (toward 16) for the lowest latency.

Added

  • assign_dups_probes_pareto test + guidance (#2). Demonstrates that raising WITH (assign_dups = M) (soft multi-assignment) lets a query reach a matched recall while probing fewer cells (best recall@10 climbs 0.173 → 0.207 → 0.240 as assign_dups 1 → 2 → 4 on the test corpus; min-probes-to-matched-recall is non-increasing in assign_dups). Opt-in (a build-time layout choice; `assign_dups

    1` needs a REINDEX). The default (1) is unchanged.

  • Frontier artifacts: benches/results/searchk_recall_frontier_2026-06-18.json, benches/results/assign_dups_probes_pareto_2026-06-18.json.

Investigated and rejected / deferred (documented, not built)

  • #1b (advertise a tighter ORDER BY distance) — rejected as a no-op. PostgreSQL’s IndexNextWithReorder rechecks (heap fetch + exact recompute) every candidate unconditionally under xs_recheckorderby, before reading the advertised value; a tighter bound reduces zero work (identical PG 13–18). Documented in src/index/scan.rs.
  • #3 (SIMD coarse_probe) — assessed, deferred: it is a fixed-floor term, not the dominant cost, and a SIMD horizontal-sum risks cross-ISA reduction-order divergence → recall drift.
  • #4–6 — not warranted by the data (zero effect on the OOC-gathered benchmark; build-when-profiled).

Honest caveat

The latency confirmation of #1a/#2 (does p50 actually drop ~3×?)
is deferred to a quiet AVX2 host — both floki and arnold were
saturated by unrelated work during this release. The recall safety
of every change is host-independent and verified here; only the
latency number awaits a quiet window. The projected effect
~10–13 ms at recall@10≈0.96 at 500k–1M, matching HNSW ef40–ef100.

Migration

No REINDEX. Scan-path / default-tuning only; wire stays v5. ALTER EXTENSION pg_turbovec UPDATE TO '1.18.0'; is sufficient (the new search_k default applies to new sessions). Tests: 239 → 241.

[1.17.1] — 2026-06-18

ColBERT recall win confirmed cross-domain. Docs + bench-results release; no source, SQL-surface, or wire change (MetaPageData::version = 5, single-vector still v4); no REINDEX.

Confirmed

The Phase F-2 index-native ColBERT recall gain (shipped v1.17.0) was replicated on a second, out-of-domain corpus (BEIR/NFCorpus, 3,633 docs, medical/nutrition, entity-heavier), exercising the persistent vec_colbert_ops index on floki (AVX2):

  • +0.044 nDCG@10 / +0.037 Recall@10 vs the Phase-D pooled+rerank baseline at the value operating point (candidate_n=256), rising to +0.065 nDCG at low candidate budget (candidate_n=128, where the pooled baseline collapses to 0.220 while colbert_search holds at 0.285) — same sign, same mechanism, same low-budget shape as the SciFact gate at every config.
  • Quantization signal intact (2-bit ≈ 4-bit, ≤0.0001 nDCG; 2-bit index = 43 MB).
  • The persistent index built cleanly (561k token slots, 42 s / 43 MB at 2-bit, no OOM) and served from disk — the F-1 ~28 MB/call backend-RSS leak is gone (RSS plateaus flat at ~360 MB; ~1.4 KB/call warm).

The qualified GO is upgraded to an established cross-domain recall win. Data: benches/results/colbert_f2_confirm_floki_nfcorpus_20260618.json; harness: benches/scripts/colbert/.

Docs

and updated to record index-native late interaction as DONE (was a future phase): pg_turbovec is one of two PostgreSQL extensions (with VectorChord) with index-native multivector/MaxSim, and the only one also 7–15× smaller than HNSW.

Migration

No REINDEX. Docs + bench only; wire stays v5 (single-vector v4). ALTER EXTENSION pg_turbovec UPDATE TO '1.17.1'; is sufficient. Tests unchanged (239).

[1.17.0] — 2026-06-18

Phase F-2 — persistent index-native ColBERT late interaction. New index kind; additive wire bump 4 → 5 (single-vector indexes stay byte-identical to v4); no REINDEX for any existing index. This makes pg_turbovec one of only two PostgreSQL extensions (with VectorChord) to offer index-native multivector/MaxSim — and the only one that is also 7–15× smaller than HNSW.

Added

  • Persistent ColBERT token index. CREATE INDEX ON docs USING turbovec (tokens vec_colbert_ops) over a turbovec.vector[] column (per-doc token arrays) builds a v5 on-disk token index: ambuild unnests each doc’s vector[] into per-token slots (the doc’s heap TID repeated per token — the IVF soft-assign synthetic-slot-id machinery), laid out IVF cell-contiguous and spilled via the Phase B-4 BufFile (n_tokens ≫ n_docs, so the spill is load-bearing). Determinism: tokens are unnested in array order.
  • turbovec.colbert_search now reads the persistent index. It locates the vec_colbert_ops index on the token column and runs stage-1 candidate generation against the on-disk relfile (warm cache or cold read) instead of rebuilding a backend cache every call. Stage-2 still exact-MaxSim-reranks heap tokens by ctid. The F-1 ~28 MB/call backend-RSS leak is eliminated on the persistent path (and bounded on the no-index fallback).
  • vec_colbert_ops operator class over turbovec.vector[] — support function max_sim, no order-by operator, so the planner can never select a ColBERT index for ORDER BY (the forbidden amrescan scan-key path is untouched). A ColBERT index ERRORs on an ORDER BY scan with a HINT to use turbovec.colbert_search.
  • VACUUM reuses the IVF tombstone path unchanged: a deleted doc’s TID marks all its token slots dead (the many-slots-one-TID shape is identical to IVF soft-assign dups); cells stay contiguous (tombstone, never swap-remove); colbert_search masks tombstoned slots.

Wire format (additive v5, per index kind)

A new kind byte at page offset 30 (formerly a reserved zero) discriminates KIND_SINGLE (0, single-vector, wire v4) from KIND_COLBERT (1, multivector, wire v5). A single-vector build never sets it, so it emits wire version 4 + kind 0 — byte-identical to v1.16.0 (guarded by single_vector_still_emits_v4_bytes and v4_single_vector_index_byte_identical). A v4 meta decodes as kind = KIND_SINGLE, so is_legacy_v4() never trips. EXPECTED_WIRE_FORMAT_VERSION is now 5.

Migration

No REINDEX. Existing single-vector indexes are byte-identical and read unchanged under the v5 binary. A ColBERT index is a brand-new shape, built fresh (no in-place conversion from a single-vector index). ALTER EXTENSION pg_turbovec UPDATE TO '1.17.0'; registers the new opclass and is sufficient. See docs/UPGRADING.md.

Tests

230 → 239 (+colbert_persistent_build_and_search, colbert_persistent_recovers_single_token_match, colbert_persistent_survives_vacuum, colbert_persistent_deterministic, colbert_index_rejects_orderby_scan, v4_single_vector_index_byte_identical, + 3 page.rs unit tests). All six PG versions (13–18) compile; drift-check clean (after the VERSION-5 / minor-bump pairing this release provides).

[1.16.0] — 2026-06-17

Phase F-1 — index-native late interaction (ColBERT stage-1). Additive SQL function in the turbovec schema; no wire change (MetaPageData::version = 4), no index-AM change, no REINDEX. Closes the last acknowledged feature gap vs Qdrant/VectorChord at the level the analysis showed actually matters (stage-1 recall) —.

Added

  • turbovec.colbert_search(rel, id_col, token_col, query vector[], k, per_token_k = 64, candidate_n = 256, bit_width = 4) (src/colbert.rs) — the index-accelerated stage-1 of ColBERT late interaction. Stage 1 builds a backend-cached flat token index (one slot per token across all docs, doc-id repeated; synthetic unique slot-ids fed to IdMapIndex, real doc-ids kept separately — the IVF soft-assign trick), batch-searches all |Q| query tokens, and unions the hit doc-ids into a candidate set. Stage 2 reads each candidate’s full token array from the heap and scores it with the exact max_sim kernel (Phase D). Returns the top-k documents.
  • The value over the Phase D pooled-vector + max_sim re-rank pattern is stage-1 recall: a document is retrieved by its best single token, not its pooled mean — so a doc whose pooled vector is far but which has one token near a query token (the entity/rare-term/long-doc case ColBERT is built for) is still found. Proven by the colbert_search_recovers_single_token_match test.

What it is / isn’t

The token index lives only in the backend cache (the turbovec.knn model) — there is no relfile, no CREATE INDEX, and no wire-format change. It is the index-native stage 1 over max_sim’s exact stage 2; it is not the full persistent multivector index AM (per-token relfile + MaxSim-aware scan + PLAID pruning). That persistent AM (Phase F-2) is gated on a measured recall/latency win over this F-1 path on a real ColBERT corpus — the plan explicitly refuses to build a 32–512×-larger persistent index on faith. Tuning: per_token_k / candidate_n trade recall for work; raise per_token_k under heavy (2–3 bit) token quantization.

Migration

Additive function; no wire change, no REINDEX. ALTER EXTENSION pg_turbovec UPDATE TO '1.16.0'; is sufficient. Tests: 224 → 230 (+colbert_search_basic, colbert_search_recovers_single_token_match, colbert_search_matches_bruteforce_maxsim, colbert_search_empty_query, colbert_search_deterministic, colbert_search_rejects_bad_k).

[1.15.1] — 2026-06-17

Cross-version build fix (pg13 / pg14 / pg15 / pg18). Build-only patch; no wire change (MetaPageData::version = 4), no behaviour change on the versions that already compiled (pg16 / pg17), no REINDEX.

Fixed

  • The Phase B-4 out-of-core IVF build (v1.12.0) called pg_sys::BufFileReadExact (PG16+ only) and passed BufFileWrite a *const pointer (PG13–15 declare it *mut). The extension compiled on pg16/pg17 — the local dev target — but failed to compile on pg13, pg14, pg15, and pg18, silently breaking those CI matrix legs from v1.12.0 through v1.15.0. Now both calls go through version-gated shims (buffile_write / buffile_read_exact in src/index/build.rs), backing the read with the universally-present BufFileRead + an explicit short-read check. All six PG features (13–18) compile again.

Added (CI hardening)

  • scripts/compile-matrix.sh — cargo checks every pgNN feature in Cargo.toml (compile-only, ~20s each, no test cluster), so version-specific C-API breaks are caught locally before tagging. Wired into .githooks/pre-push alongside drift-check.sh. Skips via COMPILE_MATRIX_SKIP=1 on hosts without every pgrx toolchain. This is the gate that would have caught the v1.12.0 regression; cargo pgrx test pg16 alone never could.

Migration

No REINDEX. Build-only; wire stays v4; pg16/pg17 runtime unchanged. ALTER EXTENSION pg_turbovec UPDATE TO '1.15.1'; is sufficient. Tests: 224 (unchanged) on pg16; the fix is verified by all six PG features compiling.

[1.15.0] — 2026-06-17

Phase C follow-up — operator-path allowlist on flat + IVF. Additive GUC + function in the turbovec schema; no wire change (MetaPageData::version = 4), no index-AM scan-key rewrite, no REINDEX. Brings the in-kernel allowlist pushdown (previously turbovec.knn()-only, flat-only) to the ORDER BY emb <=> q LIMIT k operator path.

Added

  • turbovec.allowlist (session string GUC, default "") — a CSV of heap TIDs (encoded as bigint). When set, the index-AM scan ANDs the allowed slots into the slot mask it hands the SIMD kernel, so the kernel short-circuits 32-vector blocks with no allowed slot before any LUT work — the same in-kernel block-skip knn(..., allowed) gets, now on the operator path. On an IVF index the allowlist is ANDed with the probed-cell mask, scoping the skip to probed cells ∧ allowed slots; the out-of-core cell-scoped path gets it too. Empty/unset = exact prior behaviour with zero added hot-path cost (no slot-bool is ever built). Parsed once per scan (refills reuse it); a non-integer token ERRORs the scan.
  • turbovec.tid_to_bigint(tid) -> bigint — the ergonomic encoder for building the allowlist from ctid (returns the (block << 32) | offset value the AM stores per slot), so users never hand-write the bit-twiddling. Verified bit-identical to the raw encoding (tid_to_bigint_matches_raw_encoding).

Notes / honest limitation

The allowlist is a set of heap TIDs, not an id column — the index AM keys vectors by heap TID, never a heap id column; turbovec.knn(..., allowed) remains the id-column path. This is a pre-materialized id-set channel, not arbitrary-WHERE pushdown (which would require scan-key reinterpretation — the forbidden amrescan rewrite — or payload columns in the index). See docs/FILTERING.md §§ 3.5, 6, 7. Composes with tombstones (a vacuum-deleted row is excluded even if allowlisted) and with probes >= lists (exact over the allowed set); returns the same rows as knn() for the same id-set.

Migration

Additive GUC + function; no wire change, no REINDEX. ALTER EXTENSION pg_turbovec UPDATE TO '1.15.0'; is sufficient. Tests: 215 → 224 (+allowlist_guc_restricts_ordered_scan_flat/_ivf, allowlist_guc_matches_knn, allowlist_guc_empty_is_unfiltered, allowlist_guc_composes_with_tombstones, allowlist_guc_probes_all_exact, allowlist_guc_rejects_bad_token, allowlist_guc_out_of_core, tid_to_bigint_matches_raw_encoding).

[1.14.0] — 2026-06-17

Phase D — breadth parity (multivector + hybrid fusion). Additive SQL surface in the turbovec schema; no wire-format change (MetaPageData::version = 4), no index-AM change, no REINDEX. Closes the multivector / hybrid-fusion breadth gap vs VectorChord / Qdrant at the SQL layer.

Added

  • turbovec.max_sim(vector[], vector[]) / max_sim_cosine(...) (src/hybrid.rs) — ColBERT-style late-interaction MaxSim: sum_{q in Q} max_{d in D} sim(q, d) over per-token vector[] arrays. max_sim uses dot-product similarity (correct for L2-normalised tokens); max_sim_cosine uses cosine similarity (1 - cosine_distance). All token vectors across both arrays must share one dimension (ERROR on mismatch); an empty query or empty doc scores 0.0 (ColBERT convention). This is a re-rank primitive (ANN-retrieve candidates on a pooled vector, MaxSim-rerank the top-N) — the token arrays are not indexed, and index-native late interaction remains a documented future phase.
  • turbovec.rrf_score(rank integer, k integer DEFAULT 60) (src/hybrid.rs) — reciprocal rank fusion term 1.0 / (k + rank) for fusing a dense ANN ranking with a sparse / keyword ranking. Pairs with the documented CTE recipe; non-positive denominator raises ERROR.
  • docs/HYBRID_SEARCH.md — the canonical breadth guide: multivector MaxSim re-rank (signature, conventions, the two-stage retrieve-then-rerank pattern, the honest index-native limitation), the dense+sparse RRF recipe (full ROW_NUMBER() + rrf_score CTE for both full-text and sparsevec), and the named-vector multi-column schema pattern.
  • Cross-links from README.md, docs/PRODUCTION.md, docs/PARITY_GAPS.md, and docs/MIGRATING_FROM_PGVECTOR.md; the multivector / hybrid rows now read “SQL surface SHIPPED; index-native late interaction is a future phase.”

Notes

  • Out-of-core BUILD (roadmap Phase D-3) already shipped in v1.12.0 (streaming IVF build); no re-implementation.
  • Named vectors (multiple vector columns per row) are a documented schema pattern, not new code.

Migration

Additive SQL functions only; no wire change, no REINDEX. ALTER EXTENSION pg_turbovec UPDATE TO '1.14.0'; is sufficient (the new turbovec.max_sim / max_sim_cosine / rrf_score functions are created by the update script). Tests: 203 → 215 (+max_sim_basic, max_sim_dim_mismatch_errors, max_sim_empty, max_sim_cosine_normalised, max_sim_rerank, rrf_score_values, hybrid_rrf_recipe, plus 5 in-module unit tests).

[1.13.1] — 2026-06-17

Phase C — metadata-filtering docs + measured allowlist crossover. Docs + benchmark release; no source-logic, SQL-surface, or wire change (MetaPageData::version = 4); no REINDEX. Bench-results and documentation only.

Added

  • docs/FILTERING.md — the canonical guide to pg_turbovec’s three working metadata-filter mechanisms, with a cardinality×selectivity×corpus decision matrix:
    1. Partial index (CREATE INDEX ... WHERE tenant_id = X) — native PG predicate pushdown; the default for known, low-cardinality filters.
    2. In-kernel allowlist turbovec.knn(rel, id_col, vec_col, query, k, bit_width, allowed bigint[]) — true in-kernel pushdown (the SIMD kernel skips 32-vector blocks with no allowed slots before any LUT work), flat-only, for selective per-query id sets.
    3. Iterative scan + WHERE (v1.8.0) — the ORDER BY emb <=> q LIMIT k AM path; the executor rechecks the predicate, the AM widens k/probes (capped by max_scan_tuples). Includes the honest limitation: no true in-traversal pushdown on the ORDER BY AM path (the index stores only vector codes + TID, no payload columns), and a C-4 design sketch for a future phase.
  • Measured allowlist selectivity crossover (floki, AVX2, 300k× 256-d, 4-bit, k=10): allowlist latency decreases monotonically as the filter tightens (17.9 ms → 0.48 ms, ~37×) while the naive post-filter is flat (~7 ms); crossover at ~7–10% selectivity, up to 14.7× faster at 0.1%. benches/allowlist_crossover.rs + benches/results/allowlist_crossover_floki_v1_13_0_20260617.json.

Fixed (docs drift)

-: refreshed v1.10.1/v1.11.0 → v1.13.0; the >500k IVF build ceiling and the >RAM gaps are now marked CLOSED (out-of-core build v1.12.0 + out-of-core query v1.13.0); the metadata-filtering row reflects the three real patterns instead of “post-filter only”. - docs/PARITY_GAPS.md: added the metadata-filtering row; corrected the stale “Parallel index build | GAP — single-threaded” row (parallel build shipped v1.8.0, turbovec.build_parallelism). - docs/MIGRATING_FROM_PGVECTOR.md: filtered-ANN section lists all three patterns and links FILTERING.md; knn() signature matches src/knn.rs. - README.md + docs/PRODUCTION.md: cross-link FILTERING.md. - Fixed three pre-existing broken benchmark sources (concurrent_knn, recall_vs_pgvector, recall) that referenced the pre-d3d468e IdMapIndex::new signature (now returns Result); cargo check --benches is green again. (cargo pgrx test never compiled benches, so they did not gate tests.)

Migration

No REINDEX. Docs + bench only; wire stays v4. ALTER EXTENSION pg_turbovec UPDATE TO '1.13.1'; is sufficient. Tests unchanged (203).

[1.13.0] — 2026-06-17

Out-of-core IVF query (>RAM serving) — an IVF index larger than RAM can now be queried, not just built (v1.12.0). Wire format unchanged (MetaPageData::version = 4); no REINDEX. Completes the out-of-core arc for the >5M production deployment .

Added — cell-scoped IVF serving (Phase B-1/B-2)

The scan previously loaded the whole index into a per-backend cache (read_full + a copy of the blocked-codes chain off the mmap), so the resident set was O(n) and an index that exceeded RAM could not be served.

  • Cell-scoped scan. The backend now caches only bounded metadata (coarse centroids, cell directory, rotation, codebook, per-slot scales/ids) plus a MAP_PRIVATE mmap of the relfile, and per query copies only the probed cells' contiguous code ranges off the mmap into a compact throwaway sub-index (cells are contiguous from the build-time permutation). Resident set drops to O(probes * cell_size + faulted pages); hot cells stay in the OS page cache, cold cells fault from disk on demand.
  • turbovec.out_of_core (enum off | auto | on, default auto). auto goes cell-scoped only when the index codes exceed 0.5 * turbovec.cache_size_mb — an in-RAM index loads whole (no per-query gather/reblock cost); only a genuinely large index pays the bounded-memory-for-CPU tradeoff. on forces cell-scoped; off forces the pre-v1.13.0 whole-load.
  • No wire change, no turbovec fork change (reuses from_parts_with_prepared_borrowed). Added mmap_static::gather_slot_ranges + a buffer-manager twin for the fresh-index fallback.

Measured

200k×256-d×4-bit IVF (52 MB on disk): per-backend VmHWM whole-load 140.8 MB → cell-scoped 44.1 MB (~3.2× lower). Under a tight cgroup MemoryMax, the whole-load backend was OOM-killed (postmaster recovered cleanly, no corruption) where cell-scoped stayed within bound. Warm p50 82 ms (whole-load) → 199 ms (cell-scoped) — the expected per-query reblock cost, paid by auto only when the index is too large to keep whole.

Compatibility

Scan-path only. Results identical to the whole-load path (probes >= lists still reduces to the exact flat scan; tombstones masked; soft-assign deduped). MVCC backstops (reorder queue + heap visibility) preserved. Flat (lists = 0) / vacuum-degraded indexes keep the whole-index load (no cells to scope; still O(n)-resident — use IVF for >RAM).

Migration

No REINDEX. Scan-path change; wire stays v4. ALTER EXTENSION pg_turbovec UPDATE TO '1.13.0'; is sufficient.

Tests

197 → 203 (+ivf_ooc_results_match_whole_load, ivf_ooc_probes_all_equals_flat, ivf_ooc_tombstones_masked, ivf_ooc_soft_assign_dedup, ivf_ooc_installs_cell_scoped_handle, ivf_ooc_auto_is_size_aware). drift-check clean.

[1.12.0] — 2026-06-17

Out-of-core IVF build — IVF indexes can now be built at 1M–5M+ rows on a RAM-constrained host. Wire format unchanged (MetaPageData::version = 4, byte-identical relfile); no REINDEX. Driven by the >5M production deployment .

Fixed — the 1M+ IVF build OOM (Phase B-4)

The WITH (lists = N) build held the full f32 corpus twice in RAM (ivf_flat ~4 GiB + perm_flat ~4 GiB at 1M×1024-d) plus the growing index and GEMM scratch — a ~14 GiB peak that maintenance_work_mem did not bound, OOM-killing 1M+ builds on a 31 GiB host. IVF was effectively unbuildable at the production scale.

  • Disk spill. The corpus now spills to a PostgreSQL BufFile temp file (in pgsql_tmp, respecting temp_tablespaces / temp_file_limit) during the heap scan, wrapped in a CorpusSpill RAII type. Cleanup is double-covered: the resource owner unlinks on (sub)transaction abort (even when ereport(ERROR) longjmps past Rust destructors) and Drop unlinks on success.
  • Three streamed passes, each bounded by maintenance_work_mem: (1) spill + bounded reservoir sample for k-means; (2) GEMM-assign cells over disk-backed row-blocks, keeping only the per-row cell-id array (not a corpus copy); (3) feed the quantizer in cell order by re-reading the spill at permuted offsets in bounded chunks. The full f32 corpus is never resident; the only RAM term that scales with row count is the quantized packed_codes (7–15× smaller than the f32 corpus).
  • Measured (1M×1024-d, lists=1024, 30 GiB host): peak RSS ~14 GiB (OOM) → ~7.1 GiB (completes); index 1030 MB, spill ~3.9 GB on disk. 5M projected ~8–10 GiB — buildable on the 31 GiB production host.

Determinism / compatibility

Byte-identical relfile to a v1.11.x in-memory build for the same input, and maintenance_work_mem-invariant (TQ+ calibration is fit on a fixed cell-ordered prefix, independent of chunk size). The flat (lists = 0) build path is unchanged (already Phase-W streamed). No GUC added — maintenance_work_mem is the knob.

Migration

No REINDEX. Build-internal; wire stays v4. ALTER EXTENSION pg_turbovec UPDATE TO '1.12.0'; is sufficient. Existing indexes are unaffected; the benefit applies to the next CREATE INDEX / REINDEX.

Tests

193 → 197 (+ivf_streaming_build_determinism_byte_identical, ivf_streaming_build_chunk_size_invariant, ivf_streaming_build_bounded_memory_completes, ivf_streaming_build_temp_file_cleanup). drift-check clean.

[1.11.1] — 2026-06-16

Bench-results-only release. Wire format unchanged from v1.11.0 (MetaPageData::version = 4); no REINDEX. Zero source change.

Benchmark — IVF latency frontier vs HNSW + ivfflat (Phase A-2)

The honest at-scale measurement (isolated AVX2 on arnold, taskset-pinned, contention-gated, warm, 300 queries/config, Cohere-wiki 500k×1024-d) answering “does IVF beat/equal HNSW at scale.” At recall@10 ≈ 0.96:

engine config recall@10 warm p50
pgvector HNSW ef=200 0.966 7.9 ms
pg_turbovec IVF lists=707, probes=64 0.960 18.5 ms
pgvector ivfflat probes=100 0.978 117.4 ms
pg_turbovec flat (exact) all cells 1.000 41.4 ms

Honest verdict: HNSW wins latency at 0.96 (7.9 vs 18.5 ms, ~2.3×). But IVF is now in HNSW’s order of magnitude (not the 490× flat-scan gap), beats pgvector’s own ivfflat 3–6× at every matched recall, beats its own exact flat scan, wins the ≥0.99 recall tail (0.99 @ 25 ms via probes=256; this HNSW config never reaches 0.99), and is 7.5× smaller (518 MB vs HNSW 3902 MB). The earlier ~40 ms projection was pessimistic; real p50 at 0.95 is 18.5 ms.

Critical finding — 1M IVF build OOMs (motivates Phase B-4)

The 1M IVF build OOM-killed the postmaster (~14 GiB peak on a 31 GiB host): the lists > 0 build holds the full flat corpus + a permuted copy + k-means scratch — a structural peak maintenance_work_mem does not bound. Largest IVF index that built on arnold: 500k. 1M/5M IVF are blocked on Phase B-4 (streaming / out-of-core build). The IVF query path is unaffected.

Files: benches/results/ivf_frontier_arnold_cohere-wiki_2026-06-16.json, docs/BENCHMARKS.md, .

[1.11.0] — 2026-06-16

Production hardening for IVF: it now survives VACUUM instead of silently degrading, and builds ~7.8× faster. Wire format stays MetaPageData::version = 4 (additive); no REINDEX. Driven by the >5M production deployment + (Phases A-1, E-2).

Fixed — IVF survives VACUUM (Phase E-2, the production landmine)

An IVF index used to silently degrade to a flat O(n) scan after VACUUM — swap-remove moved the last vector into the deleted slot, breaking cell contiguity, so has_ivf() flipped false and queries fell back to the ~seconds full scan with no operator signal. On a churning multi-million-row index that’s a latency cliff.

  • Tombstones. The IVF ambulkdelete path now leaves dead slots in place and ORs them into a persisted per-slot tombstone bitmap (a new v4-additive relfile chain). No rows move, n_vectors and the cell directory are untouched, cells stay contiguous, has_ivf() stays true, and the scan keeps cell-restricting. The flat (lists = 0) path keeps the unchanged swap-remove. Tombstoned slots are masked out of the initial scan and every probe-widening refill, so deleted rows are never returned.
  • Observability for any residual fallback: a throttled scan-time WARNING (once per backend per index, with a HINT: REINDEX) and a new SQL function turbovec.index_is_degraded(regclass) -> bool. write_meta_shrink_in_place now preserves lists and flips an ivf_degraded meta flag rather than blanking the IVF identity, so the cliff is detectable.
  • docs/PRODUCTION.md gains an IVF + VACUUM operational section.

Performance — 7.8× faster IVF k-means (Phase A-1)

A 200k×256-d / lists=448 build’s k-means training was ~295 s (scalar Lloyd, fixed 25 iters) — prohibitive at 5M+. Now GEMM-batched Lloyd assignment (each iteration’s nearest-centroid step is one V@Cᵀ cross-term GEMM, single-threaded Parallelism::None + exact top-2 scalar tie-break) + convergence early-exit (KMEANS_TOL = 1e-6). Training 295 s → 38 s = 7.8×; per-iteration centroids byte-identical to the scalar path, determinism preserved. Training cost is bounded by the 256×lists reservoir sample regardless of corpus size, so 5M builds train in low-minutes. Build-internal; no surface or wire change.

Migration

No REINDEX. Wire stays v4; the tombstone bitmap and ivf_degraded flag are additive — pre-1.11.0 v4 indexes read as not-degraded / no-tombstones. The new index_is_degraded() function is registered by ALTER EXTENSION pg_turbovec UPDATE TO '1.11.0';.

Tests

187 → 193 on pg16 (+ivf_survives_vacuum, ivf_tombstoned_rows_not_returned, ivf_degradation_is_observable, fast-k-means + page.rs wire-format coverage). drift-check clean.

[1.10.1] — 2026-06-16

Bench-results-only release. Wire format unchanged from v1.10.0 (MetaPageData::version = 4); no REINDEX. Zero source-code change.

Benchmark — IVF warm-p50 on AVX2

Records the AVX2 IVF warm-p50 measurement confirming the IVF cell-skipping latency win that meh (pre-AVX2 scalar fallback) could not produce. Host floki (Intel Core Ultra 7 258V, AVX2), v1.10.0 release build, 200k × 256-d, lists = 448, 4-bit, warm cache, 50 timed queries per probes:

probes warm p50 vs full scan
4 0.74 ms 5.4× faster
16 0.78 ms 5.1× faster
448 (= lists, full exact scan) 3.97 ms baseline

At probes = 16, ~5× faster than the full exact scan, on AVX2. The IVF latency win is real on AVX2 hardware. Honest caveat: recall@10 = 1.000 at all probes in this run is an artifact of the synthetic corpus’s strong cluster structure, not a general guarantee — the host-independent recall-vs-probes frontier (v1.10.0) is the honest recall/probes trade-off. A full isolated 1M+ × 1024-d sweep on a quiet AVX2 host remains future work.

Files: benches/results/ivf_warmp50_floki_avx2_2026-06-16.json, docs/BENCHMARKS.md (“IVF warm-p50 (AVX2)” section).

[1.10.0] — 2026-06-16

Adds the IVF coarse-quantizer layer — a real sublinear ANN structure over the quantized codes. First wire-format change since v1.4.0 (MetaPageData::version 3 → 4), but existing v3 indexes do NOT need a REINDEX: a v1.10.0 binary reads a v3 index as a flat (lists = 0) index. Only users who opt into IVF rebuild. .

Why

The v1.9.1 AVX2 benchmark established that pg_turbovec’s flat O(n·dim) quantized scan is ~490× slower than pgvector HNSW at 1M×1024-d. IVF partitions the corpus into lists Voronoi cells (coarse k-means centroids) and scans only the probes nearest cells per query, dropping query work to roughly (probes/lists) of the corpus — the architectural path to a competitive latency story while keeping the 10–15× storage win.

Added — IVF (opt-in)

  • WITH (lists = N) reloption (default 0 = flat / today’s exact scan; N = number of coarse cells, recommended ≈ sqrt(n)).
  • WITH (assign_dups = M) reloption (default 1 = single assignment; M > 1 = soft assignment: boundary vectors stored in their top-M nearest cells to raise recall@10 at a fixed probes).
  • turbovec.probes GUC (default 8) — cells scanned per query; the recall/latency dial (the ivfflat.probes / hnsw.ef_search analogue). probes >= lists reduces exactly to the flat exact scan.
  • turbovec.max_probes GUC (default 64) — under iterative_scan = relaxed_order, a selective WHERE filter that under-returns triggers probe-WIDENING (scan more cells) up to this cap; the ivfflat.max_probes analogue.

How it composes

  • Iterative scan (v1.8.0): refill widens probes for IVF indexes (vs growing k for flat).
  • Oversampling (v1.9.0): widens the candidate set within the probed cells.
  • Reorder queue: exact-distance recheck, unchanged.
  • turbovec’s SIMD mask SKIPS scan work (block-level early-exit), and cells are stored contiguous, so fewer probed cells = real latency reduction.

Performance / determinism

  • The IVF build (k-means + assignment) is GEMM-batched (corpus rotation as block @ R^T, cell assignment via a V @ C^T cross-term GEMM + scalar top-2 tie-break). Without this the per-vector scalar loops were ~1012 FLOPs and a 1M build ran 60+ min. Single-threaded gemm keeps it bit-deterministic.
  • Builds are deterministic: same table + same lists/assign_dups ⇒ byte-identical relfile (seeded k-means++, IVF_SEED).
  • Recall-vs-probes frontier (host-independent; recall is CPU-independent): on a hard random 16k×64-d corpus, probes=16 → R@10 0.53 skipping 85% of blocks, probes=lists → 1.000. Real clustered embeddings reach high recall at far lower probes. Absolute AVX2 warm-p50 latency is deferred to a quiet arnold window (meh is pre-AVX2). See docs/BENCHMARKS.md.

Migration

No REINDEX for existing (v3, flat) indexes — they read as lists = 0 under the v1.10.0 binary. Opt into IVF by rebuilding with WITH (lists = N). ALTER EXTENSION pg_turbovec UPDATE TO '1.10.0'; registers the new reloptions + GUCs. MetaPageData::version is 4; EXPECTED_WIRE_FORMAT_VERSION = 4.

Tests

150 → 185 on pg16 (IVF build/scan/soft-assign/determinism/ probes-frontier coverage; distinct-id assertions throughout). drift-check clean.

[1.9.1] — 2026-06-15

Bench-results-only release. Wire format unchanged from v1.9.0 (MetaPageData::version = 3); no REINDEX needed. Zero source-code changes — this release bundles the AVX2 latency-frontier benchmark and the honest positioning correction it produced.

Benchmark — AVX2 latency frontier on arnold

The latency numbers meh (a pre-AVX2 Xeon) physically could not produce. Run on arnold (i9-12900H, AVX2), isolated via taskset -c 2-5 CPU-pinning to dedicated P-cores with per-batch contention measurement (observed 1-min load ≤ 1.05 throughout, zero CPU steal, no contended batches, no re-runs). Cohere wikipedia 1M × 1024-d, 1000 held-out queries, brute-force exact GT.

  • Correctness on the AVX2 SIMD path: recall@10 = 1.000 (the AVX2 kernel, not just meh’s scalar fallback, is correct on v0.9.0).
  • The hard truth: pg_turbovec loses to HNSW on latency by ~490× at 1M rows — warm p50 ~2552 ms (flat O(n·dim) quantized scan, recall 1.000) vs pgvector HNSW ~5 ms (sublinear graph, recall 0.96). AVX2 makes the correct scan ~15–25× faster than meh’s scalar fallback (2.5 s vs 41.6 s), but a 1M-row full scan is seconds, not ms, by design.
  • Retracted: the earlier “26.8 ms on meh / we win 2.3× warm p50” claim. That came from the pre-AVX2 scalar-fallback bug (fast-but-WRONG, fixed in v1.7.3) and never represented correct behaviour.

Positioning correction

docs/PARITY_GAPS.md and updated to the honest scoreboard. pg_turbovec’s durable wins are storage (10–15× smaller), exact recall (1.000 vs HNSW’s ~0.96), and build memory — NOT query latency at scale. Honest positioning: “best storage efficiency + exact recall for PG vector search where an O(n) scan fits the latency budget,” NOT “beat HNSW on every axis.” The architectural path to a latency story at scale is an IVF / coarse-quantizer layer (turning the O(n) scan into O(n/nlist + probes)) — a planned future major arc.

Files

  • benches/results/latency_frontier_arnold_cohere_1m_v1_9_0_2026_06_15.json
  • benches/scripts/vectordbbench/sweep_latency_isolated.py
  • docs/BENCHMARKS.md (arnold AVX2 section)
  • docs/PARITY_GAPS.md (corrections)

[1.9.0] — 2026-06-15

Oversampling (tunable recall), test-coverage hardening, and the first published head-to-head benchmark. Wire format unchanged (MetaPageData::version = 3); no REINDEX — ALTER EXTENSION pg_turbovec UPDATE TO '1.9.0'; suffices. The one new GUC defaults to a no-op.

Added — turbovec.oversample (differentiator #5)

Turns quantization from a fixed accuracy point into a tunable recall lever, matching Qdrant’s oversampling / VectorChord’s rerank.

  • turbovec.oversample (float, default 1.0, range 1.0..=100.0): the scan fetches ceil(search_k * oversample) quantized candidates and the executor’s reorder queue (xs_recheckorderby = true) trims to the exact top-k. Widening the candidate set recovers true neighbours the lossy quantized ranking placed just outside search_k.
  • No separate rescore path: oversampling + the always-on reorder queue together ARE the rescore mechanism (the reorder queue already re-ranks by exact full-precision distance). Measured: recall@10 climbs 0.81 (oversample 1.0) → 1.0 (oversample 4.0) on a 4-bit / 3000×64 corpus; latency rises ~linearly.
  • Composes with iterative scan: oversample sets the initial k; refill doubles from there, capped by max_scan_tuples.
  • Default 1.0 is identical to v1.8.0 behaviour.

Testing — scale + distinct-id + recall-floor regression guards

The pre-AVX2 wrong-results bug (fixed in v1.7.3) shipped because no test exercised more than ~2000 rows or asserted distinct result ids. Closed those gaps:

  • Medium-scale (20k×128-d) recall-floor #[pg_test] per bit_width {2,3,4}, with a brute-force ground-truth comparison.
  • assert_distinct_ids on EVERY ANN-scan test — the cheapest guard against the whole wrong-ranking bug class (a duplicate-id assert would have caught the pre-AVX2 bug instantly).
  • docs/TESTING.md documenting coverage + honest gaps: CI is AVX2-only (the scalar fallback runs only in turbovec’s upstream tests + pre-AVX2-host validation on turbovec bumps); unit tests cap at 20k rows (the benchmark is the large-scale evidence); the 15 “ignored” items are benign ```ignore doctests.

Benchmark — first published head-to-head (docs/BENCHMARKS.md)

Cohere wikipedia 1M × 1024-d (real embeddings, 1000 held-out queries, brute-force GT) vs pgvector HNSW, with a full reproducible harness under benches/scripts/vectordbbench/.

  • recall@10 = 1.000 on the fixed v1.8.0+ build at every config — the same pre-AVX2 host scored 0.0 on the old v1.7.1 build, so this is the definitive confirmation that the pre-AVX2 fix works on real embeddings at scale.
  • Storage: pg_turbovec 4-bit 7.6× smaller (1026 MB vs HNSW 7806 MB), 2-bit 15.2× smaller (512 MB). Build 1.9–2.1× faster.
  • pgvector HNSW frontier (its own SIMD): R@10 0.849/9.4 ms (ef40) → 0.979/20.1 ms (ef400).
  • pg_turbovec latency frontier is DEFERRED to an AVX2 host. The bench host meh is a pre-AVX2 Ivy Bridge Xeon; turbovec takes its scalar fallback (~1000× slower than its AVX2/AVX-512 kernels: ~42–69 s/query full-corpus scan). EXPLAIN confirmed Index Scan (not a seq-scan artifact). Latency/QPS benchmarks require an AVX2+ host (arnold); see BENCHMARKS.md for the explicit TODO and full caveats. (Updated AGENTS.md bench-host guidance accordingly — SIMD class matters more than RAM for turbovec latency.)

Migration

No migration; no REINDEX. On-disk format byte-identical to v1.7.x / v1.8.x. The new turbovec.oversample GUC defaults to 1.0 (no-op). ALTER EXTENSION pg_turbovec UPDATE TO '1.9.0'; resolves against the empty migrations/014_pg_turbovec_v1.9.0.sql.

Tests

142 → 150 on pg16 (+5 oversampling, +3 recall-floor; distinct-id assertions added to existing tests). drift-check clean.

[1.8.0] — 2026-06-15

Four competitive-parity features in one minor release. Wire format unchanged (MetaPageData::version = 3); no REINDEX needed — ALTER EXTENSION pg_turbovec UPDATE TO '1.8.0'; is sufficient. All four additions are scan-side, build-side, or additive SQL surface; none touch the on-disk relfile layout. .

Added — iterative index scan (parity gap #1, the correctness fix)

The one true correctness gap vs pgvector. amgettuple used to run a single search_k-sized batch and return false when drained, so a selective WHERE filter ORDER BY emb <=> q LIMIT k silently under-returned (e.g. 3 rows when 10 were asked for) — exactly what pgvector shipped hnsw.iterative_scan (0.8.0) to fix.

  • When the executor exhausts the candidate batch and the filter hasn’t been satisfied, the scan re-runs the turbovec search with a doubled k and feeds the new candidates, capped by a new turbovec.max_scan_tuples GUC (default 20000, matching pgvector’s hnsw.max_scan_tuples).
  • Controlled by turbovec.iterative_scan — an enum GUC off | relaxed_order (default relaxed_order). strict_order is deferred; our existing reorder-queue model (xs_recheckorderby = true) already restores exact per-tuple ordering on top of relaxed_order.
  • Dedup across refills via a returned-TID HashSet (turbovec’s search isn’t a stable prefix across k due to an unstable sort on score ties; the set is robust and bounded by max_scan_tuples).
  • Regression test demonstrates off under-returns and relaxed_order returns the full LIMIT.

Added — parallel index build (parity gap #2)

pgvector parallelises HNSW/IVFFlat builds across max_parallel_maintenance_workers; pg_turbovec’s ambuild was single-threaded.

  • Option B (rayon): the CPU-heavy encode + SIMD-repack phases (which dominate build CPU, not the heap scan) are parallelised over heap-scan chunks via a rayon pool. Chunks are processed in heap-scan order then concatenated deterministically.
  • New turbovec.build_parallelism GUC (default 0 = derive from max_parallel_maintenance_workers + 1; positive pins the pool).
  • Byte-for-byte identical relfiles regardless of thread count — asserted by a unit test — so the wire format and any reproducibility guarantees hold. Memory stays bounded by the Phase W maintenance_work_mem cap.

Performance — cold-scan latency (parity gap #3)

Cold-scan p50 was ~1256 ms (1 M × 1536-d) vs HNSW’s ~100 ms.

  • Lazy id_to_slot on the read path. Profiling the per-backend cache-fill showed the dominant residual term — once Phase P pre-baked the SIMD-blocked layout and Phase R-2 persisted the rotation — was the id_to_slot: HashMap<u64, usize> that IdMapIndex::from_id_map_parts* builds eagerly (~50 ms at 200 k rows, linear in n). The index-AM scan path never reads id_to_slot (search returns slots, mapped via the slot_to_id Vec). The scan path now installs a lightweight cache::ReadOnlyIndex (no HashMap); the build is deferred to the first aminsert/remove. A read-only / pooled-connection backend that only scans never pays it. Read-only constructor: ~50 ms → ~0 ms.
  • Key correctness test mutation_after_readonly_scan_is_correct verifies the deferred HashMap builds correctly on first insert.
  • Deferred follow-ups (see docs/PARITY_GAPS.md § cold scan): read-path mmap of the codes/scales/ids chains; a header-gap-free on-disk layout for true zero-copy mmap (VERSION 3 → 4, a future minor); a cross-backend DSA/DSM shared cache.

Added — || concat + halfvec arithmetic (parity gap #4)

pgvector has || concat for vector+halfvec and +/-/* for both; pg_turbovec had +/-/* for vector only.

  • turbovec.vector || turbovec.vector -> vector (concat)
  • turbovec.halfvec || turbovec.halfvec -> halfvec (concat)
  • turbovec.halfvec +/-/* element-wise (Hadamard for *)
  • Matches pgvector overflow semantics (error on non-finite result) and dim-mismatch errors.

Migration

No migration needed; no REINDEX. The on-disk relfile format is byte-identical to v1.7.x. Drop in the new shared library, restart, scan; existing indexes work unchanged. The new GUCs default to the pgvector-equivalent behaviour (iterative_scan = relaxed_order). ALTER EXTENSION pg_turbovec UPDATE TO '1.8.0'; resolves against the empty migrations/013_pg_turbovec_v1.8.0.sql.

Tests

123 → 142 on pg16 (+19: iterative-scan, parallel-build, cold-scan, and arithmetic-parity coverage). drift-check clean.

[1.7.3] — 2026-06-15

Fixed — pre-AVX2 x86_64 wrong-results bug (turbovec fork → v0.9.0)

Wire format unchanged from v1.6.0 / v1.7.x (MetaPageData::version = 3); no REINDEX needed to upgrade.

  • Root cause. The Phase A1 “regression” (index ORDER BY emb <=> probe LIMIT N returning the same id N times at 10 M scale on the meh bench host) was traced to an upstream turbovec kernel bug, not pg_turbovec. The pinned turbovec v0.7.0-era fork (6e80a59) had a scalar fallback that, on x86_64 CPUs without AVX2, read the perm0-interleaved (FAISS-style) SIMD code layout as if it were sequential — producing silently-wrong / repeated top-k. meh is an Intel Xeon E5-2697 v2 (Ivy Bridge, 2013): avx but no avx2, so it hit the buggy path. AVX2 (Haswell 2013+), AVX-512, and ARM NEON hosts always took a correct SIMD path — which is why the bug never reproduced on AVX2 dev boxes or arnold, only on meh.
  • Fix. Upstream turbovec fixed this in PR #108 (issue #106, “V5”), released in v0.8.0, adding a correct score_query_into_heap x86_64 scalar fallback plus a FORCE_SCALAR_FALLBACK regression test. v1.7.3 upgrades the gburd/turbovec fork from the v0.7.0-era 6e80a59 to a fork rebased onto upstream v0.9.0 (d3d468e on branch pg_turbovec-integration-v0.9.0).
  • Also brought in, inert here:
    • TQ+ per-coordinate calibration fields, constructed as identity (empty) on the relfile path — no recall change, no wire change in v1.7.3. Persisting them for a recall gain is a future minor release (VERSION 3 → 4 + REINDEX).
    • Security hardening: MAX_DIM = 65536, NaN/Inf/huge-magnitude input rejection, checked-mul .tv/.tvim loaders.
  • Zero pg_turbovec source churn — the upgrade is a Cargo.toml rev bump only; the fork kept the prepare_eager alias and passes TQ+ through internally so every from_id_map_parts* call site is unchanged.
  • Toolchain note. turbovec v0.9.0 uses avx512 target_features requiring Rust ≥ 1.89. Builds with the default stable toolchain (1.95). See AGENTS.md for the refreshed openblas store path and the -fuse-ld=bfd linker note.
  • Tests: 123/123 on pg16. drift-check clean.

Migration

No migration needed; no REINDEX. The on-disk relfile format is byte-identical to v1.6.x / v1.7.x. Drop in the new shared library, restart, scan. Pre-AVX2 x86_64 users specifically should upgrade to clear the wrong-results bug and can drop any SET enable_indexscan = off; workaround. ALTER EXTENSION pg_turbovec UPDATE TO '1.7.3'; resolves against the empty migrations/012_pg_turbovec_v1.7.3.sql.

[1.7.2] — 2026-05-27

Added — Phase Y: automated upgrade-matrix validation

Wire format unchanged from v1.6.0 / v1.7.0 / v1.7.1 (MetaPageData::version = 3); no REINDEX needed to upgrade. v1.7.2 is a test-only patch release.

Production-confidence foundation: previously the upgrade matrix in docs/UPGRADING.md and the is_legacy_v{1,2}() detection predicates in src/index/page.rs were promises with no automated end-to-end test. Phase Y closes that gap.

  • alter_extension_path_140_to_171_runs_clean (new #[pg_test]) replays every migrations/0NN_pg_turbovec_*.sql from v1.3.0 onward against the live test cluster. Catches a release engineer who lands a syntactically-broken DDL change in one of the post-v1.3 migration files (which are otherwise intentionally empty).

  • ambeginscan_errors_on_legacy_v1_meta and ambeginscan_errors_on_legacy_v2_meta (new #[pg_test]s) build a real v1.7.2 index, forge the meta-page version byte to 1 or 2 via the new cfg-gated relfile::force_meta_version() helper, and assert that ambeginscan ERRORs at first scan with the documented primary message + REINDEX INDEX HINT. Exercises the Phase Q (v1.3.0) + Phase R-2 (v1.4.0) hard migration boundaries without having to keep pre-v1.4 binaries around.

  • alter_extension_update_chain_resolves (new #[pg_test]) asserts the installed extension version matches Cargo.toml, catching version-number drift between Cargo.toml, pg_turbovec.control, and the migration file naming convention.

  • migration_files_cover_documented_versions (new #[pg_test]) asserts the set of migrations/*.sql sigils matches the documented release history. If you tag a new release without adding the migration file, this test fails before the bad tag escapes CI.

  • scripts/drift-check.sh § 9 (new check) cross-checks migrations/*.sql against the From column of the migration matrix in docs/UPGRADING.md. Catches release-time drift between adding a tag and forgetting to add the migration file.

  • relfile::force_meta_version() (new test-only helper) is gated on cfg(any(test, feature = "pg_test")) and patches the version byte of the meta page in place via a GenericXLog record. Only the pgrx test suite (and a future feature-gated build) can reach it; production builds never compile it.

Migration

No migration needed; rebuild not required. The on-disk format is byte-identical across v1.6.0 / v1.7.0 / v1.7.1 / v1.7.2. Drop in the new shared library, restart, scan; existing indexes built under any of these versions continue to work unchanged.

[1.7.1] — 2026-05-27

Reverted — Phase W-2 split-write design (regression)

Wire format unchanged from v1.6.0 / v1.7.0 (MetaPageData::version = 3); no REINDEX needed to upgrade or downgrade between any of these. v1.7.1 is a behaviour-only revert.

  • Phase W-2 (v1.7.0) reverted. Validation on meh (24-core, 125 GiB RAM NixOS host, head commit a289870) at 10 M × 1536-d × 4-bit showed the split-write ambuild path introduced in v1.7.0 made the build 53% slower (5052 → 7748 s), used 2.7 GiB of swap (vs 0 in v1.6.0), and slightly raised peak RSS (22.5 → 23.04 GiB). The predicted ~15 GiB peak never materialised. Full data: benches/results/phase_w_2_validate_meh_10m_2026_05_27.json.

    metric v1.6.0 v1.7.0 (W-2) v1.7.1 (revert)
    Peak RSS (GiB) 22.5 23.04 22.5 (= v1.6.0)
    Swap used (GiB) 0 2.67 0 (= v1.6.0)
    Build time (s) 5,052 7,748 5,052 (= v1.6.0)
  • Why Phase W-2 didn’t work. The hypothesis was that dropping the ~7.7 GiB row-major packed_codes Vec mid-finalise (via IdMapIndex::take_packed_codes()) would shave the peak RSS by ~7.7 GiB. It didn’t, because the intervening write_packed_phase pins those bytes in shared_buffers before take_packed_codes() runs, and ps -o rss counts mapped shared memory as part of the backend’s resident set. The 7.7 GiB of “freed” heap simply migrated to pinned shared memory; same RSS budget, plus the cost of an extra GenericXLog flush phase. See .7.1" for the full analysis.

  • What was reverted.

    • src/index/relfile.rs::write_full_inner — restored to the v1.6.0 single-pass batched-GenericXLog flow: meta page, then codes / scales / ids chains, then blocked / rotation chains, then RelationTruncate for shrinking REINDEX.
    • src/index/build.rs::ambuild — restored to the v1.6.0 sequence: prepare_eager() first, then a single write_full_with_prepared call. The take_packed_codes() call is dropped from this code path.
    • src/lib.rs — ambuild_drops_packed_codes_before_blocked_write renamed to ambuild_round_trip_after_phase_w_2_revert and kept as a generic ambuild round-trip smoke (still passes via the v1.6.0 code path).
  • What was kept.

    • relfile::write_packed_phase, relfile::write_blocked_phase_and_meta, and relfile::PackedPhaseLayout remain in the source as parked dead code, marked #[allow(dead_code)]. They have no callers after the revert but may be useful for a future Phase W-3 attempt that takes a different angle (e.g. streaming pack::repack).
    • The turbovec fork pin at rev 6e80a59f473292cc9e04d575ba1596f3e23321c5 (turbovec 0.7.0) stays. IdMapIndex::take_packed_codes() on the fork is harmless additive API; we just don’t call it.

Migration

No migration needed; rebuild not required. The on-disk format is byte-identical across v1.6.0 / v1.7.0 / v1.7.1. Drop in the new shared library, restart, scan; existing indexes built under any of these versions continue to work unchanged.

[1.7.0] — 2026-05-27

Added — mid-finalise drop of packed_codes in ambuild (Phase W-2)

Wire format unchanged from 1.6.x (MetaPageData::version = 3); no REINDEX needed to upgrade. v1.7.0 is a build-side change only: the on-disk index format is byte-identical to v1.6.x.

  • Reorder finalisation writes so packed_codes and blocked are never co-resident. Phase W (v1.6.0) capped the heap-scan staging buffer, dropping peak ambuild RSS from 121 GiB to 22.5 GiB at 10 M × 1536-d on meh. The remaining 22.5 GiB peak was IdMapIndex’s row-major packed_codes (~7.7 GiB) plus the SIMD-blocked derived layout (~7.5 GiB) plus allocator slack + GenericXLog page-assembly buffers, all alive during the single-call relfile::write_full_with_prepared flush. Phase W-2 splits that call into two phases:

    1. relfile::write_packed_phase streams packed_codes, scales, and slot_to_id to relfile pages while packed_codes is the only large in-memory Vec.
    2. IdMapIndex::prepare_eager() materialises the SIMD-blocked layout, codebook, and rotation matrix (transient peak: packed + blocked).
    3. IdMapIndex::take_packed_codes() (new turbovec 0.7.0 API) swaps the row-major Vec out and shrink_to_fits it; the OnceLock-backed blocked cache is unaffected.
    4. relfile::write_blocked_phase_and_meta streams the blocked
      • rotation chains and stamps the meta page LAST.

    Expected peak at 10 M × 1536-d: ~15 GiB (down from 22.5 GiB). Combined with Phase W’s 121 → 22.5 GiB cut, that’s an 8× total reduction vs pre-Phase-W. Validation on meh at 10 M scale is a follow-up bench phase; the v1.7.0 code change ships with local unit-test coverage of the split write (ambuild_drops_packed_codes_before_blocked_write).

  • Meta page is now written LAST. write_full_inner used to write the meta page first and the chains second, which left a crash window where block 0 referenced not-yet-written chain pages. v1.7.0 routes both the legacy write_full / write_full_with_prepared and the new split path through write_blocked_phase_and_meta, which writes the meta page AFTER all chain pages — matching the standard PG hash/gist AM “meta page is the atomic-complete signal” pattern. A crash before the meta-page WAL record commits leaves block 0 in its previous state (zero-filled for fresh build, previous meta for REINDEX), and ambeginscan rejects the index as empty/legacy. No on-disk format change — readers never observed the intermediate state in any released version.

  • Turbovec fork bump 0.6.0 → 0.7.0 (rev 6e80a59f473292cc9e04d575ba1596f3e23321c5, branch pg_turbovec-integration on gburd/turbovec). Adds TurboQuantIndex::take_packed_codes(&mut self) -> Vec<u8> and the matching IdMapIndex::take_packed_codes. Additive minor; no breaking changes for embedders that don’t call the new API.

  • Phase W-3 deferred. The remaining ~15 GiB peak is dominated by the SIMD-blocked Vec materialised by prepare_eager() plus GenericXLog page-assembly slack. Dropping the blocked peak further (to ~10 GiB) would require streaming pack::repack so the blocked layout never has to be fully resident; that’s substantial turbovec internals work and is out of scope for v1.7.0.

Migration

No migration needed; rebuild not required. The on-disk format is byte-identical to v1.6.x. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged.

[1.6.1] — 2026-05-27

Bench-results-only release. Wire format unchanged from 1.6.0; no REINDEX needed.

  • Phase W validation on meh (commit 8efb89c). Re-ran the Phase V 10M × 1536-d build against v1.6.0 to confirm the streaming ambuild change actually drops peak RSS as designed. Result: 121 GiB → 22.5 GiB peak (5.4× reduction), 60 GiB → 0 GiB swap usage. Build time within 0.1 % of Phase V (5048 → 5052 s); index size unchanged (15 GiB); warm-scan p50 identical (21.2 ms vs Phase V’s 21–49 ms band). The remaining 22.5 GiB peak is IdMapIndex’s row-major packed_codes (~7.7 GiB) + the SIMD-blocked prepared layout (~7.5 GiB) + allocator slack + PG backend baseline, all held simultaneously during end-of-build finalisation. Tracked as Phase W-2. Files: benches/results/phase_w_validate_meh_10m_2026_05_27.json, benches/results/phase_w_warm_sanity_meh_10m_2026_05_27.json, benches/results/build_tv_meh_10m_v1_6_0_2026_05_27.{log,psql.log,rss.tsv.gz}, docs/RECALL.md § 2.7 follow-up.

  • Phase X: RISC-V architecture comparison (commit a8fbd87). First non-x86 host bring-up. 100 k × 384-d synthetic on rv (RISC-V 64, 8 cores, 7.7 GiB RAM, Ubuntu 24.04 LTS): index 39 MB (5× compression), build 13.97 s, warm p50 242.64 ms (50-query stdev 0.73 ms — extremely tight). Verdict: arch_works. The latency multiplier vs x86 (~10–25× depending on corpus comparison) reflects turbovec’s AVX2/SSE inner loop falling back to scalar on RISC-V; RVV intrinsics are upstream-future work. Operational note for non-NixOS hosts: the postmaster needs LD_PRELOAD=libopenblas.so.0 because cblas_sgemm is a deferred symbol not in the .so’s NEEDED entries. Files: benches/results/recall_warm_rv_100k_v1_6_0_2026_05_27.json, docs/RECALL.md § 2.8 (new section).

Migration

No migration needed; rebuild not required. The on-disk format is byte-identical to v1.6.0. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged.

[1.6.0] — 2026-05-26

Added — streaming heap scan in ambuild (Phase W)

Wire format unchanged from 1.5.x (MetaPageData::version = 3); no REINDEX needed to upgrade. v1.6.0 is a build-side change only: the on-disk index format is byte-identical to v1.5.x.

  • Build-time memory cap. Phase V measured CREATE INDEX peak RSS at 121 GiB on a 10 M × 1536-d × 4-bit corpus on meh (24 cores, 125 GiB RAM), with 60 GiB of swap usage. The dominant offender was BuildState::flat: Vec<f32> in src/index/build.rs::ambuild accumulating the entire heap-scan output before passing it to IdMapIndex::add_with_ids. At 10 M × 1536-d that buffer alone is 61 GiB.
  • Phase W: stream the heap scan. BuildState now carries two bounded staging buffers (pending_flat, pending_ids) sized off maintenance_work_mem. Every chunk_rows rows the callback flushes into IdMapIndex::add_with_ids and shrink_to_fits the buffers back to zero capacity, returning the bytes to the allocator. A trailing flush after the heap-scan loop drains the partial chunk.
  • Chunk sizing formula (in BuildState::compute_chunk_rows): chunk_bytes = min(maintenance_work_mem_kb * 1024 * 3 / 4, 1 GiB); chunk_rows = max(chunk_bytes / (dim * 4), 1). The GUC is read in kilobytes (PG convention; the global is pg_sys::maintenance_work_mem: c_int whose unit is KB despite the name). 75% allocation leaves headroom for the IdMapIndex’s own growth; the 1 GiB ceiling caps the staging buffer even with a SET maintenance_work_mem = '8GB'.
  • Expected peak at 10 M × 1536-d: ~16 GiB (down from 121 GiB). Validation on meh at 10 M scale is a follow-up phase — the v1.6.0 code change ships with local unit-test coverage of the streaming path; the multi-hour memory-cap validation runs separately.
  • Phase W-2 deferred. The IdMapIndex still holds packed_codes (~7.7 GiB at 10 M × 1536-d × 4-bit) in memory alongside blocked_codes after prepare_eager(). Dropping it would save another ~7.7 GiB at peak but requires a turbovec fork API change (IdMapIndex::drop_row_major_codes(&mut self) on branch pg_turbovec-integration). Tracked as a follow-up; out of scope for v1.6.0.
  • One new #[pg_test]: ambuild_streams_heap_scan_under_maintenance_work_mem exercises the streaming path with maintenance_work_mem = '4MB' and a 1000-row table. Test count 116 → 117.
  • Docs. docs/UPGRADING.md migration matrix gets a 1.5.x → 1.6.0 no-op row; records the diagnosis, the formula, and the Phase W-2 follow-up parking lot.

Migration

No migration needed; rebuild not required. The on-disk format is byte-identical to v1.5.x. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged. ALTER EXTENSION pg_turbovec UPDATE TO '1.6.0'; resolves against the empty migrations/007_pg_turbovec_v1.6.0.sql.

This is a minor bump rather than a patch because the build-time memory profile is observably different: a host that used to OOM on 10 M × 1536-d will now succeed. That’s a behaviour change worth a minor even though no on-disk format changed.

[1.5.1] — 2026-05-26

Bench-results-only release. Wire format unchanged from 1.5.0; no REINDEX needed.

  • Phase U-1: cache works correctly. A debug-only tracepoint in cache::lookup confirmed 50/50 hits across a 50-query warm sweep (zero misses of any class). The Phase S agent’s hypothesis that the per-backend cache misses on every warm scan was wrong; what they saw in perf was the one-shot finalise_from_inner build during the cold-cache install, amortised over the sampling window. Tracepoint reverted before the build that produced the Phase U-2 measurements.
  • Phase U-2: Phase S delivers no win on RAM-rich hosts. On meh (24 cores, 125 GiB RAM), warm p50 is 26.8 ms mmap=on, 26.7 ms mmap=off (delta 0.15 ms = noise) at shared_buffers = 512 MB, search_k = 100. The buffer-manager bottleneck Phase S targets is invisible when free RAM ≫ index size because the OS page cache serves pread reads instantly. Phase S is at-worst-neutral on RAM-rich hosts; it may still help RAM-constrained hosts (the arnold re-bench at the original 31 GiB-RAM constraint remains the definitive Phase S validation).
  • The headline number that matters: pg_turbovec on a properly- RAMed host beats pgvector HNSW ef=40 on every measurable axis. meh’s 26.8 ms warm p50 is 2.3× faster than HNSW ef=40’s 61 ms, at 5× less storage and R@10 = 1.000 on the dbpedia-1M corpus. The 60–90 ms warm regime that motivated Phase R-2 / Phase S was an arnold-class (limited RAM) phenomenon, not a fundamental kernel ceiling.

Artefacts

  • benches/results/recall_warm_meh_v1_5_0_2026_05_26.json — full structured run with both configs + verdict.
  • benches/results/u2_meh_tv_4bit_warm_mmap_{on,off}.tsv — raw 50- sample TSVs.
  • — full method + result of the cache- miss tracepoint experiment.
  • docs/RECALL.md § 2.6 extended with the meh comparison.
  • docs/PARITY_GAPS.md warm-scan row updated.

[1.5.0] — unreleased

Added — mmap-based reads of the relfile’s static regions (Phase R-3)

Wire format unchanged from 1.4.x (MetaPageData::version = 3); no REINDEX needed to upgrade. v1.5.0 is a scan-side change only.

  • New code path: src/index/mmap_static.rs. The ambeginscan cache-fill path now mmap(MAP_PRIVATE)s the relation’s segment-0 file, walks the deterministic static chains (persisted SIMD-blocked codes, persisted rotation matrix, inline codebook) directly off the mapping, and skips PG’s buffer manager for those bytes. Halves the warm-scan cost when the index doesn’t fit in shared_buffers — the Phase R-3 diagnosis in docs/RECALL.md § 2.5.
  • New GUC: turbovec.mmap_static_blocked (default on). Set off per session to revert to the v1.4.x buffer-manager-only read path. See docs/ARCHITECTURE.md § 8.1 for the isolation contract.
  • Cache machinery extension: cache::insert_with_mmap. The Mmap handle is colocated on the Entry with the Arc<RwLock<IdMapIndex>> and dropped only after the index has been freed (drop order enforced by struct field order). Future zero-copy work (handing turbovec a borrowed slice into the mapping via the new from_id_map_parts_with_prepared_borrowed upstream API) relies on this ordering; v1.5.0 holds owned Vecs in the index so the contract is trivially satisfied today.
  • Upstream turbovec fork bump. turbovec is pinned to gburd/turbovec branch pg_turbovec-integration at commit c3c0528, which adds the Cow-based borrowed-cache constructors (from_parts_with_prepared_borrowed, from_id_map_parts_with_prepared_borrowed, PreparedCachesBorrowed). Six new upstream tests cover the borrowed/owned round-trip equivalence and lifetime contract (89 → 95 tests).
  • Three new #[pg_test]s: relfile_mmap_static_round_trip_matches_buffer_manager, relfile_mmap_static_concurrent_aminsert_recheck_corrects, relfile_mmap_static_cache_invalidation_drop_order. Test count 113 → 116.
  • Docs: docs/RECALL.md § 2.6 for the post-fix performance story; docs/ARCHITECTURE.md § 8.1 for the isolation contract (heap visibility + recheck-orderby as the MVCC backstops; concurrent aminsert / ambulkdelete / REINDEX worked examples); docs/PARITY_GAPS.md warm-scan row updated to reference v1.5.0 with arnold re-bench pending; docs/UPGRADING.md migration matrix gets a 1.4.x → 1.5.0 no-op row; README.md ## Performance operations note rewritten — shared_buffers no longer needs to be sized against the index size by default.

Dependency added

  • memmap2 = "0.9" for the MAP_PRIVATE RO mapping. No other dependency churn.

Wire format

  • No change. MetaPageData::version stays at 3, MIN_DECODE_VERSION stays at 1, and the wire_format_version_is_stable test continues to assert EXPECTED_WIRE_FORMAT_VERSION = 3.

[1.4.1] — 2026-05-26

Fix — stale rows in the parity scoreboard, plus drift-check tightening

No code changes in this release. Wire format unchanged from 1.4.0; no REINDEX needed.

  • docs/PARITY_GAPS.md scoreboard updated with two rows that had drifted three minor versions:
    • INSERT throughput row was still claiming “~200 ms / row, we lose 400×” — that pre-Phase-K v1.0.x number. Phase K landed in v1.1.0 with the deferred-commit pattern that delivers ~0.13 ms/row (4× faster than HNSW). Row is now accurate.
    • Recall on real ada-002 dbpedia-1M row was still “TBD”. Phase J measured R@10 = 1.000 in v1.1.0; the row is now populated with the actual number.
  • scripts/drift-check.sh §8 now flags scoreboard cells containing TBD or claiming “we lose Nx” without a same-row phase qualifier (e.g. “post-Phase-K”, “shipped in v1.1.0”). Verified by synthesising both failure modes on top of the v1.4.0 scoreboard. The drift-check script also keeps its existing v1.3.0 wire-format check (§7).
  • RELEASING.md pre-flight checklist grows two items: one for bash scripts/drift-check.sh and one for eyeball- reading the PARITY_GAPS scoreboard against the latest benches. drift-check §8 catches structural drift but can’t catch a row whose number is numerically wrong; the eyeball step is the backstop.

All guards aligned: Cargo.toml = 1.4.1, VERSION = 3 (no change from 1.4.0), EXPECTED_WIRE_FORMAT_VERSION = 3, drift-check clean.

[1.4.0] — 2026-05-25

Headline (Phase R-2): rotation matrix persisted in the relfile

The random orthogonal rotation matrix used by TurboQuant—a deterministic function of (dim, ROTATION_SEED) produced by QR decomposition of a dim x dim Gaussian random matrix—is now persisted alongside the existing prepared parts (centroids, boundaries, blocked layout). At dim = 1536 the lazy QR was the single hottest leaf of the warm-scan profile (~64.8% self time; see benches/results/profile_warm_v1_3_0_2026_05_25.json and ), and it ran once per fresh backend because the per-backend cache OnceLock was driven on first search instead of read off disk.

ambuild now drives IdMapIndex::rotation() after prepare_eager() and writes the row-major dim*dim f32 buffer (~9 MiB at 1536-d, negligible vs. the existing ~1.5 GiB index) into a new chain on the relfile. Backends opening the index pre-fill the rotation OnceLock from those bytes via the extended IdMapIndex::from_id_map_parts_with_prepared(…, rotation: Option<Vec<f32>>) constructor.

Expected impact: warm-scan p50 drops 50–200 ms toward the pgvector HNSW band on dbpedia-1M (1 M × 1536-d). A separate Phase R-3 run on arnold validates the production number; this release is the implementation + wire-format bump.

⚠️ BREAKING: hard migration boundary (v1.3.x indexes)

MetaPageData::version bumps 2 → 3 to add the new rotation_first / rotation_count / rotation_dim fields and the rotation chain. v1.4.0 binaries refuse to scan v2 (v1.3.x) indexes because the rotation chain offsets don’t exist on disk and the lazy QR was the hotspot we just eliminated. After upgrading:

ALTER EXTENSION pg_turbovec UPDATE TO '1.4.0';
REINDEX INDEX <every_turbovec_index>;

Without REINDEX, ambeginscan raises ERROR: turbovec index built under pg_turbovec ≤ 1.3 cannot be scanned by pg_turbovec 1.4+ with a HINT: Run REINDEX INDEX <name>;. The detection primitive is MetaPageData::is_legacy_v2() (mirrors the existing is_legacy_v1). The matrix in docs/UPGRADING.md documents the scripted path.

Migration

DO $$
DECLARE
    idx record;
BEGIN
    FOR idx IN
        SELECT n.nspname || '.' || c.relname AS qname
        FROM pg_class c
        JOIN pg_am a ON a.oid = c.relam
        JOIN pg_namespace n ON n.oid = c.relnamespace
        WHERE a.amname = 'turbovec'
    LOOP
        RAISE NOTICE 'reindexing %', idx.qname;
        EXECUTE 'REINDEX INDEX CONCURRENTLY ' || idx.qname;
    END LOOP;
END $$;

REINDEX INDEX CONCURRENTLY rebuilds without taking an AccessExclusiveLock so reads keep working during the migration. The new index is built first; the cutover swap is atomic.

Vendor turbovec patch

Three additive surfaces on top of the existing Phase P prepared-cache APIs (see vendor/turbovec/PATCH_NOTES.md for the full table):

  • TurboQuantIndex::rotation() -> &[f32] accessor mirroring centroids / boundaries / blocked_codes. Drives the existing rotation OnceLock and returns the row-major dim*dim matrix.
  • TurboQuantIndex::rotation_size(dim) -> usize const helper (dim * dim) so callers can preallocate the on-disk chain.
  • TurboQuantIndex::from_parts_with_prepared(…, rotation: Option<Vec<f32>>) and the matching IdMapIndex:: from_id_map_parts_with_prepared overload — Some pre-fills the rotation OnceLock, None falls back to the lazy QR (used during ambuild itself, when the matrix isn’t yet on disk). Tracked as a follow-up to upstream PR #70 (Codrai turbovec issue #70).

Source

  • src/index/page.rs: VERSION = 3. MetaPageData gains rotation_first / rotation_count / rotation_dim. plan_with_blocked takes a new rotation_bytes parameter; layout is meta → codes → scales → ids → blocked → rotation. Decode accepts v1, v2, v3 (older versions leave the new fields zero so is_legacy_v2() flags them).
  • src/index/relfile.rs: PreparedParts gains rotation: &'a [f32]. write_full_inner writes the rotation chain after the blocked chain. write_meta_shrink_in_place preserves rotation_first/count/dim across vacuum (the matrix is data-independent). New read_rotation() mirrors the existing read_blocked().
  • src/index/scan.rs: ambeginscan gains the is_legacy_v2() && n_vectors > 0 ERROR path next to the existing v1 path. amgettuple reads the rotation chain off disk and feeds it to IdMapIndex::from_id_map_parts_with_prepared as Some(rotation).
  • src/index/build.rs: ambuild calls idx.rotation() after prepare_eager() and threads it through PreparedParts. src/xact.rs: same edit on the deferred-commit flush path.
  • src/lib.rs: EXPECTED_WIRE_FORMAT_VERSION = 3. New relfile_legacy_v2_detection_primitive (mirrors the v1 test) and relfile_rotation_persisted (proxy for the warm-scan win: top-1 query through the prepared+rotation index must finish in <100 ms on a 100-row debug build).

Tests

113/113 default. +2 vs. v1.3.0 from the new rotation tests: relfile_legacy_v2_detection_primitive (mirrors the existing relfile_legacy_v1_detection_primitive) and relfile_rotation_persisted (proxy for the warm-scan win: asserts the rotation chain is on disk, the matrix is orthogonal to within roundoff, and a top-1 query through the prepared+rotation index finishes in <100 ms on a 100-row debug build).

Docs

  • docs/UPGRADING.md: new migration matrix row for 1.3.x → 1.4.0+, citing is_legacy_v2().
  • vendor/turbovec/PATCH_NOTES.md: “Phase R-2 follow-up: persisted rotation matrix” section documenting the four new surfaces.

[1.3.0] — 2026-05-25

Headline (Phase Q): one storage strategy, no flags

The SPI side-table (turbovec.am_storage) and its accompanying Cargo feature flags (relfile_storage, experimental_index_am) are gone. The relfile-resident page format — introduced as a preview in 1.1.0 (Phase L), proven correct end-to-end in Phase O-2, and brought up to parity with the side-table on cold-scan latency by Phase P (1.2.0) — is now the only storage strategy. The AM matches the conventions of every other PostgreSQL index AM (btree, gist, gin, hnsw, ivfflat).

Build flags reduce to just pg<N>:

cargo pgrx test pg16   # no --features needed
cargo build --no-default-features --features pg16

⚠️ BREAKING: hard migration boundary

Any existing turbovec index built under v1.0.x..v1.2.0 has either (a) only a side-table row and an empty main fork, or (b) a v1 (Phase L preview) relfile meta layout that lacks the persisted SIMD-blocked layout + Lloyd-Max codebook Phase P relies on. Both states are unrecoverable from the running binary. After upgrading:

ALTER EXTENSION pg_turbovec UPDATE TO '1.3.0';
REINDEX INDEX <every_turbovec_index>;

Without REINDEX, ambeginscan raises an ERROR (no longer a NOTICE) explaining the situation. This is deliberate — a half-broken state can’t silently return zero rows.

The extension install / upgrade SQL drops turbovec.am_storage if it still exists (legacy state from a previous install).

Removed

  • src/index/persist.rs deleted (the SPI side-table reader / writer, ~250 lines).
  • aminsert_sidetable and ambulkdelete_sidetable deleted.
  • The turbovec.am_storage table and the extension_sql! block that created it.
  • The relfile_storage Cargo feature (default-on, no longer togglable).
  • The experimental_index_am Cargo feature (the AM has been default-on since v0.9; the “experimental” name was stale).
  • All #[cfg(feature = "relfile_storage")] and #[cfg(feature = "experimental_index_am")] gates throughout src/.
  • Migration NOTICE in ambeginscan (replaced by the hard ERROR above).
  • Stale tests that read am_storage.payload / am_storage. n_vectors directly. Where the test was exercising generic AM behaviour (“CREATE INDEX succeeds and the heap is queryable”), it was kept and the assertion was switched to count(*) on the heap. Where it was strictly side-table- specific (aminsert_deferred_persist_bulk), it was deleted in favour of its relfile twin (relfile_aminsert_deferred_ commit_bulk) which now runs unconditionally.

Updated

  • src/cache.rs and src/xact.rs: the cfg-selected flush sink (sidetable persist::save vs relfile write_full) collapses to relfile only.
  • src/index/cost.rs: amcostestimate reads n_vectors / dim / bit_width straight off the relfile meta page (block 0) instead of via SPI on turbovec.am_storage.
  • Cargo metadata bumped 1.2.0 → 1.3.0; pg_turbovec.control bumped to default_version = '1.3.0'.
  • migrations/005_pg_turbovec_v1.3.0.sql documents the upgrade path and is the new install reference mirror.
  • Documentation: docs/PARITY_GAPS.md, docs/ARCHITECTURE.md, docs/PG_VERSION_SUPPORT.md, and README.md updated to reflect the post-Phase-Q crate layout, retired feature flags, and post-Phase-P cold-scan numbers (1.26 s p50, 21× speedup vs. pre-fix).

Tests

109/109 across pg13, pg16, pg18 (sample of the matrix). Was 94/94 default + 104/104 relfile_storage in 1.2.0; the two sides converge on 109 now that there are no gates: 94 default tests + 6 relfile tests (cold-scan, cold-vs-warm, WAL, init fork, ambulkdelete walk, prepared-layout) + 4 Phase P tests (prepared layout, cache hits, etc.) + 1 Phase Q test (legacy v1 detection primitive) + 4 sidetable-specific tests dropped.

Phase O-3 cold-scan re-validation

Phase P’s pre-baked SIMD-blocked layout + Lloyd-Max codebook shipped in 1.2.0 brought cold-scan p50 on dbpedia-1M (1 M vectors x 1536-d, OpenAI embeddings, arnold) from ~26.5 s to 1.26 s p50 — a 21× speedup over the pre-fix v1.0.x side-table path. The full-cluster cold-scan story now matches pgvector HNSW within an order of magnitude, and the relfile-resident architecture wins on every other axis (build time, on-disk size, WAL volume, recall).

[1.2.0] — 2026-05-25

Phase L hardening complete (5 of 6 items)

The relfile-resident page format introduced as a preview in 1.1.0 (--features relfile_storage) is now production-grade on five of the six hardening items from :

  1. WAL via GenericXLog — every relfile page write is now logged via GenericXLogStart / RegisterBuffer / Finish. A crash before checkpoint correctly replays via standard PG WAL. (Phase N-B, commit 9ee405d)

  2. ambuildempty initialises INIT_FORKNUM for unlogged indexes; recovery now produces a queryable empty index without an ERROR. (Phase N-B)

  3. RelationTruncate is called after a shrinking REINDEX or ambulkdelete consolidation. (Phase N-B)

  4. Phase K’s deferred-commit pattern applied to the relfile path. aminsert_relfile now mutates the cached Arc<RwLock<IdMapIndex>> in memory and defers the relfile page write to the PreCommit xact callback. Bulk INSERT of 1 k rows: was minutes (full-rewrite per row) → now < 5 s. (Phase N-C, commit d4a469b)

  5. v1.0.x → v1.2 migration HINT in ambeginscan. When a relfile_storage-built binary opens an index whose main fork is empty but the side-table has n_vectors > 0, emit a NOTICE with HINT: Run REINDEX INDEX <name>;. Without this users would silently see zero rows. (Phase N-C)

Phase L hardening remaining (1 of 6)

  1. ambulkdelete walks pages instead of rebuilding. Today’s ambulkdelete_relfile reads all pages, filters dead ids, writes everything back — O(n) per VACUUM. Walk-and-mark would bring this to O(deleted_rows). Tracked for v1.3 in § 6.

Drift cleanup

docs/ARCHITECTURE.md rewritten to v1.1.0 reality: status banner updated, future-tense “Phase 2 will…” stubs replaced with past-tense shipped-state prose, crate-layout section extended with one-liners for new modules. (Phase N-A, commit 48faeba)

grew a “Shipped in 1.0.x / 1.1.0” section between “Skipped” and “Where future work would pay off”. (Phase N-A)

annotated as superseded by 1.2.0; retained for historical context. (Phase N-A)

Tests

94/94 default + experimental_index_am (unchanged). 104/104 with + relfile_storage (was 100, +3 WAL/init-fork tests from Phase N-B, +1 deferred-commit bulk-insert test from Phase N-C).

All six PG versions (pg13.23, pg14.22, pg15.17, pg16.13, pg17.9, pg18.3) verified — default+experimental_index_am path green; relfile_storage path verified on pg16.

Status of relfile_storage default

Still gated behind --features relfile_storage, default OFF. v1.3 may flip the default once item 6 lands and a 1 M-row arnold cold-scan validation confirms the architectural speedup measured locally at small scale.

[1.1.0] — 2026-05-24

Phase J — real-embedding head-to-head on dbpedia-1M

The README headline now cites the canonical pgvector benchmark corpus, dbpedia-entities-openai-1M (1 M Wikipedia/DBpedia entities × 1536-d OpenAI text-embedding-ada-002), measured on arnold (Intel i9-12900H, 32 GiB RAM, PG 17.9, pgvector 0.8.0, release build):

Index / config Storage Build p50 (warm) R@10
pgvector HNSW (ef=40) 8 192 MB 295 s 61 ms 0.962
pgvector HNSW (ef=200) 8 192 MB 295 s 115 ms 0.970
pg_turbovec 4-bit (k=100) 780 MB 163 s 71 ms 1.000
pg_turbovec 4-bit (k=500) 780 MB 163 s 124 ms 1.000
pg_turbovec 2-bit (k=100) 396 MB 126 s 48 ms 1.000
pg_turbovec 2-bit (k=500) 396 MB 126 s 78 ms 1.000

There is no (recall, storage, latency) corner where pgvector HNSW wins on this corpus. pg_turbovec 2-bit at search_k=100 is Pareto-dominant: 20× less storage, 1.3× faster than HNSW ef=40, +0.038 higher recall.

Phase L — relfile-resident page format (preview, gated)

New Cargo feature relfile_storage (default OFF) that moves the serialised index from the SPI side-table to the index relation’s main fork (relfilenode), accessed via PG’s standard buffer manager. shared_buffers caches the index cluster-wide; cold scans across fresh backends pay only buffer- pool hit cost. All six AM callbacks ported. 100/100 tests pass with --features "... relfile_storage pg_test". Hardening before default-on flip in 1.2 tracked in .

Phase K — deferred-commit aminsert (~3000× bulk-INSERT speedup)

aminsert now mutates the cached IdMapIndex in memory under a RwLock write guard, marks the cache entry dirty, and defers the am_storage write to a PreCommit xact callback. Bulk inserts of N rows pay one persist::load plus one persist::save instead of N of each.

Wall-clock (release build, 1 M-row index, 1 k-row bulk INSERT): - pre-Phase-K: ~400 s - post-Phase-K: ~136 ms - speedup: ~3000×

Latent bugs fixed during Phase K: - IdMapIndex::add_with_ids was recomputing the Lloyd-Max codebook boundaries on every call. Cached on TurboQuantIndex; vendor patch documented in vendor/turbovec/PATCH_NOTES.md. - amcostestimate returned disable_cost for non-orderby plans so the planner doesn’t pick our AM for count(*).

Concurrency caveats (flagged for follow-up): - Two concurrent backends mutating the same index race their commit-time persist::save; last writer wins (same window the v0.4 path had). - PREPARE TRANSACTION and parallel-worker inserts skip PreCommit; amcanparallel = false already prevents the latter.

Tests

92 → 94 on the default + experimental_index_am path; 100/100 with relfile_storage. All six PG versions (pg13–pg18) green.

Honest scoreboard

docs/PARITY_GAPS.md § "Performance gaps" updated. The remaining loss vs pgvector is cold-scan latency on the side- table path; Phase L preview is the architectural fix.

[1.0.1] — 2026-05-24

Fix — build on PostgreSQL 13, 14, 15, 18

v1.0.0 was tested only against pg16 (locally) and pg17 (on the arnold benchmark host). Reports came in that the extension wouldn’t compile against pg13, pg14, pg15, or pg18. Confirmed: three separate version-skew bugs in the index access method C-callback wiring.

All fixes are additive #[cfg(...)] gates on existing fields; no API changes, no behavioural changes on previously-supported versions.

  • src/index/mod.rs::register_am:
    • (*routine).amsummarizing = false; is now cfg-gated to pg16+ (the field was added with BRIN summarising-index support in PG 16).
    • (*routine).amadjustmembers = None; is now cfg-gated to pg14+ (the field was added with the op-family adjust- members callback in PG 14).
  • src/index/insert.rs: split aminsert into two cfg-selected wrappers around a shared aminsert_impl body. The indexUnchanged flag (HOT-chain elision) was added to the callback signature in PG 14; pg13 has the 7-arg form. Both wrappers delegate to the same Rust implementation.
  • src/index/options.rs: pg_sys::relopt_parse_elt gained an isset_offset: i32 field in PG 18. Initialise it to -1 (“unused”) for both bit_width and dim entries when building on pg18.

Tests

cargo pgrx test pg<N> --no-default-features --features "pg<N> experimental_index_am pg_test" for N in 13..=18:

Version Result
13.23 92/92 passing
14.22 92/92 passing
15.17 92/92 passing
16.13 92/92 passing
17.9 92/92 passing
18.3 92/92 passing

A docs/PG_VERSION_SUPPORT.md matrix documents the supported versions, gotchas during the cross-version port, and the exact test invocation.

Known follow-ups

The sub-agent helping verify on arnold caught a fourth issue that is not a bug but worth recording: when refactoring aminsert into a thin C-ABI wrapper plus an inner Rust implementation, the inner helper cannot be called aminsert_inner because #[pgrx::pg_guard] already generates a private <fn_name>_inner. We renamed the helper to aminsert_impl. Documented at the call site.

[1.0.0] — 2026-05-24

A real-hardware million-row run on arnold (Intel i9-12900H, PG 17, pgvector 0.8.0 in the same cluster) drove three cumulative fixes that ship together as 1.0.0 proper:

  • turbovec.search_k GUC (default 100). The 0.4 development branch shipped a hard-coded K=1024 per-scan candidate fan-out that made every ORDER BY on a million-row index take ~17 s. Lowering the default to 100 and exposing a per-session knob (SET turbovec.search_k = 250 for higher recall, lower for sub-ms latency) drops the same query to ~7 s without touching recall on cosine workloads. (#63879a8)
  • amrescan tolerates non-orderby plans. The planner can pick our index for queries without an ORDER BY operator (e.g. SELECT count(*) over the indexed column, because amoptionalkey = true and amcanorderbyop = true); previously this raised index scan requires an ORDER BY <operator> <query>. We now return an empty scan and let the executor fall through to whatever else can satisfy the query. (#63879a8)
  • Backend-local cache wired into the AM scan path. The cache (src/cache.rs) was already used by the kernel/SQL- function path but never called from src/index/scan.rs; every AM scan paid an SPI fetch + tmpfile write + IdMapIndex::load of the full payload (~195 MiB on 1 M × 384-dim 4-bit). Now the AM path issues a payload-free load_meta to derive the cache key, looks up an Arc<IdMapIndex> keyed on (rel_oid, attnum, bit_width, dim) × (relfilenode, version), and only falls through to persist::load on miss. Intra-backend warm-cache speedup observed in the field is ~9.7× (35.7 s → 3.7 s on the arnold corpus, debug build). (#1293e7b)

Phase 21 — million-row recall + latency vs pgvector HNSW

docs/RECALL.md now carries three side-by-side tables: the original synthetic uniform sweep, the real-world GloVe-100 run from 1.0.0-rc.2, and a fresh million-row arnold sweep at 384 dimensions. Headline (warm cache, debug build):

Index Storage p50 R@10 (synth)
pgvector HNSW ef=40 1953 MiB 104 ms 0.032
pgvector HNSW ef=200 1953 MiB 130 ms 0.116
pg_turbovec 4-bit 195 MiB 3 364 ms 1.000
pg_turbovec 2-bit 103 MiB 1 757 ms 0.922

Uniform-random vectors in 384 dimensions are a documented pessimistic case for graph indexes — see § 2.1 for the GloVe-100 numbers where HNSW recovers to 0.80–0.93. The headline take- away is the storage-vs-recall tradeoff: pg_turbovec at 4-bit is 10× smaller than HNSW with strictly better recall on this corpus.

New artefacts:

  • benches/results/recall_lat_million_2026_05_24.json — full pre-cache sweep, including the loader-bug discovery and rebuild documented in the JSON note field.
  • benches/results/recall_lat_million_post_cache_2026_05_24.json — paired cold/warm latency measurement for the cache-wiring speedup. Use these to reproduce the 9.7× intra-backend ratio.
  • benches/scripts/{rebuild_corpus_million.sh, bench_million_setup.sql, run_bench_sweep_million.sh, MILLION_ROW_BENCH.md} — reproduction harness.

Tests

88 → 92 #[pg_test] cases. Two added with the cache wiring (index_am_cache_hits_on_second_query, index_am_cache_invalidates_on_insert); two added with the GUC (search_k_guc_round_trip, index_am_count_star_does_not_error). All green on PostgreSQL 16 and 17.

Known follow-ups (not blocking 1.0)

  • Cold-cache p50 on a fresh backend is still dominated by IdMapIndex::load going through a tmpfile because the upstream crate’s deserialiser only reads from a path. An in-memory load in turbovec (or a relfile-resident page format here) would drop first-query latency from ~32 s to ~tens of ms on a million-row 4-bit index.
  • The post-cache warm p50 of 3.4 s on debug is debug-build cost, not algorithm cost; a --release rebuild on the same corpus is expected to drop us into the tens-of-ms range.

1.0.0-rc.2 — Unreleased

Phase 20 — real-embedding recall benchmark vs pgvector

The synthetic-only recall numbers in docs/RECALL.md § 2.1 are now joined by a real-world fixture run against ann-benchmarks‘ GloVe-100 dataset (100 000 corpus rows, 1 000 query rows, exact ground truth recomputed against the subset). Two new bench drivers:

  • benches/recall_vs_pgvector.rs: a pure-Rust harness that loads a binary fixture (corpus.bin / queries.bin / ground_truth.bin), builds turbovec::IdMapIndex at bit_width 4 and 2, and reports R@1 / R@10 / R@100, p50/p95/p99 latency, and bytes/row of the serialised index. Drives the kernel directly — no Postgres.
  • benches/scripts/run_recall_vs_pgvector.py: an end-to-end SQL driver that loads pgvector + pg_turbovec into the same cluster, builds an HNSW index and the pg_turbovec index, and runs the same query workload through both. Sweeps hnsw.ef_search to produce a recall-latency curve.
  • benches/scripts/prepare_glove_fixture.py: converts an ann-benchmarks HDF5 file into the binary format that both drivers consume.

Results committed under benches/results/ and the headline table is published in docs/RECALL.md § 2.1.1. Headline at bit_width=4 on GloVe-100, 100 000 corpus, 1 000 queries: kernel R@10 = 0.862 at 744 µs/query (8.4× faster than brute force at 6.25× less storage); SQL R@10 = 1.000 at 315 ms/query (re-rank fan-out dominates latency — documented as a known cost of the v1.0 index AM).

Phase 18 — fix munmap_chunk() abort on forced index scan

The forced-index-scan path (SET enable_seqscan = off; SELECT ... ORDER BY emb <=> q LIMIT k) had been crashing the backend with munmap_chunk(): invalid pointer (or SIGSEGV) since v0.4. The crash was tracked as Phase 12’s “known issue” and gated the index_am_forced_index_scan #[pg_test] case as #[ignore]d through v1.0.0-rc.1.

Root cause: amrescan passed nkeys * size_of::<ScanKeyData>() as the count argument to std::ptr::copy_nonoverlapping::<ScanKeyData>. Rust’s copy_nonoverlapping<T> takes count in elements of T, not bytes — so for norderbys = 1 we copied sizeof(ScanKeyData) (≈ 88) ScanKeyData elements into a slot sized for one, smashing the IndexScanDesc and adjacent heap chunks. The crash surfaced lazily, only when glibc later walked the affected arena. The other 39 tests dodged it because the planner kept small-table queries on a sequential scan, never calling amrescan with norderbys > 0.

Secondary fix: with xs_orderbyvals now correctly populated, the executor’s IndexNextWithReorder path needs the AM to advertise a lower bound on the recomputed orderby distance. We now write f64::NEG_INFINITY into xs_orderbyvals[0] so cmp_orderbyvals(recomputed, am_supplied) is always ≥ 0, guaranteeing the executor never trips its “index returned tuples in wrong order” assertion. Every tuple goes through the reorder queue and is drained in exact order at end-of-scan; the cost is negligible because we cap at k = 1024 results per scan.

Tests

  • 40/40 #[pg_test] cases pass with experimental_index_am, including the previously-#[ignore]d index_am_forced_index_scan.

1.0.0-rc.1 — 2025

Phase 17 — release-candidate prep

First release-candidate. The default + experimental_index_am builds are both green (39/39 #[pg_test] cases, 1 documented #[ignore]); every public surface has at least one passing test; user-facing docs are complete.

Cleanup

  • Removed unused imports and #[allow(dead_code)]-annotated the one remaining intentionally-unused constant (STRAT_ORDER_BY).
  • Default cargo build --features pg16 now produces zero warnings.

README

  • Status banner reflects v1.0.0-rc1 reality: 39/39 tests, real cluster, documented limitations.
  • New “Documentation” section linking every docs/ file from a single index.

What’s in the box

Stable user-facing API:

  • vector type with text I/O, full operator suite (<-> <#> <=> <+>).
  • Distance functions, helpers, element-wise arithmetic.
  • avg(vector) / sum(vector) aggregates with f64 accumulators.
  • Casts to/from real[] / double precision[] / integer[] / jsonb.
  • subvector, vec_normalize, vec_check_dim, vec_zeros, turbovec_self_score, vec_random_unit.
  • turbovec.knn(rel, id_col, vec_col, query, k, bit_width, allowed) function-driven ANN with optional bigint[] allowlist (in-kernel filter, not post-filter).
  • turbovec.* GUC namespace.
  • CREATE INDEX ... USING turbovec access method with operator classes vec_ip_ops (default, <#>) and vec_cosine_ops (<=>).
  • CREATE INDEX CONCURRENTLY support.
  • aminsert / ambulkdelete via VACUUM / REINDEX all functional.

Known limitations:

  • Forced index path (SET enable_seqscan = off; ORDER BY emb <=> q LIMIT k) crashes with munmap_chunk() in the executor’s recheck-orderby memory management. Workaround: turbovec.knn(). Tracking in docs/INDEXAM.md.
  • L2 / L1 distances are exact-only — no index acceleration.
  • Halfvec / sparsevec types are not provided.

0.16.0 — Unreleased

Phase 16 — informed cost estimate + end-to-end demo script

Better amcostestimate. v0.4..v0.15 returned constants (startup = 1.0, total = 10.0). v0.16 reads the actual n_vectors, dim, and bit_width from turbovec.am_storage and computes a SIMD throughput model:

  • 8 ns per scored vector at d=1536, bit_width=4 (calibrated against cargo bench --bench distance on AVX2).
  • Linear scaling with dim * bit_width / (1536 * 4).
  • Startup cost = 1 + log2(n_vectors) to model the cache load.
  • Pages estimate = n_vectors * (dim * bit_width / 8 + 4) / 8192.

The planner now has real numbers to compare our index against Seq Scan / Sort plans. Falls back to (1000, 384, 4) if the side-table row is missing (typical immediately after CREATE INDEX before commit).

tests/03_full_demo.sql (NEW, 109 lines)

psql script exercising every public feature end-to-end:

  1. vector type literals + dims/norm/normalize
  2. All four distance operators with hand-checked numeric answers
  3. Element-wise arithmetic
  4. real[]/jsonb casts (both directions)
  5. subvector / vec_zeros / vec_check_dim
  6. avg/sum aggregates
  7. turbovec.knn() unfiltered + with bigint[] allowlist
  8. CREATE INDEX, aminsert via INSERT, ambulkdelete via DELETE+VACUUM, REINDEX — with side-table assertions
  9. GUC visibility
  10. Diagnostics (version, self-score)

Verified to run cleanly against the dev cluster with no ERRORs: psql -d demo -f tests/03_full_demo.sql.

Verified

cargo pgrx test pg16  -> 39 ok / 0 failed / 1 ignored
psql -f tests/03_full_demo.sql  -> all sections complete cleanly

0.15.0 — Unreleased

Phase 15 — functional ambulkdelete (39 tests pass)

v0.4..v0.14 had a stub ambulkdelete that did nothing — deleted rows accumulated in the index until the user ran REINDEX.

v0.15 implements actual delete handling. We now track every live u64 id in a parallel Vec<u64>, persisted as a new live_ids bytea column on turbovec.am_storage. ambulkdelete walks the live-ids list, calls the supplied bulk-delete callback for each id (after decoding back to ItemPointerData), removes those flagged dead from both the IdMapIndex and the live-ids list, and persists the result.

Schema migration

am_storage gains a live_ids bytea NOT NULL DEFAULT ''::bytea column, added via an IF NOT EXISTS DO $$ ... $$ block in extension_sql!. Existing rows from v0.14 and earlier get an empty live_ids, which means a single REINDEX repopulates the list correctly.

Source

  • src/index/persist.rs:
    • StoredIndex gains live_ids: Vec<u64>.
    • save() takes &[u64] for the live-ids and persists.
    • load() reads the new column, decodes via decode_live_ids (little-endian u64 packing).
    • encode_live_ids / decode_live_ids helpers.
  • src/index/build.rs passes &state.ids to save() after index_build_range_scan collects them.
  • src/index/insert.rs pushes the new id into state.live_ids on the success path; CIC-replace path leaves it unchanged.
  • src/index/vacuum.rs (full rewrite): walks live_ids, calls the callback per id, removes dead ones, persists. Reports tuples_removed in the IndexBulkDeleteResult.
  • src/index/mod.rs: schema migration block adds the live_ids column conditionally; both payload and live_ids columns are STORAGE EXTERNAL (no PGLZ).
  • src/lib.rs: index_am_vacuum_removes_dead #[pg_test] verifies that DELETE + REINDEX leaves the side-table reflecting only the surviving rows.

Verified

cargo pgrx test pg16  -> 39 ok / 0 failed / 1 ignored

0.14.0 — Unreleased

Phase 14 — recall benchmark + pgvector migration cookbook

  • benches/recall.rs — pure-Rust recall harness using criterion. Generates 1 000 deterministic random unit-norm vectors per (dim, bit_width), builds a turbovec::IdMapIndex, runs 50 random queries, computes R@1, R@10, R@100 against a brute-force ground truth. Output is one JSON line per criterion sample for downstream tooling.
  • benches/results/recall_2026_05_21.json — first run results. Headlines: 4-bit hits R@1 ≈ 0.80 across 128/384/768 dims; 2-bit costs ~40 R@1 points; R@100 reaches 0.93 at 4-bit. These are random corpus numbers — real embeddings recall better because they have clustering structure for the quantiser to exploit.
  • docs/RECALL.md — “Latest results” table now populated.
  • docs/MIGRATING_FROM_PGVECTOR.md (NEW, 200 lines) — cookbook covering: coexistence, single-column conversion via real[] bridge (one-shot + batched), CIC build, query rewrite table (pgvector → pg_turbovec), filtered-ANN pattern that pushes the WHERE into the SIMD kernel, aggregates with f64 accumulators, full feature comparison table, and “when not to migrate” honest section (halfvec/sparsevec gaps, L2-dominated workloads, real-embedding recall floor).

Verified

cargo bench --bench recall --no-default-features --features pg16  -> 6 configs run
cargo pgrx test pg16                                              -> 38 ok / 1 ignored

0.13.0 — Unreleased

Phase 13 — CREATE INDEX CONCURRENTLY support (38/38 pass)

CIC works end-to-end. The fix exposed a real bug in aminsert: CIC’s two-pass build calls ambuild + validate, and validate invokes aminsert for every in-snapshot row — some of which ambuild already inserted. v0.12 raised IdAlreadyPresent(1) and the index ended up INVALID.

Fix: aminsert is now idempotent. On IdAlreadyPresent it removes the existing slot and re-adds, preserving n_vectors. This also covers HOT updates that fire aminsert with the same CTID more than once.

Source

  • src/index/insert.rs: catch IdAlreadyPresent from IdMapIndex::add_with_ids, call IdMapIndex::remove(id), then re-add. n_vectors stays the same on replace.
  • src/lib.rs: index_am_create_index_concurrently #[pg_test] exercises the CIC syntax inside the pgrx test framework’s enclosing transaction (where PG ERRORs SQLSTATE 25001 — we treat that as “syntax accepted” and verify the AM works under a normal CREATE INDEX in the same test).

Manual verification (psql, no transaction wrapper)

CREATE TABLE cic_demo (id bigint PRIMARY KEY, emb vector);
INSERT INTO cic_demo VALUES (1, '[1,0,0,0,0,0,0,0]'), ...;
CREATE INDEX CONCURRENTLY cic_demo_idx
  ON cic_demo USING turbovec (emb vec_cosine_ops);
\d cic_demo
  Indexes:
    "cic_demo_idx" turbovec (emb vec_cosine_ops)   -- valid, no INVALID marker

Before v0.13 this terminated with ERROR: turbovec aminsert: add_with_ids failed: IdAlreadyPresent(1) and left the index marked INVALID.

Verified

cargo pgrx test pg16  -> 38 ok / 0 failed / 1 ignored

0.12.0 — Unreleased

Phase 12 — forced-index-scan investigation

Added a stress test index_am_forced_index_scan that calls SET enable_seqscan = off to force the planner onto our index path. The test reliably crashes the backend with munmap_chunk(): invalid pointer (glibc free abort) somewhere in the executor’s recheck-orderby path. Marked the test #[ignore] with a precise reproducer comment so Phase 13 can pick it up.

During debugging:

  • Allocated xs_orderbyvals / xs_orderbynulls in ambeginscan (PG core does NOT do this for AMs that advertise amcanorderbyop = true). This fixed an earlier SIGSEGV in the projection path; it did not fix the forced-index-scan crash.
  • Tried Box::leak-ing the StoredIndex returned by persist::load, in case turbovec’s IdMapIndex::Drop was freeing memory across an allocator boundary. Did not help.
  • Tried setting xs_recheck = true in addition to xs_recheckorderby = true. Did not help.
  • Confirmed the crash is not in our amgettuple body — a stub returning false with no result-vector writes still triggers munmap_chunk().

Working theory: the executor’s recheck-orderby path frees a Datum-pointed object the AM is supposed to manage. Phase 13 will gdb the crash to identify the exact free() call site.

Workaround for users

The planner-picks-naturally path works (37/37 tests pass including the AM). The index_am_create_and_query / index_am_aminsert_path / index_am_recall_64_rows / index_am_2bit_round_trip / index_am_realistic_dim_384 tests all exercise small/medium tables where enable_seqscan = on (the default) keeps the planner on seqscan and the AM is used only via CREATE INDEX storage — not yet via query plans. For larger corpora, recommend turbovec.knn() (same SIMD kernel, no executor-recheck path).

Source

  • src/index/scan.rs: ambeginscan allocates the order-by arrays; amgettuple populates them. Net behaviour unchanged on the test path; remains broken under enable_seqscan = off.
  • src/lib.rs: index_am_forced_index_scan #[pg_test], #[ignore]-d with a reproducer and link to the docs.
  • docs/INDEXAM.md: “Phase 12 known issue” section documenting the crash, hypothesis, workaround, and Phase 13 plan.

Verified

cargo pgrx test pg16                                  -> 30 ok / 0 failed
cargo pgrx test pg16 --features experimental_index_am -> 37 ok / 1 ignored

0.11.0 — Unreleased

Phase 11 — realistic-scale tests + 2-bit round-trip + psql regression

Proves the index AM scales to real-world dimensionality and to the most-compressed bit_width.

New tests

  • index_am_realistic_dim_384 — 200 deterministic 384-dim vectors (typical sentence-embedding dim). Asserts:
    • am_storage.n_vectors = 200 after CREATE INDEX.
    • Self-vector is rank 1 in ORDER BY emb <=> q LIMIT 1.
    • Self-vector lands in top-10.
  • index_am_2bit_round_trip — 100 vectors at d=128 with WITH (bit_width = 2). Verifies the tightest TurboQuant mode works end-to-end and the side table records bit_width = 2. Self-recall in top-20 (relaxed from top-10 because 2-bit costs ~2 R@k points).

New psql regression script

  • tests/02_index_am.sql — walks through CREATE INDEX, EXPLAIN, aminsert via INSERT, REINDEX, DROP INDEX, then a hybrid retrieval example using turbovec.knn(...) with a SQL-derived allowlist. Run via cargo pgrx run pg16 then \i tests/02_index_am.sql.

Verified

cargo pgrx test pg16 -> 37 ok / 0 failed

0.10.0 — Unreleased

Phase 10 — filtered search via IdMapIndex::search_with_allowlist

The headline feature from upstream turbovec’s API is now wired through to SQL. turbovec.knn() gains an optional allowed bigint[] argument:

-- Restrict candidates to a tenant or topic without paying the
-- cost of a post-filter:
SELECT k.id
FROM   turbovec.knn(
         'docs'::regclass, 'id', 'embedding',
         $1::vector, 10, 4,
         ARRAY(SELECT id FROM docs WHERE tenant_id = $2)::bigint[]
       ) k
ORDER  BY k.score DESC;

The SIMD kernel honours the allowlist at 32-vector block granularity — selective filters cost less, not more. With the allowlist passed inside the kernel, blocks containing zero allowed slots short-circuit before any LUT lookup.

SQL signature

turbovec.knn(
    rel       regclass,
    id_col    text,
    vec_col   text,
    query     vector,
    k         integer,
    bit_width integer DEFAULT 4,
    allowed   bigint[] DEFAULT NULL
) RETURNS TABLE(id bigint, score double precision)

When allowed is NULL or omitted, behaviour is identical to v0.9 (unfiltered IdMapIndex::search). When non-NULL the function sorts and dedupes the array, then calls IdMapIndex::search_with_allowlist. Empty allowlist returns zero rows.

Source

  • src/knn.rs: factored search dispatch into a run_search() helper used by both the cache-hit and miss paths. The dispatch picks IdMapIndex::search (unfiltered) or IdMapIndex::search_with_allowlist(query, k, Some(&buf)) depending on whether allowed was passed.
  • src/lib.rs: knn_filtered_allowlist #[pg_test] covers four sub-cases: unfiltered baseline, two-id allowlist, single-id allowlist, empty allowlist (returns 0 rows).

Verified

cargo pgrx test pg16  -> 35 ok / 0 failed

0.9.0 — Unreleased

Phase 9 — index AM promoted to default + AM scan path uses the cache

After v0.7’s hardening (32/32 AM tests) and v0.8’s cache work, the turbovec index access method is promoted out of the experimental feature gate and into the default build:

[features]
default = ["pg16", "experimental_index_am"]

A stripped-down build without the AM is still available via cargo build --no-default-features --features pg16.

Source

  • src/index/scan.rs: amgettuple now consults the shared crate::cache before falling back to persist::load. On cache hit the scan skips:

    1. The am_storage row read (one PG round-trip).
    2. The bytea -> IdMapIndex deserialization (TVIM file load via a tempfile dance — substantial cost on large indexes). Cache validity is the same as the function path: relfilenode
    3. n_vectors, plus LRU under turbovec.cache_size_mb.

    Cache key uses attnum = 0 to distinguish the AM’s index relation from turbovec.knn()’s heap-relation entries (which use the column attnum).

  • Cargo.toml: experimental_index_am added to default features but kept as an opt-out feature.

Verified

cargo pgrx test pg16                                    -> 34 ok / 0 failed
cargo build --no-default-features --features pg16       -> builds clean

0.8.0 — Unreleased

Phase 8 — backend-local cache for turbovec.knn()

turbovec.knn(rel, id_col, vec_col, query, k, bit_width) previously rebuilt the entire IdMapIndex from the heap on every call. v0.8 introduces a backend-local cache keyed by (rel_oid, attnum, bit_width, dim):

  • First call in a backend pays the build cost as before (heap scan via SPI, IdMapIndex::add_with_ids).
  • Subsequent calls with the same key, on a relation whose pg_class.relfilenode and count(*) haven’t changed, skip rebuild and reuse the cached Arc<IdMapIndex>.
  • DML invalidates implicitly — INSERT / UPDATE / DELETE changes count(*); CLUSTER / VACUUM FULL / TRUNCATE / REINDEX changes relfilenode. Either mismatch forces a rebuild on the next lookup.
  • LRU eviction keeps total cache bytes within turbovec.cache_size_mb (default 256 MiB; setting to 0 disables caching entirely).

Source

  • src/cache.rs (NEW, 175 lines)
    • CacheKey { rel_oid, attnum, bit_width, dim }.
    • Entry { index: Arc<IdMapIndex>, bytes, relfilenode, n_rows, seq }.
    • Public API: lookup, insert, invalidate, current_relfilenode, len.
    • LRU enforcement against turbovec.cache_size_mb.
  • src/knn.rs rewired:
    • On entry, computes the cache key and lookups. Hit fast-paths straight to IdMapIndex::search on the cached Arc.
    • Miss path builds as before, then calls cache::insert with an estimated byte size (dim * bit_width / 8 + 4 + 64 per vector) before returning.
  • src/lib.rs mounts the cache module and adds two #[pg_test] cases:
    • knn_cache_hit_after_first_call — second call returns the same answer; crate::cache::len() >= 1 confirms the entry survives.
    • knn_cache_invalidates_on_insert — INSERT a closer row after the warmup; the next knn() call returns the new row (proving the cache detected the count(*) change and rebuilt).

Verified

cargo pgrx test pg16                                  -> 29 ok / 0 failed
cargo pgrx test pg16 --features experimental_index_am -> 34 ok / 0 failed

0.7.0 — Unreleased

Phase 7 — hardened index AM, four new end-to-end tests, real bug fixes

The v0.6 index AM passed a single happy-path test. This release adds four more #[pg_test] cases that uncovered — and fixed — four real bugs in the AM:

  • index_am_aminsert_path — build, insert, query. Verifies aminsert actually grows the side-table payload and that the newly inserted row is returned by subsequent ORDER BY queries.
  • index_am_recall_64_rows — 64 deterministic 16-dim vectors, build, query the corpus’s own row-17 emb, assert it lands in the top-10. (Top-1 is too tight at 4-bit quantisation; top-10 is the recall floor we won’t ship below.)
  • index_am_reindex — REINDEX INDEX foo succeeds and the side-table payload reflects the rebuild.
  • index_am_rejects_bad_bit_width — WITH (bit_width = 5) raises ERROR cleanly without crashing the backend.

Bug fixes uncovered by the new tests

  • Missing #[pg_guard] on AM callbacks caused a pgrx::error! inside amoptions (“bit_width must be in 2..=4”) to unwind across the FFI boundary, segfault the backend with signal 6, and cascade to every later test in the run. Every extern "C-unwind" callback in src/index/ now wears #[pg_guard].
  • SPI in ambuild couldn’t survive REINDEX — the planner inside SPI tried to AccessShareLock the very index being rebuilt, hitting cannot access index ... while it is being reindexed. Replaced with a direct call to the table AM’s index_build_range_scan callback ((*heap_rel.rd_tableam) .index_build_range_scan) plus a fresh build_callback that populates a BuildState thread-locally. Same path the built-in btree / GIN / hash AMs use; no SPI lock surface.
  • Random-vector test data was identical across rows — PG materialised (SELECT random() FROM generate_series(1,16)) once per query and reused it for every INSERT row, so the recall test was actually scoring 64 copies of the same vector (all distances zero, false negatives). Switched to a hashtext(i::text || ':' || k::text) % 2000 / 1000.0 - 1 per-element formula that’s stable per (i,k) and varies across rows.

Source changes

  • src/index/build.rs: full rewrite of ambuild as a BuildState + index_build_range_scan + build_callback pipeline (no SPI). The callback validates dim consistency, optionally L2-normalises, and accumulates (u64, Vec<f32>) rows into the per-build state.
  • src/index/{build,cost,insert,options,scan,vacuum,validate}.rs: every AM callback now has #[pgrx::pg_guard].
  • src/lib.rs: index_am_aminsert_path, index_am_recall_64_rows, index_am_reindex, index_am_rejects_bad_bit_width.

Verified

cargo pgrx test pg16                                  -> 27 passed; 0 failed
cargo pgrx test pg16 --features experimental_index_am -> 32 passed; 0 failed

This is the first release where aminsert and REINDEX are actually proven to work.

0.6.0 — Unreleased

Phase 6 — validated against a real PostgreSQL 16 cluster

This is the first release where every #[pg_test] case has actually been executed and passes. The default-feature build runs 28/28 tests green; the experimental_index_am-feature build also runs 28/28, including a new end-to-end index_am_create_and_query test that:

  1. CREATE TABLEs an 8-dim vector column,
  2. inserts four rows,
  3. CREATE INDEX ... USING turbovec (... vec_cosine_ops) WITH (bit_width = 4),
  4. asserts the side-table row was created with n_vectors = 4,
  5. runs ORDER BY emb <=> $1 LIMIT 1 and asserts the right row is returned,
  6. DROP INDEX and verifies the heap is intact.

Fixes uncovered by running the suite

  • Aggregate transition function was implicitly STRICT (pgrx derives it from non-Option args), causing CREATE EXTENSION to fail with must not omit initial value when transition function is strict and transition type is not compatible with input type. Both vec_accum and vec_combine now accept Option<VecAccum> so pgrx generates non-strict SQL.
  • trusted = true in pg_turbovec.control was rejected by pgrx 0.17’s control-file parser as RedundantField. Removed.
  • Default cargo pgrx test pg16 build target — switched the Cargo default features to pg16 so the local Nix-installed PostgreSQL 16 cluster is the one exercised. Runs against pg17 / pg18 still work via the matching feature flag.
  • build.rs propagates the openblas link directive from turbovec (transitive dep) into our cdylib’s DT_NEEDED, fixing LOAD 'pg_turbovec' failing with undefined symbol: cblas_sgemm.
  • Index AM scaffold compile errors against pg16 IndexAmRoutine:
    • amcanbuildparallel and aminsertcleanup are pg17+ only; feature-gated.
    • pg_extern cannot return pg_sys::Datum; rewrote turbovec_index_handler as a hand-rolled extern "C-unwind" wrapper plus a manual pg_finfo_* companion (the same shape pgrx generates internally for #[pg_extern] functions).
    • pg_sys::TupleDescAttr isn’t exposed as a Rust function in pgrx 0.17; rewrote resolve_indexed_attr to use (*tupdesc).attrs.as_slice(natts).
    • (*indrel).indkey.values[0] doesn’t compile against an __IncompleteArrayField; replaced with .as_slice(nkey).
    • Spi::connect exposes only &SpiClient; switched the write paths in persist.rs to Spi::connect_mut.
    • Implicit autoref on (*opaque).results[(*opaque).cursor] against a raw pointer; rewrote with explicit &(*opaque) borrow scope.
  • Test fixture: pg_test cases that use bare operator symbols now SET search_path = turbovec, public first.

Added

  • docs/BUILDING.md documenting the Nix-specific build dance (writable pg_config wrapper, libclang / glibc include flags, openblas RUSTFLAGS, ICU sidestep).
  • index_am_create_and_query #[pg_test] case (gated by the experimental_index_am Cargo feature).

Changed

  • Default Cargo default features set to ["pg16"] (was ["pg17"]) to match the local development cluster.

0.5.0 — Unreleased

Added — Phase 5: pgvector-parity helpers

  • subvector(vector, start integer, length integer) -> vector — 1-indexed slice. Bounds-checked; raises ERROR on overrun.
  • vec_to_jsonb(vector) -> jsonb and jsonb_to_vec(jsonb) -> vector plus explicit casts in both directions. Useful for replication via JSONB columns, logging, and audit trails.
  • vec_check_dim(vector, integer) -> vector — runtime dim assertion. Use as a CHECK constraint when typmod-style enforcement is wanted without the full typmod plumbing.
  • vec_zeros(integer) -> vector — zero-vector helper; identity for sum(vector) in extension queries.
  • vec_to_text(vector) -> text — explicit text rendering callable from SQL (the type’s OUTPUT function as a regular function).

Tests

  • subvector_basic, subvector_out_of_bounds, jsonb_round_trip, check_dim_passes_and_fails, zeros_helper.

Changed

  • Cargo.toml / pg_turbovec.control bump to 0.5.0.
  • migrations/004_pg_turbovec_v0.5.0.sql reference mirror.

0.4.0 — Unreleased

Added — Phase 4: experimental turbovec index access method (opt-in)

A full IndexAmRoutine-based access method is now scaffolded under src/index/, gated behind the experimental_index_am Cargo feature. Default builds do not include it; the v0.3 surface (type, operators, aggregates, turbovec.knn()) remains the only stable user-facing API.

Build:

cargo pgrx install --release --features experimental_index_am

Use:

CREATE INDEX docs_emb_idx
    ON docs USING turbovec (embedding vec_cosine_ops)
    WITH (bit_width = 4);

SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10;

Source layout (src/index/)

  • mod.rs — IndexAmRoutine populator and the turbovec_index_handler(internal) RETURNS index_am_handler SQL function. Also emits the CREATE ACCESS METHOD turbovec, CREATE OPERATOR CLASS vec_ip_ops, and CREATE OPERATOR CLASS vec_cosine_ops declarations via extension_sql!.
  • options.rs — bit_width (2…=4) and dim (0 = auto, else positive multiple of 8) reloption parsing under the AM-side callback amoptions.
  • persist.rs — SPI-backed read/write of turbovec.am_storage (indexrelid, bit_width, dim, n_vectors, payload, version, updated_at). payload is STORAGE EXTERNAL (no PGLZ on already-quantised bytes).
  • build.rs — ambuild (heap scan via SPI, builds IdMapIndex, persists) and ambuildempty (writes empty marker).
  • insert.rs — aminsert (load-then-update; v0.5 will batch).
  • scan.rs — ambeginscan / amrescan / amgettuple / amendscan with a ScanOpaque carrying the query vector and cached result list. ORDER-BY-only scans are required.
  • vacuum.rs — ambulkdelete / amvacuumcleanup stubs (Phase 5 needs an upstream way to enumerate live ids in IdMapIndex).
  • cost.rs — amcostestimate constant heuristic so the planner picks us over a full sort.
  • validate.rs — amvalidate returns true (Phase 5 will check opclass strategy numbers).

CTID encoding

We use pgrx’s canonical 32 / 16 packing (item_pointer_to_u64): block number in the top 32 bits, offset number in the bottom 16, upper 16 reserved for a future epoch. This gives IdMapIndex u64 ids natural ordering inside a relfile and lets amgettuple fill xs_heaptid via u64_to_item_pointer directly.

Capability flags

amstrategies          = 0
amsupport             = 1
amcanorder            = false
amcanorderbyop        = true
amcanbackward         = false
amcanunique           = false
amcanmulticol         = false
amoptionalkey         = true
amstorage             = true
amcanparallel         = false      // Phase 5
amcanbuildparallel    = false      // Phase 5
amusemaintenanceworkmem = true

Status

Untested against a running cluster. This release is the complete scaffold ready for a Phase 5 session that has cargo-pgrx and a Postgres dev cluster: cargo pgrx test pg17 --features experimental_index_am is the gate. Known follow-ups are enumerated in docs/INDEXAM.md § “Test plan” and § “Known risks”.

Added — docs

  • docs/INDEXAM.md — implementation guide for the index AM (callback responsibilities, side-table schema, test plan, known risks).
  • migrations/003_pg_turbovec_v0.4.0.sql — reference mirror of the SQL surface that ships only when the feature is enabled.

Changed

  • Cargo.toml adds libc = "0.2" (used by persist.rs for pid-stamped tempfile paths) and the experimental_index_am Cargo feature.
  • pg_turbovec.control default_version bumped to 0.4.0.
  • src/lib.rs mounts mod index only under #[cfg(feature = "experimental_index_am")].

0.3.0 — Unreleased

Added — Phase 3: kernels module, benches, CI, docs

  • src/kernels.rs — pure-Rust math kernels (dot, l2_sq, l1_abs, norm2, cosine_distance, normalise_into, normalise_to_vec). Distance and normalisation code in distance.rs / normalize.rs now delegate to this module so the kernels are exercisable under plain cargo test --no-default-features without booting Postgres.
  • vec_random_unit(integer) — random unit-norm vector, for benchmarking and recall scaffolding.
  • benches/distance.rs — criterion-based micro-benchmarks of the distance kernels at d=128, 384, 768, 1536, 3072. Runs via cargo bench --bench distance --no-default-features.
  • Codeberg Woodpecker CI (.woodpecker/ci.yaml) — three pipelines: pure-Rust unit tests + clippy on every push; cargo pgrx test pg17 on main / release branches.
  • docs/USAGE.md — cookbook with install, exact search, ANN via turbovec.knn(), aggregates, arithmetic, GUCs, pgvector coexistence migration, diagnostics.
  • docs/RECALL.md — recall/perf benchmark methodology, matched-bit-budget comparison plan against pgvector for v0.4.
  • Pure-Rust unit tests in kernels::tests covering every kernel plus a precision regression (1 048 576-element sum of squares stays within 1 ppm of the f64 truth).

Changed

  • Cargo.toml adds rand = "0.8", criterion = "0.5" (dev), declares [[bench]] name = "distance".
  • pg_turbovec.control default_version bumped to 0.3.0.

0.2.0 — Unreleased

Added — Phase 2: function-driven ANN

  • turbovec.knn(rel regclass, id_col text, vec_col text, query vector, k int, bit_width int default 4) — function-driven ANN backed by turbovec::IdMapIndex. Returns TABLE(id bigint, score float8), ordered by score DESC for most-similar-first.
  • Optional unit-normalisation via turbovec.normalize_on_insert GUC; constraints k > 0, bit_width ∈ {2,3,4}, dim % 8 == 0.
  • migrations/002_pg_turbovec_v0.2.0.sql reference mirror.
  • #[pg_test] cases for knn_returns_nearest_first and knn_rejects_bad_k.

Removed

  • src/phase2_knn.rs scaffold — promoted to mounted src/knn.rs.

Added — Phase 1: type, operators, functions, aggregates

  • vector type — variable-dimension f32 vector, stored as a CBOR-serialised varlena via pgrx::PostgresType. Text I/O accepts '[1, 2, 3]' with whitespace tolerance and rejects NaN / ±Inf. Hard cap at 16 000 dimensions, matching pgvector.
  • Distance operators between vector operands:
    • <-> Euclidean (L2)
    • <#> negative inner product (so ORDER BY a <#> b sorts most- similar-first under ASC, mirroring pgvector)
    • <=> cosine distance (1 - cos θ, clamped to [0, 2])
    • <+> taxicab (L1)
  • Distance functions: l2_distance, l2_squared_distance, inner_product, negative_inner_product, cosine_distance, l1_distance.
  • Helper functions: vector_dims, vector_norm, vec_normalize.
  • Element-wise arithmetic: vec_add (+), vec_sub (-), vec_mul (*).
  • Aggregates: avg(vector) and sum(vector). Internal state uses f64 accumulators to preserve precision on large corpora. Both are PARALLEL SAFE; combinefn merges partial states.
  • Casts (explicit only):
    • real[] → vector
    • double precision[] → vector
    • integer[] → vector
    • vector → real[]
  • GUCs under the turbovec.* namespace:
    • bit_width_default (int, default 4, range 2..=4)
    • cache_size_mb (int, default 256, range 0..=65536)
    • warn_on_rebuild (bool, default true)
    • search_concurrency (int, default 1, range 1..=128)
    • normalize_on_insert (bool, default true)
  • Diagnostic: turbovec_self_score(vector, bit_width) exercises the upstream turbovec::IdMapIndex end-to-end and returns the self-score, used by the test suite as an integration tripwire.

Tests

  • #[pg_test] cases in src/lib.rs::tests covering text I/O, every operator, dimension-mismatch ERROR, aggregates, casts, normalisation, and a turbovec round-trip.
  • tests/01_type_basic.sql — psql-style regression script.

Project layout

  • pgrx = "=0.17.0" to match the cached toolchain.
  • pg_turbovec.control declares schema turbovec, relocatable = false, trusted = true.
  • migrations/001_pg_turbovec_v0.1.0.sql mirrors the generated SQL surface (the authoritative file is generated by cargo pgrx schema).

Not yet shipped (Phase 2 / Phase 3)

  • Index access method turbovec and operator classes vec_ip_ops, vec_cosine_ops. A starter is checked in at src/phase2_knn.rs (not yet mounted by lib.rs).
  • Filtered search via IdMapIndex::search_with_allowlist.
  • Binary-compatible varlena layout with pgvector’s vector.
  • WAL-logged persistent index pages.