Contents
- pgvector parity gap tracker
- Performance gaps (the honest scoreboard)
- Types
- Distance operators
- Arithmetic & concatenation operators
- Index scan features
- Aggregates
- Functions
- Index AMs
- Operator classes
- Gaps found against zvec (2026-09-21 source review)
- Known ceiling: IVF builds are not memory-bounded — FIXED (Phase Z6)
- Phase plan
pgvector parity gap tracker
What pgvector offers (as of 0.8.x) and where pg_turbovec stands.
Performance gaps (the honest scoreboard)
This section enumerates known performance regressions of pg_turbovec vs pgvector. They are correctness-OK in every case; the trade-off is that the wins (10× less storage, exact recall) come paired with these losses.
2026-09-25 re-measurement — this SUPERSEDES the 2026-06-15 correction below. The 2026-06-15 “~2552 ms / we LOSE ~490×” line was itself flawed: the harness put the query vector in an ORDER BY subquery (
... <=> (SELECT emb FROM query_set WHERE qid=$N)), which adds ~90 ms of InitPlan/materialize overhead OUTSIDE the index scan on BOTH engines (measured: HNSW ef80 is 2.3 ms with a literal query vector vs 93.9 ms with the subquery, same plan), and it timed only the Index-Scan node. A corrected end-to-end benchmark (top-levelEXPLAIN(ANALYZE)Execution Time, literal query vectors, one warm session per arm; c7i.8xlarge AVX2, Cohere-wiki 1M × 1024-d, 100 held-out queries, exact top-10 GT;benches/results/rebench_20260925/) shows: flat-bw4 is 5.2 ms at R@10 = 1.000 and BEATS HNSW at R@10 ≥ 0.95 (HNSW 8.6 ms) and is 3× faster at ≥0.98 (HNSW 16.6 ms, and HNSW needs ef=400 to reach 0.98 while flat is already at 1.000). IVF-bw1 is within 2.4–2.8× of HNSW at 58× smaller storage; IVF-bw4 is the weakest arm (flat wins over IVF for bit_width ≥ 2 at this scale, as our own guidance says). So at 1M × 1024-d, warm, single-connection, all-in-RAM, pg_turbovec flat is competitive-to-better than HNSW on latency AND wins 14.6–58× on storage AND hits perfect recall. The remaining real gaps are COLD latency (flat cold ≈ 1.9 s — separate work, parallel repack) and QPS-under-load / >10M scale (unmeasured). The pre-AVX2 scalar-fallback caveat below still stands as history.2026-06-15 correction (SUPERSEDED 2026-09-25, kept for the record). An isolated, AVX2, contention-controlled benchmark on
arnold(Cohere-wiki 1M × 1024-d; seedocs/BENCHMARKS.md) overturned the earlier “we win warm p50” claim. pg_turbovec is a flat quantized full-scan index:O(n·dim)per query. At 1M rows its warm p50 was reported ~2.5 s (AVX2) vs pgvector HNSW’s ~5 ms — but that 2.5 s was inflated by the subquery measurement artifact described in the 2026-09-25 note above; the true end-to-end flat-bw4 number is ~5 ms. The old “26.8 ms on meh / we win 2.3×” numbers were produced by the pre-AVX2 scalar-fallback bug (fixed in v1.7.3) that returned fast-but-WRONG results, so they never represented correct behaviour. pg_turbovec’s wins are storage (14.6–58×), exact recall (1.000 vs HNSW’s ~0.96), build memory, AND — as of the 2026-09-25 re-measure — warm latency at 1M × 1024-d. The wrong choice remains COLD low-latency ANN before the per-backend cache warms; that is separate (parallel-repack) work.
| Metric (1 M × 384-d cosine, release build, arnold) | pgvector HNSW | pg_turbovec | Status |
|---|---|---|---|
| Storage | 1 953 MiB | 195 MiB (4-bit) | ✅ we win 10× |
| Build time | 8 m 13 s | 33 s | ✅ we win 15× (at 384-d; 1.9–2.1× at 1024-d) |
| Warm scan p50 (1 M × 384-d, GloVe) | 100 ms | 22 ms (v1.0.0) | ✅ we win 5× |
| Warm scan p50 (1 M × 1024-d, Cohere-wiki, AVX2, END-TO-END, 2026-09-25) | 4.3 ms @0.90 · 8.6 ms @0.95 · 16.6 ms @0.98 | flat-bw4: 5.2 ms @ R@10 1.000 (beats HNSW at ≥0.95, 3× faster at ≥0.98); IVF-bw1: 12–21 ms @0.90–0.96 | ✅ flat WINS at high recall; IVF-bw1 within 2.4–2.8×. Correctly measured: top-level Execution Time, literal query vectors (a subquery in the ORDER BY inflated the OLD 2552 ms figure by ~90 ms on both engines), one warm psql session per arm. Full iso-recall table + corrected harness in benches/results/rebench_20260925/. The old “~2552 ms / LOSE 490×” row was a measurement artifact and is RETRACTED. |
| Cold scan p50 (after backend restart) | ~100 ms | 566 ms (1 M × 1024-d 4-bit, v2.10.3, c7i.8xlarge AVX-512) | ⚠️ v2.10.3 cut cold-backend p50 3.1× (1766→566 ms) by parallelizing the cold-open repack. A follow-up profile (benches/results/finish_20260926/) then split the residual 566 ms: read_chain (buffer-manager copy of the 534 MB codes chain) is 280 ms / 99.3% of the read+repack cost, repack is now 2 ms / 0.7%. The repack (CPU) win is fully banked; the residual is buffer-manager I/O (~65k serial ReadBufferExtended+memcpy), which is NOT safely parallelizable from Rust threads inside a backend. Warm queries in the same backend are ~30 ms. A further cold win needs a prefetch / bulk-read change (L-effort, gated on its own measurement); mmap is off the table per BUFFER_CACHE_ONLY_DESIGN.md. Earlier history: 1 256 ms (1 M × 1536-d, post-Phase-P) → v1.7.3 deferred the id_to_slot HashMap off the read path → v2.10.3 parallelized repack. |
| INSERT throughput (per row, into a 1 M-row index) | ~0.5 ms (HNSW O(log n)) | 0.13 ms (post-Phase-K, deferred-commit on the relfile path) | ✅ we win 4× — v1.0.x had ~200 ms/row (full re-serialise per row) and we lost 400×; v1.1.0 (Phase K) shipped the deferred-commit pattern that mutates the cached Arc<RwLock<IdMapIndex>> per-row and persists once at xact commit, taking 1k-row bulk inserts from ~400 s to ~136 ms. v1.3.0 (Phase Q) extended the same pattern to the relfile path. |
| Recall on uniform-random | 0.03 | 1.000 | ✅ (but synthetic; real-world recall varies) |
| Recall on real OpenAI ada-002 (dbpedia-1M) | ~0.962 (ef_search=40) / ~0.970 (ef_search=200) | R@10 = 1.000 at default turbovec.search_k=100 |
✅ we win by 0.030–0.038. See docs/RECALL.md §2.2 for the full Phase J head-to-head; the 4-bit and 2-bit configurations both hit 1.000 because TurboQuant’s rotation + Lloyd-Max coding preserves rank order for real ada-002 embeddings (the rotation-then-reconstruct cycle is near-lossless on the workload). |
Cold-cache latency — the relfile-resident page format
v1.0.x..v1.1.0 stored the serialised index in a side-table
(turbovec.am_storage) read via SPI on first access. Every
fresh PostgreSQL backend paid the full SPI fetch + HashMap
construct cost (~6.8 s on 1 M × 384-d), then cached the result
in a per-backend Arc<IdMapIndex>. Connection pools that
create-and-destroy backends, or VACUUM workers, hit this every
time.
pgvector’s HNSW lives in the index relation’s main fork and is
cached in shared_buffers cluster-wide — first scan after a
restart is the same ~100 ms as the warm scan.
Status: shipped. The relfile-resident page format (Phase L, preview in v1.1.0) plus the persisted SIMD-blocked layout + Lloyd-Max codebook (Phase P, v1.2.0) close the cold-scan gap: dbpedia-1M cold p50 is 1 256 ms post-Phase-P, a 21× speedup over the v1.0.x side-table baseline. Phase Q (v1.3.0) removed the side-table path entirely; the relfile is the only storage strategy and the AM matches every other PostgreSQL index AM (btree, gist, gin, hnsw, ivfflat).
The remaining gap to pgvector HNSW (~1.2 s vs ~100 ms) is
bounded by the cost of reading the codes + scales + ids +
blocked-layout chains off disk into the per-backend index
plus, until v1.7.3, the O(n) id_to_slot HashMap build.
v1.7.3 (parity gap #3): lazy id_to_slot on the read path.
Profiling the cache-fill (200 k × 256-d, debug) showed the
dominant residual term was the id_to_slot:
HashMap<u64,usize> that IdMapIndex::from_id_map_parts*
eagerly materialises in finalise_from_inner — ~50 ms at
200 k rows, scaling linearly with n, dwarfing the
read_full (~16-22 ms) and read_blocked+read_rotation
(~12-18 ms) data copies. The scan path never reads
id_to_slot: search(q, k) with allowlist = None only ever
indexes slot_to_id[slot] (a Vec). So the AM scan path now
installs a cache::ReadOnlyIndex (the inner positional
TurboQuantIndex + the slot_to_id Vec, no HashMap), and
the HashMap build is deferred to the first aminsert /
remove, which rebuild a full IdMapIndex via am_install.
A read-only / pooled-connection backend that only ever scans
never pays the HashMap build. With the fix the read-only
constructor drops from ~50 ms to ~0 ms in the profiled debug
build. Wire format unchanged (scan-side only).
Deferred follow-ups (not in v1.7.3):
- Read-path mmap of codes / scales / ids. Today only the
static regions (blocked codes + rotation) are mmap’d; the
codes/scales/ids chains still go through
read_full(the buffer manager) becauseambulkdeleteswap-removes them in place. On a read-only cold scan they could be mmap’d RO too (same MVCC backstop as the static regions: heap visibility +xs_recheckorderby). Codes are the bulk of the index (≈ 768 MiB at 1 M × 1536-d × 4-bit), so this removes the largest remaining data copy on the real cold path. TheReadOnlyIndex::from_prepared_parts_borrowedconstructor already acceptsCow::Borrowed, so the wiring is additive once the relfile path resolution + per-page header-gap handling is extended to those chains. - Zero-copy mmap (wire-format change). Each chain page
carries a 24-byte PG
PageHeaderDataprefix, so the chain bytes are not contiguous in the mmap and must be copied off once at cache-fill. A header-gap-free on-disk layout would let the SIMD kernel read straight from the mmap with no copy at all. That is aMetaPageData::version3 → 4 wire bump and belongs in a v1.8 / v2.0 minor, not a scan-side patch. - Cross-backend shared cache. Cluster-wide caching of the index parts in a PG DSA/DSM segment keyed by relfilenode so the second backend onward maps an already-built structure instead of rebuilding. Biggest win for pooled workloads but the most invasive (DSA allocator, REINDEX invalidation, concurrency); XL effort, tracked as a follow-up.
INSERT throughput — the deferred-commit pattern
v1.0 aminsert did a full SPI fetch + full re-serialise per
row, costing ~2× 195 MiB of TOAST I/O per inserted row on a
1 M-row index. A bulk INSERT ... SELECT of 1 M rows would
have taken ~55 hours.
Status: shipped. Phase K (v1.1.0) introduced the deferred-
commit pattern: aminsert mutates the cached
Arc<RwLock<IdMapIndex>> in place, marks the entry dirty,
and registers a PreCommit xact callback that persists once
at the end of the transaction. Phase N-C (v1.2.0) extended
this to the relfile path. A 1 k-row bulk INSERT on a
turbovec-indexed table now finishes well under 5 s on debug
builds (was ~400 s pre-Phase-K).
For large INSERT ... SELECT we still pay one full relfile
rewrite at commit time, which is O(n_vectors) — confirmed by
measurement (benches/results/parity_20260925/item4_bulk_insert_design.md):
the PreCommit flush re-reads and rewrites all three chains
(codes/scales/ids) from offset 0, so the cost is per-commit, not
per-touched-row. Bulk-build at ROWS-per-COMMIT scale is
order-of-magnitude better than the pre-Phase-K hot loop, but
pgvector’s HNSW remains O(log n) per insert. Tracked as future work.
Guidance, two cases:
- One-shot bulk load: load into the heap first, then
CREATE INDEX — never the other way around.
- Continuous high-ingest into an already-large index (can’t stop
to CREATE INDEX): the rewrite is per-commit, so fewer, larger
transactions amortize it directly (batch many rows per commit). For
an IVF index, turbovec.ivf_max_delta_pct bounds how far the
appended tail grows before the index degrades to a flat scan;
REINDEX on a cadence restores cell pruning. A true incremental /
amortized-consolidation write path is L/XL persist-path work under
the corruption HARD MANDATE and is deferred pending demand.
Recall tuning
Two knobs together form the recall-vs-latency frontier:
turbovec.search_k(default 100) — how many candidates the kernel returns.turbovec.oversample(default 1.0, v1.8.x+) — the candidate-set widener. The scan fetchesceil(search_k * oversample)candidates ranked by the lossy quantized distance, and the always-on reorder queue (xs_recheckorderby) re-ranks them by exact full-precision distance, trimming to the true top-k under the LIMIT. This recovers true neighbours the quantized ranking placed just outsidesearch_k, turning quantization from a fixed accuracy point into a tunable frontier (Qdrantoversampling/ VectorChord rerank).
Measured (4-bit, 3000×64, search_k=10, 8 query seeds,
benches/results/oversample_recall_curve_2026_06_15.json):
| oversample | recall@10 | p50 (ms) |
|---|---|---|
| 1.0 | 0.8125 | 3.81 |
| 1.5 | 0.9625 | 3.86 |
| 2.0 | 0.9875 | 3.94 |
| 4.0 | 1.0000 | 4.06 |
| 8.0 | 1.0000 | 4.70 |
Recall climbs monotonically to 1.0 as oversample grows; latency
rises roughly linearly with the candidate count. There is no separate
turbovec.rescore GUC: oversampling plus the reorder queue together
are the rescore mechanism (the reorder queue already re-ranks every
returned tuple by exact distance, so an AM-side rescore would be
redundant). oversample composes with iterative scan — it sets the
initial k, iterative refill grows it from there.
On the 384-d synthetic corpus, K=100 gave R@10 = 1.000 because the uniform distribution makes ~all candidates within rounding of each other. On real-world embedding distributions (1536-d ada-002, GloVe-100), recall depends on K:
- Low K (50-100): low latency (10s of ms), recall ~0.85-0.92.
- High K (500-2000): higher latency (50-100s of ms), recall approaches 1.0.
Phase M (post-Phase J) will pick a default that hits ~0.95 on dbpedia-1M without breaking the warm-p50 latency story.
Types
| pgvector type | pg_turbovec status |
|---|---|
vector (FP32) |
✓ - turbovec.vector |
halfvec (FP16) |
✓ - turbovec.halfvec |
sparsevec |
✓ - turbovec.sparsevec |
bit (binary) |
✓ - turbovec.bitvec (named differently to avoid colliding with PG core’s built-in bit) |
Distance operators
| Op | pgvector | pg_turbovec |
|---|---|---|
<-> L2 |
✓ | ✓ (vector, halfvec, sparsevec; exact only on AM) |
<#> neg-IP |
✓ | ✓ (indexed for vector) |
<=> cosine |
✓ | ✓ (indexed for vector) |
<+> L1 |
✓ | ✓ (vector, halfvec, sparsevec; exact only on AM) |
<~> Hamming (binary) |
✓ | ✓ (bitvec) |
<%> Jaccard (binary) |
✓ | ✓ (bitvec) |
Arithmetic & concatenation operators
Element-wise add/subtract, the Hadamard (element-wise) product, and
concatenation. pgvector errors on a non-finite result coordinate
(value out of range: overflow); pg_turbovec matches this — +/-/*
require equal dimensions and raise on a non-finite result, and ||
errors if the combined dimension exceeds MAX_DIM (16 000). pgvector
does not define arithmetic for sparsevec, so neither do we.
| Op | pgvector | pg_turbovec |
|---|---|---|
+ element-wise (vector) |
✓ | ✓ |
- element-wise (vector) |
✓ | ✓ |
* Hadamard (vector) |
✓ | ✓ |
|| concat (vector) |
✓ | ✓ |
+ element-wise (halfvec) |
✓ | ✓ |
- element-wise (halfvec) |
✓ | ✓ |
* Hadamard (halfvec) |
✓ | ✓ |
|| concat (halfvec) |
✓ | ✓ |
| arithmetic (sparsevec) | ✗ (not offered) | ✗ (parity: not offered) |
Index scan features
| Feature | pgvector 0.8.2 | pg_turbovec status |
|---|---|---|
| ANN index scan | ✓ (HNSW, IVFFlat) | ✓ (turbovec AM) |
| Iterative / streaming scan | ✓ hnsw.iterative_scan, ivfflat.iterative_scan, max_scan_tuples, scan_mem_multiplier, max_probes |
✓ (v1.8.0; default flipped v1.20.1) — turbovec.iterative_scan (off | relaxed_order, default off — see the v1.20.1 perf-fix note in docs/UPGRADING.md: under the old relaxed_order default, PG’s reorder queue can never pop early because we advertise NEG_INFINITY, so every default-config query paid the AM’s full refill schedule regardless of LIMIT, a measured 450x tax). Opt into relaxed_order for a selective WHERE filter ORDER BY emb <=> q LIMIT k: amgettuple re-runs the turbovec search with a doubled k and feeds the new (deduplicated) candidates, capped by turbovec.max_scan_tuples (default 20000, matches pgvector). Ordering across refill batches is restored by the existing xs_recheckorderby reorder queue. pgvector’s strict_order is future work (our reorder queue already delivers exact ordering on top of relaxed_order). |
Bitmap index scan (amgetbitmap) |
✓ | ✗ (not applicable to ANN ordering) |
| Metadata filtering | post-filter + iterative + partial idx | three patterns — partial index (native PG pushdown), in-kernel allowlist via turbovec.knn(..., allowed) (flat) and the turbovec.allowlist session GUC on the ORDER BY operator path (flat and IVF: cell-scope ∧ allowlist; selective filters get cheaper), iterative scan + recheck. Remaining gap: no true in-traversal pushdown of an arbitrary live WHERE predicate on the AM path (the index stores only vector codes + TID, no payload columns). Full guide + measured crossover: docs/FILTERING.md. |
| Parallel index build | ✓ (maintenance workers) | ✓ (v1.8.0) — turbovec.build_parallelism drives a rayon pool over the quantize/pack stage; relfiles are byte-identical to a serial build. |
| Quantization tuning | manual re-rank CTE | turbovec.search_k (candidate count) plus turbovec.oversample (v1.8.x+): fetch ceil(search_k * oversample) quantized candidates, the always-on reorder queue re-ranks by exact distance — oversampling + reorder queue are the rescore mechanism, matching Qdrant oversampling / VectorChord rerank. Recall@10 climbs to 1.0 as oversample grows (see § Recall tuning). |
CREATE INDEX CONCURRENTLY |
✓ | ✓ (standard AM path) |
Build progress (pg_stat_progress_create_index) |
✓ phased | partial (no custom phase labels) |
Aggregates
| Aggregate | pgvector | pg_turbovec |
|---|---|---|
avg(vector) |
✓ | ✓ |
sum(vector) |
✓ | ✓ |
avg(halfvec) |
✓ | ✓ |
sum(halfvec) |
✓ | ✓ |
sum(sparsevec) |
✓ | ✓ |
Functions
| Function | pgvector | pg_turbovec |
|---|---|---|
l2_distance |
✓ | ✓ |
inner_product |
✓ | ✓ |
cosine_distance |
✓ | ✓ |
l1_distance |
✓ | ✓ |
vector_dims(vector) |
✓ | ✓ |
vector_dims(halfvec) |
✓ | ✓ |
vector_dims(sparsevec) |
✓ | ✓ |
vector_norm(vector) |
✓ | ✓ |
vector_norm(halfvec) |
✓ | ✓ |
subvector |
✓ | ✓ |
to_vector(text) |
✓ | ✓ (also to_vec) |
to_vector(text, integer, boolean) |
✓ | ✓ |
array_to_vector(real[]) |
✓ | ✓ (cast + array_to_vec) |
array_to_vector(real[], integer, boolean) |
✓ | ✓ |
vector_to_float4(vector, integer, boolean) |
✓ | ✓ |
binary_quantize(vector) |
✓ | ✓ |
hamming_distance(bitvec, bitvec) |
✓ | ✓ |
jaccard_distance(bitvec, bitvec) |
✓ | ✓ |
l2_normalize(vector) |
✓ | ✓ (also vec_normalize) |
vector_concat(vector, vector) |
✓ | ✓ (also || operator) |
halfvec_concat(halfvec, halfvec) |
✓ | ✓ (also || operator) |
max_sim / max_sim_cosine (ColBERT MaxSim) |
✗ | ✓ — SQL re-rank over vector[]; see HYBRID_SEARCH.md |
rrf_score (reciprocal rank fusion) |
✗ | ✓ — 1/(k+rank) hybrid-fusion helper; see HYBRID_SEARCH.md |
turbovec_check(regclass) (index integrity) |
✗ | ✓ — read-only, ownership-checked; reports wire version, kind, n_vectors vs slot count, duplicate-id / is_corrupt health, tombstone density (v1.28.4), plus a reason string and full graph-adjacency (CSR) structural validation for kind = graph. See PRODUCTION.md § Monitoring |
index_is_degraded(regclass) (IVF fallback) |
✗ | ✓ — reports whether an IVF index degraded to a flat O(n) scan |
Index AMs
| AM | pgvector | pg_turbovec |
|---|---|---|
ivfflat |
✓ (Lloyd k-means) | ✓ - WITH (lists = N), TurboQuant-quantized, byte-deterministic, out-of-core |
hnsw |
✓ | WITH (graph = true) exists but is DEPRECATED (v2.5.0) and scheduled for removal. A real Vamana build + beam scan with verified recall, VACUUM-able and insert-able (v1.24.0), parallel-built (v1.26.0), with a tunable beam (turbovec.graph_ef, v2.2.0) — but the reason it was built never materialised. Measured at matched recall it loses on every user-visible axis: SIFT-1M/128d @R@10≥0.95 costs 26.2 ms/qps@8 299 versus flat 0.98 ms/1380 and IVF 1.8 ms/2039; at GIST-1M/960d ≥0.95 and GIST-10M/960d ≥0.98 the target is unreachable at any graph setting (10M ceiling 0.873 at 181 ms) while IVF hits 0.983 at 28.4 ms. Its apparent sublinearity holds only at iso-beam (p50 1.11× for a 10× corpus, but recall falls 0.605→0.472); at iso-recall the curves diverge, never cross. Also 57–90× slower to build, larger on disk, and no out-of-core path. Use flat below ~1M and WITH (lists = N) (IVF) at scale — it is IVF, not the graph, that beats flat’s O(n) wall. Note the “60× parallel build speedup” was an artefact: graph_build_partitions_decide coupled shard count to thread count, and shards cost recall (GIST-1M R@10 0.920 at P=4 → 0.605 at P=83); threads at recall-preserving P buy <5×. |
turbovec |
n/a | ✓ - TurboQuant flat (the default, lists=0) |
Operator classes
| Opclass family | pgvector | pg_turbovec |
|---|---|---|
vector_l2_ops (ivfflat + hnsw) |
✓ | ✓ - vec_l2_ops (uses recheck-orderby; quality matches cosine for unit-norm vectors) |
vector_ip_ops |
✓ | ✓ (vec_ip_ops, default) |
vector_cosine_ops |
✓ | ✓ (vec_cosine_ops) |
vector_l1_ops (hnsw) |
✓ | ✓ - vec_l1_ops (recheck-orderby; candidate-set quality is approximate, recheck makes final order exact) |
halfvec_*_ops |
✓ | ✓ via expression index: CREATE INDEX ... USING turbovec ((emb::vector) vec_cosine_ops) |
sparsevec_*_ops |
✓ | ✓ via expression index, same pattern (note: dense-cast cost on each row may dominate for very high-dim sparse) |
bit_hamming_ops |
✓ | ✗ - TurboQuant kernel doesn’t fit Hamming-space ANN; use the exact <~> operator (no index) |
bit_jaccard_ops |
✓ | ✗ - same |
Gaps found against zvec (2026-09-21 source review)
Read-only comparison of the local alibaba/zvec checkout (d88357b, v0.7.0)
against pg_turbovec deac2d9. This was a source review, not a benchmark — no
performance claim below is measured. zvec is an in-process embedded vector DB,
so several of its “features” are things PostgreSQL already supplies us; those are
listed separately so they do not get mistaken for work.
Real gaps, in priority order
1. IVF degradation on insert is not reportable — FIXED (Phase Z1).
An aminsert into an IVF index cannot place the row in its cell without an O(n)
reshuffle, so both paths append and fall back to a flat scan. The BQ path always
preserved lists and stamped ivf_degraded; the TurboQuant path blanked
lists, so index_was_ivf() went false and the degradation was silent —
reconcile_and_write_flush planned its meta via plan_with_blocked, which
hardcodes lists: 0. Now it captures the on-disk lists under the held rewrite
lock and stamps both fields, leaving the coarse/cell-dir offsets at zero so the
scan takes the flat fallback deterministically. Tests:
ivf_flush_degradation_is_reportable,
ivf_soft_assign_index_rejects_insert_and_is_not_corrupt (429 passed, all legs).
Two things this turned up, worth knowing before touching the insert path:
- The two insert paths differ in when they write. BQ writes synchronously
inside
aminsert; TurboQuant only marks the cache dirty and defers to thePreCommitxact callback. A#[pg_test]always rolls back before PreCommit, so a plainINSERTin a test never exercises the TurboQuant flush — drive it throughxact::flush_to_relfile_for_test. - An
assign_dups > 1index is effectively READ-ONLY, and this is not a Z1 regression. Soft assignment repeats an external id across cells on purpose, soslot_to_idis not a bijection; the insert path loads the index into a flatIdMapIndex::from_id_map_parts, which requiresid_to_slot.len() == slot_to_id.len(), and fails before any of our code runs. Our ownlists-gated dup check correctly skips IVF — turbovec’s internal requirement is the blocker. Separate bug worth fixing: that rejection reportscorrupt relfile pages: duplicate idsabout a perfectly healthy index (turbovec_checkverifies it clean).
2. No sparse ANN opclass. We ship sparsevec with distance operators and
casts (src/sparsevec_ops.rs), but the AM registers opclasses only over
vector and vector[] (src/index/mod.rs:203–250) — so indexed learned-sparse
retrieval (SPLADE and similar) requires densifying first, which is exactly what
sparse representations exist to avoid. zvec has native sparse FLAT and HNSW,
inner-product only (Z/src/core/interface/index.cc:1279–1323). Note PostgreSQL
GIN full-text is not a substitute: it does lexical matching, not weighted
sparse nearest-neighbour.
3. —
RESCOPED, largely a non-gap (see Phase Z3 below). Investigating this for
implementation showed my original framing was wrong on two counts.WHERE predicates do not automatically become ANN masks
A qual on a non-indexed column cannot reach the AM at all. PostgreSQL defines
a scan key as index_key operator constant where “the index key is one of the
columns of the index” (Index Scanning); a qual on any other column
becomes a Filter on the scan node, evaluated by the executor. So “push the
WHERE into the kernel” is not a thing an AM can unilaterally do — the
information never arrives. zvec can do it because it owns its planner and
storage; a PostgreSQL AM does not.
amgetbitmap is not the route. It returns an unordered TIDBitmap, and an
ANN scan’s entire value is ordering. Serving ORDER BY <-> LIMIT k from a
bitmap would force a sort over the whole candidate set — scoring everything,
which is what ANN exists to avoid. Adding amgetbitmap would buy unordered
retrieval we have no use for.
And the useful behaviour already shipped in v1.8.0. Iterative scan handles
exactly this case, demand-driven: when the executor’s post-filter drains a
batch, amgettuple re-runs the search with doubled k (widening probes for
IVF), deduplicates, and restores ordering through the xs_recheckorderby
reorder queue — capped by turbovec.max_scan_tuples. The AM never needs to see
the filter, which is why this design works at all.
What genuinely remains is smaller and is not a mask-pushdown feature: the
manual allowlist path is a measured 2.6–14.7× win below ~7 % selectivity and
a 2.6× LOSS at 100 % (docs/FILTERING.md § 3), and nothing automatically
decides which side of that crossover a query is on. Z4 supplies the missing
input (real selectivity); see Phase Z3.
4. Cost estimation ignores everything that matters for filtered ANN —
FIXED (Phase Z4). amcostestimate used corpus size, dim and bit width and
reported index_selectivity = 0.0 unconditionally. Three defects, all fixed:
IVF is now costed for the cells it actually probes (probes / lists, with a
degraded index costed as flat since it takes the flat fallback); selectivity
derives from the planner’s own rel->rows / rel->tuples, so we agree with it by
construction rather than second-guessing with our own
clauselist_selectivity; and a pre-existing unit error — seconds divided by
cpu_operator_cost — had made a 1M × 1024-d flat scan cost ~23 against
PostgreSQL’s ~73,000 for the equivalent seq scan, i.e. ~3000× too cheap, which
let an ANN path beat plans that are genuinely faster. The ns throughput model
itself validated against our own published measurement (model 5.3 ms vs
measured 6.08 ms), so only the unit was wrong.
The arithmetic lives in two pure functions (scored_vectors, scan_cpu_cost)
with unit tests, because it is not observable through EXPLAIN: the
index-scan node’s cost also carries PostgreSQL’s heap-fetch and qual costs
(~1482 on a 20k-row fixture), which swamp the ~0.5 the AM contributes.
Worth noting zvec’s equivalent is a heuristic match-ratio threshold
(Z/src/db/sqlengine/planner/optimizer.cc:32–94), not a superior general
optimiser.
The lesson worth stealing: bounded mutable delta + explicit consolidation
zvec’s answer to “writes destroy trained structure” is not incremental
insertion into a trained index — both its IVF implementations reject additions
after training (Z/src/core/interface/indexes/ivf_index.cc:152–155,
ivf_rabitq_index.cc:152–161). Instead it separates concerns by lifecycle:
writes land in a Flat-backed mutable segment
(Z/src/db/index/segment/segment.cc:4201–4259), the segment is sealed and
rolled over (Z/src/db/collection.cc:1664–1748), and indexes are built and
merged by an explicit optimize() outside the exclusive locks, publishing
atomically at the end (Z/src/db/collection.cc:913–1010).
That maps onto our problem: keep the trained IVF cells intact and search a bounded append delta alongside them, with an explicit consolidation step, rather than flattening the whole index on first insert. Costs to weigh before committing to it: query fan-out across cells+delta, deletion/tombstone interaction, MVCC and crash-safety (our chains-then-meta invariant), and on-disk-format compatibility. Sealing alone builds no ANN structure — the build still has to happen somewhere.
Explicitly NOT gaps — PostgreSQL or we already cover these
- Hybrid fusion. zvec has native C++
MultiQuerywith RRF/weighted/callback fusion (Z/src/db/collection.cc:1877–1962). We do this in SQL withturbovec.rrf_scoreplus PostgreSQL’s own joins/aggregation (docs/HYBRID_SEARCH.md). The gap is packaged ergonomics, not capability — and zvec fuses independently truncated candidate pools, which is not obviously better. - Token-level multivector. We already have persistent ColBERT indexing with
batched token retrieval and exact MaxSim re-rank (
src/colbert.rs). zvec’sMultiQueryis rank fusion and is not evidence of MaxSim support. Our real limitation is narrower:colbert_searchtakes no filter argument and the opclass has no ORDER BY operator. - Durability / WAL. Ours is PostgreSQL’s, which is a stronger contract than
zvec’s (its default WAL flush threshold is 0 and append-time flushing is
conditional —
Z/src/db/index/storage/wal/wal_file.h:27). Not comparable as a drop-in. - Scalar filtering, partial indexes, full-text. PostgreSQL’s, natively.
- WAL amplification. Already addressed: we skip unchanged full pages
(v2.3.0) and pad chain allocations (v2.4.0). The residual cost is page
visits, not pages logged — measure with the existing
PAGES_WAL_LOGGEDcounter before assuming otherwise.
Deliberately deferred
- RaBitQ / IVF-RaBitQ, PQ-INT8. Alternative quantizers are only worth it
with training cost, raw-vector retention for refinement, and rebuild cost all
counted. We already have TurboQuant 2/¾-bit plus centered sign-BQ. Note
zvec gates RaBitQ to Linux x86_64
(
Z/src/db/index/segment/segment_helper.cc:879–882), and its PQ is a DiskANN implementation detail, not a public IVF-PQ option. - DiskANN. Its build still copies the whole corpus and allocates graph
storage in memory; the internal memory limit bounds PQ chunk count, not build
RSS (
Z/src/core/algorithm/diskann/diskann_builder.cc:260–289,734–775). We already have a spill-backed out-of-core IVF build with bounded chunks (src/index/build.rs:862–875) — keep it rather than importing an in-memory graph build. Our deprecated Vamana kind is not a DiskANN equivalent and should not be revived on the strength of this.
Known ceiling: IVF builds are not memory-bounded — FIXED (Phase Z6)
Root cause: the build callback had no memory-context management. Vector
is a PostgresType stored as CBOR, so FromDatum::from_datum palloc’s a
decoded buffer per row. With no switch, no reset and no pfree, every decoded
row accumulated in the long-lived ambuild context for the whole scan — so
the CorpusSpill was writing the corpus to disk while PostgreSQL held a
decoded copy of all of it in RAM. The arithmetic matched the measurement:
2M × ~5125 B of CBOR varlena = 9.55 GiB against 10.69 GiB measured as
unexplained.
build_callback is now a wrapper that switches into a per-tuple context and
resets it after every row. CorpusSpill::new_in(cxt, dim) plus an explicit
BuildState.build_cxt keep the lazily-opened BufFile in the long-lived
context — without that, resetting would have left a dangling BufFile on row
two.
Measured (2M × 1024-d, lists = 1414,
benches/results/z6_buildmem_20260922/):
| before | after | |
|---|---|---|
| at drain entry | 12.07 GiB | 2.39 GiB (5.05× less) |
| whole-build peak | 12.16 GiB | 3.45 GiB (3.5× less) |
| wall time | 1055 s | 889 s (16 % faster) |
Index bytes unchanged (CI 444/0, including the IVF byte-identity guards).
MEASURED at 10M × 1024-d (benches/results/z6_10m_20260924/), on the same
c7i.8xlarge / 61 GiB host and the identical config that OOM-killed before
(mwm = 8GB, 16 parallel maintenance workers, lists = 3162):
| before (≤ 2.10.0) | after (2.10.1) | |
|---|---|---|
| outcome | OOM-killed at anon-rss 58.47 GiB |
completed in 69.8 min |
| peak private | — | 11.20 GiB (18 % of the host) |
dmesg OOM events |
1 | 0 |
The release notes projected 11.2 GiB term by term; the measurement came in
at 11.20 GiB, validating the root-cause model quantitatively — a
projection derived purely from “the CBOR retention is gone, these four terms
remain” would not land within 0.01 GiB unless the retention really was the
whole defect. Index verified sound: wire v8, 10M/10M slots, is_corrupt =
false, not degraded, scan_fraction = 0.00506 (= 16/3162, so cell pruning is
live), 51 ms warm at probes = 16.
How six hypotheses missed it
Every one looked in the drain (Vec doubling, the reservoir, Vec<Vec<u32>>,
prepare(), the Lloyd cross, the assign sweep), where only ~1 % of the peak
is allocated. Two process lessons, both earned the hard way:
trace_stage!reported cumulative memory, not per-stage deltas, which produced a confidently wrong attribution totrain_kmeans(retracted). It now emits0_at_drain_entryso the scan is measured separately — and that single marker localised the bug in one build.- Check whether a suspect stage is pure code before renting a host.
train_kmeanshas nopgrxreferences; it was exonerated in minutes in a local harness after four sessions of EC2 work.
Phase plan
Phase HV - add✓ done.halfvec(FP16) type.Phase SV - add✓ done.sparsevectype.Phase BV - add✓ done.bitvectype, Hamming + Jaccard.Phase L2 - indexed L2 / L1 ANN.✓ done viavec_l2_ops/vec_l1_opsand the existing recheck-orderby path.Phase D (breadth) - multivector / hybrid SQL surface.✓ done (v1.13.x).turbovec.max_sim/max_sim_cosine(ColBERT MaxSim re-rank overvector[]),turbovec.rrf_score(reciprocal rank fusion), and the named-vector schema pattern. SeeHYBRID_SEARCH.md. Remaining gap: index-native late interaction (per-token index + MaxSim traversal) is a documented future phase — MaxSim is a SQL re-rank primitive, not an index-accelerated scan.- Phase BV-IDX - binary-vector ANN index. The TurboQuant kernel doesn’t fit Hamming-space ANN; if we want indexed bitvec we’d need a separate kernel (LSH or multi-index hashing). Out of scope for the 1.0 line.
- Phase BC - binary-compatible varlena layout for
vectorso casts to/frompgvector.vectorare zero-copy. See . Phase Z1 - make ordinary (TurboQuant) IVF insert degradation reportable.✓ done.listspreserved +ivf_degradedstamped on the degrading flush, soturbovec.index_is_degraded()and theambeginscanWARNING both fire. No format change (both are existing v4 fields). See the zvec review section for the two gotchas it exposed (deferred vs synchronous insert paths;assign_dups > 1is read-only).- Phase Z2 - sparse ANN opclass over
sparsevec(inner product first), so learned-sparse retrieval stops requiring densification. Needs a kernel decision: TurboQuant does not apply to sparse, so this is sparse FLAT (and possibly an inverted/WAND posting scan), not a reuse of the existing path. - Phase Z3 - RESCOPED to a non-gap plus one small item. The original
framing (“a bitmap/allowlist callback so an ordinary
WHEREbecomes a kernel mask”) is not implementable and would not help: a qual on a non-indexed column never reaches an AM (PostgreSQL defines a scan key asindex_key operator constantover an index column), andamgetbitmapreturns an unordered bitmap, destroying the ordering an ANN scan exists to provide. The useful behaviour shipped in v1.8.0 as iterative scan, which is demand-driven and needs no view of the filter. See the rescoped § 3 above. Remaining, genuinely small: the allowlist is a measured 2.6-14.7x win below ~7 % selectivity and a 2.6x LOSS at 100 %, and nothing tells a user which side they are on. Z4 now computes that selectivity, so the open work is guidance (documenting the crossover against a real estimate) and optionally a planner-side hint - not a mask-pushdown feature. Phase Z4 - filter- and kind-aware✓ done. IVF costed for the cells it probes; selectivity from the planner’s ownamcostestimate.rel->rows / rel->tuples; and a pre-existing unit error fixed that had made a full 1M scan ~3000x too cheap. Arithmetic extracted into pure, unit-tested functions (it is not observable throughEXPLAIN- PG’s heap costs swamp it).Phase Z5 (bounded mutable delta)✓ shipped in v2.10.0, and it needed no format change: the delta length is derivable asn_live - cell_directory.total_vectors()because inserts append at the tail. Measured 11 % (in-memory) / 22 % (out-of-core) median win at 1M x 256-d – not the 64x modelled, because a full 4-bit scan there is only 1.6x a 1-cell scan. It ships on the functional contract (no O(n) cliff on the first insert) plus a real out-of-core correctness fix, not the latency delta. The >RAM regime, where the win could be materially larger, is still UNMEASURED – the attempt was blocked by the build-memory ceiling documented above.Phase Z6 - make IVF builds memory-bounded.✓ done. Root cause was the build callback retaining a CBOR-decoded copy of EVERY row in the long-livedambuildcontext (no context switch/reset anywhere in the build path), so the spill wrote the corpus to disk while PostgreSQL held all of it in RAM. Fixed with a per-tuple context. Measured 2M x 1024-d peak 12.16 -> 3.45 GiB (3.5x), drain entry 12.07 -> 2.39 GiB (5.05x), 16 % faster, index bytes unchanged. Also fixed on the way: the Lloydcrossmatrix was quadratic inlists(9.54 GiB at 3162), andturbovec.build_parallelismwas documented as speed-only when it is also worth ~0.27 GiB/thread. Full account in the section above.