The vector type: a column of floats
pgvector adds a first-class vector type. You declare a column as vector(d) where d is the fixed dimensionality — vector(384) for an all-MiniLM embedding, vector(1536) for an OpenAI text-embedding-3-small vector. Physically it is stored as an array of 32-bit floats plus a small header, so one vector(1536) occupies roughly 1536 × 4 = 6144 bytes, about 6 KB per row before overhead.
Because it is a real column type, an embedding lives in the same row as the text, the author id, and every other attribute. There is no separate collection to keep in sync, no dual-write problem, no eventual consistency between a document store and a vector store — an INSERT writes the row and its embedding atomically, in one transaction. pgvector also ships halfvec (16-bit floats, half the storage), bit, and sparsevec, but the dense vector type carries most workloads and is the one this article’s math assumes.
Distance operators and what they compute
Similarity search is really distance ranking, and pgvector exposes distance as infix operators so a query reads like ordinary SQL. The three that matter:
a <-> b L2 (Euclidean): sqrt( Σ_i (a_i - b_i)^2 )
a <=> b cosine distance: 1 - (a·b) / (|a| |b|)
a <#> b inner product: -(a·b) (negated)The inner-product operator returns the negative dot product on purpose: Postgres orders ascending, and negating turns ‘largest similarity’ into ‘smallest value,’ so ORDER BY embedding <#> query still puts the best match first. A nearest-neighbour query is then simply SELECT id FROM docs ORDER BY embedding <=> '[...]' LIMIT 10. Cosine ignores magnitude and compares direction, which suits most text embeddings; L2 accounts for magnitude; inner product is fastest and is correct when vectors are already normalized, because for unit vectors a·b and cosine similarity coincide — so matching the operator to how the model was trained changes which rows come back.
Operator classes: binding an index to a metric
Here is the pgvector detail that trips people up. An index does not accelerate every distance operator; it accelerates one metric, fixed at creation time by the operator class you name. The three classes mirror the three operators: vector_l2_ops, vector_cosine_ops, and vector_ip_ops.
So CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops) builds a structure that only answers <=> queries quickly. If your query orders by <-> against that index, Postgres cannot use it and silently falls back to a sequential scan — correct results at exact-search cost, exactly the slowness the index was meant to remove. The operator in your ORDER BY must match the operator class in your CREATE INDEX; to rank by two metrics you build two indexes. When a query is unexpectedly slow, this mismatch is the first thing to check with EXPLAIN.
Exact search and why you eventually leave it
With no vector index, pgvector answers a nearest-neighbour query exactly: it computes the distance from the query to every candidate row and keeps the top k. This is a brute-force k-NN, an O(N · d) scan, and its virtue is that recall is 100 % by definition — no approximation to be wrong about.
For small tables that is the right answer: scanning ten thousand vector(768) rows is microseconds of work, and an approximate index would add build and maintenance cost for no felt benefit. The problem is linear growth: at one million rows every query touches a million vectors, and at ten million the sequential scan dominates your request. That is where you trade a little correctness for a large speedup and reach for an approximate-nearest-neighbour (ANN) index — and pgvector lets you make that switch without leaving Postgres, the same column and query with an index underneath.
IVFFlat: partition, then probe a few cells
IVFFlat (inverted file with flat storage) is the simpler ANN index. At build time it runs k-means over a sample of your vectors to pick lists centroids, carving the space into that many Voronoi cells and assigning every vector to its nearest centroid. A query compares itself to the centroids and searches only the closest probes cells instead of the whole table.
Two knobs govern it. lists is set at index creation — a common starting heuristic is rows / 1000 for up to a million rows, then sqrt(rows) beyond that. probes is set per query with SET ivfflat.probes and trades recall for speed directly: probes = 1 is fastest and least accurate, and as probes climb toward lists you approach exact search. Concretely, on a million rows at lists = 1000 each cell holds ~1000 vectors, so probes = 1 scans ~1000 comparisons instead of a million (a 1000× cut) but misses true neighbours across a cell boundary, and recall might sit near 0.7; probes = 10 scans ten cells and often lifts recall past 0.95 — which is why you validate by measuring recall@k rather than trusting a default. The catch pgvector is explicit about: build the IVFFlat index only after the table has representative data, because the centroids are learned from whatever is present — index an empty or tiny table and the partitioning is garbage.