What a collection actually stores per object
A Weaviate collection is not one index, it is four things sharing a shard. There is an object store (an LSM-tree keyed by UUID holding the JSON properties), an inverted index for filters and BM25, the vector index itself, and the id maps that tie an external UUID to the internal integer the graph actually uses.
The consequence people miss: the vector is persisted twice. It is written alongside the object so a shard can be rebuilt from scratch, and held again in the index that serves queries. Disk carries roughly N · (4d + |props|); RAM carries the index copy plus graph links. And named vectors multiply this cleanly — giving one object a title and a body embedding means two vector spaces, two graphs, two sets of links. Cost scales as k · (4d + links · id_width) per object, not as one index that happens to be wider.
The HNSW knobs, and the ef Weaviate computes for you
Weaviate exposes the usual trio under its own names: maxConnections (edges per node above layer 0; layer 0 gets twice that), efConstruction (build-time beam width, paid once and stored as graph quality), and ef (search-time beam width, paid every query). The Qdrant article derives why the total link count per node is close to 2M — that arithmetic carries over unchanged.
What is Weaviate’s own is that ef defaults to -1, meaning dynamic: it is derived from the query, not fixed on the collection.
ef = min( max(dynamicEfMin, dynamicEfFactor · limit), dynamicEfMax )
with factor 8, min 100, max 500:
limit = 10 → max(100, 80) = 100
limit = 50 → max(100, 400) = 400
limit = 100 → min(800, 500) = 500The trap is that recall now depends on limit. Asking for 10 results and asking for 100 run different searches, so the object ranked 12th in one need not be ranked 12th in the other — and deep pagination silently saturates at the ceiling.
Thresholds counted in objects, not bytes
Two Weaviate limits are denominated in object counts, and both are therefore blind to dimensionality. flatSearchCutoff (commonly 40,000) says: if a filter’s allow-list is smaller than this, abandon the graph and brute-force the matching vectors instead. vectorCacheMaxObjects caps how many vectors the index keeps resident before it starts reading them from disk during a walk.
Convert the first one and the blind spot appears. At d = 384 float32, 40,000 objects is 61 MB — a trivial scan. At d = 1536 the same count is 40000 × 6144 B = 246 MB, four times the memory traffic for an identical setting. Qdrant makes the opposite choice and denominates its threshold in kilobytes, precisely because a linear scan is bandwidth-bound. The rule: a count-based default is calibrated for some assumed dimension, so lower it on wide vectors, and treat the cache limit the same way.
The dynamic index: flat until it is worth a graph
Weaviate can start a collection as a flat index — no graph, a disk-backed brute-force scan, optionally over binary-compressed vectors — and convert it to HNSW once the object count crosses a threshold (on the order of 10,000). The crossover arithmetic explains why the switch point sits so low.
N = 10,000 d = 1536 float32
flat scan : 10^4 × 6144 B = 61 MB touched, 1.5e7 mul-adds
HNSW @ ef=100: a few thousand distance evals ≈ 5e6 mul-adds
+ 10^4 × 65 links × 8 B ≈ 5 MB graph
+ O(N · efConstruction · log N) build
→ graph wins by ~3× here, but by ~100× at N = 10^6So the crossover is a wide band, not a point: below it the graph buys under an order of magnitude while costing build CPU, memory and index lag; above it the scan is hopeless. Defaulting to the low end is the right asymmetry — over-building a small index is cheap, and a 100 ms scan is not.
Multi-tenancy: a shard each, and the cost of many small tenants
Weaviate’s multi-tenancy is physical. Each tenant gets its own shard — its own LSM store, inverted index and HNSW graph. A tenant-scoped query never touches another tenant’s vectors, which means the filtered-search problem that dominates single-index designs simply does not arise: there is no induced subgraph to fragment, because the graph was never shared.
You pay for that in fixed per-shard overhead — memtables, segment metadata, bloom filters, file descriptors, index structs. Call it C bytes per active shard:
50,000 tenants × 2,000 objects, d = 768
vectors if all hot : 10^8 × 3072 B = 307 GB
fixed overhead : 50,000 × C; C = 2 MB → 100 GB
— before a single vector is storedWhich is why tenants have activity states: hot in memory, cold on local disk, or offloaded to object storage and lazily reloaded. The number that sizes your cluster is the count of simultaneously active tenants, never the count of tenants.
Hybrid search: alpha, and two fusions that disagree
Weaviate’s hybrid runs a BM25 query and a vector query and blends them with alpha: alpha = 1 is pure vector, alpha = 0 is pure keyword. What matters more than alpha is which fusion consumes it. rankedFusion throws the scores away and combines 1/(60 + rank) per list. relativeScoreFusion (the modern default) min-max normalizes each list, then blends the normalized scores. Same alpha, different answers:
vector sim : X .95 Y .94 Z .60 → norm X 1.00 Y .971 Z 0
bm25 : Y 20 Z 3 X 2 → norm Y 1.00 Z .056 X 0
relativeScore, α=0.5 : Y .986 X .500 Z .028
rankedFusion, α=0.5 : Y .01626 X .01613 Z .01600Both rank Y first, but rank fusion compresses the field into a 1.6% spread while relative-score preserves the margins — X’s near-tie on vector and Y’s BM25 blowout both survive. Rank fusion is robust to a miscalibrated retriever; relative-score is faithful to a well-calibrated one, and lets one runaway keyword hit dominate even at α = 0.5.