The real question is operational, not algorithmic
Every option on this list can compute approximate nearest-neighbor search over embeddings -- the algorithmic core is not where these systems actually differ for most teams. Where they differ is what it costs to run one, how it fits into a system that already has a database, and what happens at the query volume and dimensionality where the easy option stops being easy. This site's vector database deep-dive covers a single technology's internals; this article is the decision between technologies, for the specific case of an agent's memory layer -- conversation history, retrieved facts, and long-term episodic recall.
Three shapes cover almost every real deployment: a vector extension on a database you already run, a managed hosted vector database, and a self-hosted open-source vector database. Each is a genuinely different operational bet, not just a different API.
pgvector: one database, one operational surface
pgvector adds vector columns, distance operators, and approximate-index types (IVFFlat, HNSW) directly to Postgres. The pitch is almost entirely operational: if the rest of the application's data already lives in Postgres, agent memory can live in the same database, transactionally consistent with everything else, backed up by the same process, monitored by the same dashboards, and queried by the same connection pool.
The practical ceiling is real but higher than its reputation suggests: HNSW indexes in modern Postgres handle millions of vectors with sub-100ms query latency on reasonable hardware, which covers the memory store for the overwhelming majority of single-tenant or moderate-multi-tenant agent deployments. Where it genuinely strains: very high query throughput contending with the same database's transactional write load, vector counts in the tens of millions and beyond where index build/rebuild time becomes an operational event, and workloads wanting index types or filtering performance pgvector doesn't yet match a purpose-built engine on.
The failure mode to watch for is not capacity -- it's coupling. A vector workload that grows faster than the rest of the schema starts dominating the shared database's resource budget, and now a memory-layer scaling problem is also a core-application-availability problem. That coupling is the actual argument for splitting out, not raw vector count.
Managed vector databases: pay for not operating it
A managed, hosted vector database (Pinecone is the reference example of this shape) sells the elimination of the operational surface entirely: no index tuning, no capacity planning, no upgrade windows, a control plane and an API. For a team without dedicated infrastructure engineering, or one that wants to spend its engineering time on the agent's behavior rather than on database operations, this is a legitimate and often correct trade -- the cost is a recurring bill and a new external dependency in the critical path, not a comparable amount of engineer-hours.
The trade-offs worth pricing in explicitly: data now lives outside your own infrastructure boundary, which has real implications for data residency and compliance depending on what's in the embeddings' source text; latency includes a network hop to a third-party service rather than a local database round-trip; and the bill scales with usage in a way that's easy to underestimate at prototype scale and genuinely expensive at high query volume. None of these are disqualifying -- they're the actual price of the operational simplicity, and for many teams that price is worth paying.
Self-hosted open-source vector databases: control without the managed premium
A self-hosted open-source vector database (Qdrant is the reference example) sits between the two: purpose-built for vector search specifically, so it typically outperforms a general-purpose database's vector extension at scale and offers filtering/payload features tuned for retrieval workloads -- but you're running and operating it yourself, which reintroduces the operational surface pgvector's pitch eliminates and the managed option's pitch also eliminates.
This is the right choice when a team has both the infrastructure capability to run a new stateful service well (backups, upgrades, capacity planning, on-call coverage) and a genuine reason to avoid the managed option's cost or data-locality constraints. It's the wrong choice when neither is true -- a team that reaches for a dedicated vector database purely because it sounds more "production-grade" than an extension on their existing Postgres, without the operational capacity to run it properly, often ends up with a worse-operated system than either alternative.