Vector Database

A vector database stores embeddings — the lists of numbers that represent your text. It answers one question quickly: which of these are closest to this other one? That's it. Everything else it does is in service of that.

Worth saying early, because the whole category is sold hard: most projects don't need a dedicated one. The capability matters. The separate product often doesn't.

In this guide
  1. The search underneath it
  2. Exact or approximate
  3. The category is wider than the products
  4. Do you actually need a dedicated one?
  5. Filtering is where this gets hard
  6. What it doesn't replace

The search underneath it

The core operation is nearest-neighbour search. You arrive with a query vector, and the database returns the stored vectors sitting closest to it, usually the closest 10 or 20.

The naive way to do that is to compare the query against every stored vector and sort the results. This is called a flat or brute-force index, and it has one great virtue: it's exact. It finds the genuinely closest matches every time. Its cost grows linearly, so it gets slow as the collection grows — but "slow" arrives much later than people assume. Tens of thousands of vectors is nothing.

Exact or approximate

Past some size, scanning everything stops being viable, and the answer is approximate nearest-neighbour search. You build an index that finds almost the closest matches, far faster, by not looking at most of the collection.

The common approach is HNSW — hierarchical navigable small world — which arranges vectors into a layered graph and walks it from a coarse layer down to a fine one, checking only a small fraction of the data. Others cluster the vectors and search only the nearest clusters, or compress the vectors so more fit in memory.

What matters is that this is an explicit trade, not free speed. You give up some recall — the share of true nearest neighbours the search actually returns — in exchange for latency and memory. Every one of these indexes has knobs controlling where you sit on that curve, and the defaults are usually sensible. The thing to know is that approximate means approximate: a relevant chunk can be missed because of the index, not because the embedding was wrong.

The category is wider than the products

Three different things get called vector databases, and the distinction is practical.

Dedicated systems — Pinecone, Weaviate, Qdrant, Milvus, Chroma and others — are built for this from the ground up, and bring their own serving, scaling and operational story.

Extensions to databases you already run. pgvector adds vector columns and nearest-neighbour indexes to Postgres; similar extensions exist for SQLite, Redis and the big search engines. Your vectors live beside your ordinary data, in the system you already back up and monitor.

Libraries, which are not databases at all. FAISS and hnswlib give you the index and the search, in your own process. No persistence, no network service, no access control — you build those yourself, or you don't need them. Reaching for a library when you wanted a database is a real mistake, and so is the reverse.

Do you actually need a dedicated one?

If you already run Postgres and your collection is in the thousands to low millions of chunks, pgvector is usually the right first move. It's one fewer system to operate, and your filters, joins and transactions stay in one place instead of being split across two stores that can disagree.

Published comparisons generally put the crossover somewhere in the millions of vectors, though the specific numbers come from people with products to sell, so treat them as directional rather than precise. The more reliable signal is which of these you're actually hitting:

  • Scale past what one database server handles comfortably.
  • Search traffic heavy enough that you want it isolated from your main application database.
  • A deliberate decision to make search somebody else's operational problem — a legitimate reason, and often the real one.

None of those are about vector search being special. They're the ordinary reasons any workload eventually gets its own system. And you can move later: swapping the store is far less painful than the re-embedding you'd face if you changed embedding models, because the vectors themselves transfer unchanged.

Filtering is where this gets hard

Real queries are rarely just "find similar text." They're "find similar text in this customer's documents, from this year, excluding drafts." Vector search and ordinary filtering have to work together, and that's where naive setups break.

Filter after the search and you can end up with far too few results: ask for the closest 20, discard the 18 that fail your filter, and you've got 2. Filter before the search and the index has to support it, or you're back to scanning. Every serious implementation has an answer to this, and it's worth checking what that answer is before you commit — especially if your filters are narrow, because narrow filters are where the after-the-fact approach falls apart.

This is also the practical reason a database you already run can win. Postgres was doing the "where" part competently for decades before anyone stored an embedding in it.

What it doesn't replace

It isn't a home for your source documents. It holds vectors, the chunk text, and whatever metadata you'll filter or cite by — the originals belong wherever they already live. Treating the vector store as your primary record is a good way to discover you can't rebuild your index after changing anything about how you chunk or embed.

Practice interview questions on vector databases →