The short answer
A vector database is a database built to store vector embeddings, lists of numbers that represent the meaning of text, images or other data, and to quickly find the stored items whose embeddings are closest to a query's. It lets software search by similarity of meaning rather than exact keywords, which is why it sits underneath most RAG systems and semantic search features.
Key takeaways
- A vector database answers one question very fast: which stored items are most similar to this one?
- It uses approximate nearest neighbour indexes such as HNSW, trading a small amount of accuracy for large gains in speed.
- The difference between a vector database and a vector index library is everything around the search: updates, metadata filters, backups, access control.
- For many business systems, the pgvector extension in an existing PostgreSQL database is enough; a dedicated product earns its place at larger scale or with specialised needs.
- Vectors are derived from your documents, so they carry the same privacy and residency obligations as the source data.
What is a vector database, in simple terms?
A vector database finds things that mean the same, even when they share no words. A normal database is excellent at exact questions: every invoice over $10,000, every customer in Victoria. It’s poor at fuzzy ones: which support tickets describe the same fault as this one, which policy clauses deal with working from home when the clause says “remote arrangements”.
The fix is to turn each item into an embedding, a list of a few hundred to a few thousand numbers produced by an AI model, where items with similar meaning get similar numbers. Once everything is a point in that numeric space, “find similar items” becomes “find the nearest points”. A vector database is the system that stores those points and does that search quickly, at scale, while behaving like a proper database.
What is a vector database used for?
Its main use in AI is semantic search: finding the passages, products or records closest in meaning to a query. Common uses:
- Retrieval for RAG. The vector database for RAG holds your document chunks so a language model can be given the most relevant ones for each question.
- Semantic search on an intranet, help centre or product catalogue, where people search in their own words.
- Memory for AI agents, storing past notes or conversations the agent can look up later.
- Recommendations, such as related articles or similar products.
- Deduplication and matching, such as spotting near-identical tickets, invoices or records.
- Image and multimodal search, where images and text are embedded into the same space.
How does a vector database work?
It stores each vector with an ID and metadata, builds an index that makes nearest-neighbour search fast, and returns the closest matches for a query vector. In plain steps:
- Write. Your application sends a record: an ID, the vector, and metadata such as source document, date, department and access group. Most systems also store the original text.
- Index. The database organises vectors into a search structure. The most common today is HNSW (Hierarchical Navigable Small World), a layered graph described by Malkov and Yashunin in 2016. A search enters at a sparse top layer and hops towards closer neighbours through denser layers below.
- Query. Your application embeds the user’s query with the same model and asks for the top matches, for example the 20 nearest vectors where access group is “finance” and date is after 2024.
- Score. Closeness is measured with a distance function, usually cosine similarity or dot product, and results come back ranked.
The key word is approximate. Comparing a query against every stored vector gives perfect results but gets slow as the collection grows. Approximate nearest neighbour (ANN) indexes check only a small, well-chosen fraction and usually still find nearly all of the true nearest matches. You tune the trade-off. In pgvector, for example, the ef_search setting controls how much of the HNSW graph each query explores: AWS’s June 2026 production guidance calls the default of 40 “often too low for production” and suggests 100 as a starting point.
Vector database vs relational (SQL) database
A relational database finds rows that match exact conditions; a vector database finds items that are close in meaning. Most businesses need both, and they often live side by side.
| Relational (SQL) database | Vector database | |
|---|---|---|
| Example | PostgreSQL, SQL Server, MySQL | pgvector, Pinecone, Qdrant, Weaviate, Milvus |
| Stores | Rows and columns: customers, orders, invoices | Embeddings plus metadata and usually the source text |
| Typical question | ”All orders over $10,000 from Victoria in June" | "The 10 passages most similar to this question” |
| Result | Exact matches | Ranked by similarity score |
| Good for | Transactions, reporting, systems of record | Semantic search, RAG, recommendations |
The line is blurring: pgvector adds vector columns to PostgreSQL, and Azure AI Search and Amazon OpenSearch Service support vector search alongside keyword search. A graph database, or a knowledge graph, is different again: it stores explicit relationships between entities (this supplier supplies that part) and answers questions by following them, while a vector database stores meaning and answers by similarity. Some RAG systems combine the two.
Vector database vs vector index vs search engine
The terms get blurred. Here is how they differ.
| Vector index library | Vector database | Keyword search engine | |
|---|---|---|---|
| Example | FAISS | pgvector, Pinecone, Qdrant, Weaviate | Elasticsearch, OpenSearch (BM25) |
| Finds matches by | Similarity of vectors | Similarity of vectors, plus filters | Shared words, weighted by rarity |
| Updates and deletes | Limited; often rebuild | Normal inserts, updates, deletes | Normal |
| Metadata filtering | Build it yourself | Built in | Built in |
| Backups, access control, replication | Not included | Included | Included |
| Good at exact codes and names | No | Weak | Strong |
FAISS, released by Facebook AI Research and described by Johnson, Douze and Jégou in 2017, is a library you embed in your own code: very fast, but it leaves storage, updates and security to you. A vector database wraps that kind of index in database features. Many search engines now support vectors too, and many vector databases now support keyword search, which matters because combining both (hybrid search) usually beats either alone.
Vector database examples
Vector databases fall into three groups: vector support added to a database you may already run, dedicated vector databases, and managed stores built into AI platforms.
| Group | Examples | Typical fit |
|---|---|---|
| Extension to an existing database | pgvector for PostgreSQL (available on Amazon RDS and Aurora, Azure Database for PostgreSQL, Google Cloud SQL and Supabase) | Most business systems up to a few million vectors |
| Dedicated vector database, open source | Qdrant, Weaviate, Milvus | Self-hosting, heavy filtering, large scale |
| Managed vector service | Pinecone, or managed editions of the open-source engines | Large scale with a small operations team |
| Search engines with vector support | Elasticsearch, OpenSearch, Azure AI Search | Teams that already run search and want hybrid ranking |
| Managed retrieval in an AI platform | Amazon Bedrock Knowledge Bases, which manages the vector store for you | Teams building end to end on one cloud platform |
Do you need a dedicated vector database?
Often not at first. If you already run PostgreSQL, the pgvector extension adds vector columns and HNSW indexes to the database you have. That means one system to back up, secure and host in an Australian region, and the ability to join vector results with your ordinary tables in one query.
A dedicated vector database starts to make sense when:
- You have tens of millions of vectors or more, or very high query volumes.
- Searches combine similarity with heavy, complex filtering and pgvector’s performance on your data isn’t enough.
- You want vector search managed as its own scalable service, separate from your transactional database.
- You need features a specialist engine does better, such as multi-vector records or built-in hybrid ranking.
A simple decision matrix:
| Your situation | Sensible starting point |
|---|---|
| Already on PostgreSQL, under a few million vectors | pgvector |
| Using a managed AI platform end to end | The platform’s built-in store (for example Bedrock Knowledge Bases) |
| Very large scale, small ops team | A managed vector service in a suitable region |
| Strict self-hosting or air-gapped requirement | Self-hosted pgvector or an open-source engine |
| Prototype over a few hundred documents | An in-memory index, or no vector search at all |
For a product-by-product view, see pgvector vs Pinecone vs Weaviate vs Qdrant.
Common pitfalls
The most expensive mistakes are about data and operations, not the choice of product.
- Mixing embedding models. Vectors from different models, or different versions of one model, aren’t comparable. Changing models means re-embedding everything, so record the model name with every vector.
- Filtering after the search. If you fetch the top 20 matches and then drop the ones a user can’t see, you may return three results or none. Filter inside the search. Engines such as Qdrant are built around this, and pgvector 0.8.0 added iterative index scans to address it.
- Running out of memory. HNSW indexes perform well only when they fit in RAM. Once they spill to disk, latency jumps.
- No deletion path. When a source document is removed or someone exercises a privacy request, its vectors must go too. Design the link from source to vectors on day one.
- Treating vectors as anonymous. They are derived from your content and usually stored beside the original text. Apply the same residency and access rules as the source; see data residency vs sovereignty.
Worked example: sizing a small knowledge base
Say a firm has 20,000 documents averaging 10 pages, split into chunks of roughly half a page.
- 20,000 documents × 20 chunks = 400,000 chunks.
- At 1,536 dimensions stored as 4-byte floats, each vector is about 6 KB.
- 400,000 × 6 KB ≈ 2.5 GB of raw vectors.
- Add the HNSW index, metadata and chunk text, and plan for several times that on disk, with the index held in memory.
That comfortably fits a modest managed PostgreSQL instance with pgvector. It’s a useful reminder that most business knowledge bases are not “big data” problems.
Where vector databases fit with RAG and agents
A vector database is one component. In a RAG system it’s the retrieval layer that finds passages for the model to read. In an AI agent it may sit behind a search tool the agent can call. The quality of what comes out still depends on parsing, chunking and the embedding model upstream.
How All Webbed Labs approaches vector storage
Our default is PostgreSQL with pgvector in an Australian cloud region, with hybrid keyword and vector search, permission filters applied in the query, and the embedding model version stored against every vector. We move to a dedicated engine only when measured load or filtering needs justify the extra system. See our database architecture and RAG knowledge base services.
Frequently asked questions
Is a vector database a replacement for my normal database?
No. It complements it. Your transactional data, customers and orders stay in a relational database. The vector store holds embeddings plus enough metadata to filter and link back to the source. With pgvector, both can live in the same PostgreSQL instance.
Which vector database should we choose?
Start from what you already run and your scale. If you use PostgreSQL and have up to a few million vectors, pgvector is usually the simplest choice. Managed services such as Pinecone, or engines such as Qdrant and Weaviate, suit larger scale, heavy filtering or teams that want vector search as a separate service. Our comparison guide covers the trade-offs.
Can a vector database be hosted in Australia?
Yes. pgvector runs on Amazon RDS and Aurora PostgreSQL, Azure Database for PostgreSQL and Google Cloud SQL in their Australian regions, and open-source engines can be self-hosted in any region. For managed vector services, check the vendor's current region list, as it changes.
Can someone reconstruct my documents from the vectors?
Treat it as possible. Research has shown that text can be partly recovered from some embeddings, and vector stores usually keep the original chunk text next to each vector anyway. Apply the same encryption, access control and residency rules you apply to the source documents.
How many dimensions does a vector database have?
The database doesn't set the number; the embedding model does. Common models produce vectors of 256 to 3,072 dimensions, and every vector in one index must have the same number. More dimensions mean more storage and memory per item.
What is the difference between a vector database and a vector store?
In practice the terms overlap. "Vector store" is often used in AI frameworks for any component that saves embeddings and runs similarity search, which could be an in-memory index, pgvector or a dedicated product. "Vector database" usually implies the full database features: persistence, updates, filtering, backups and access control.
How much does a vector database cost?
It depends mostly on the number of vectors, their dimensions, query volume and whether the index must sit in memory. pgvector adds no licence cost to a PostgreSQL instance you already pay for; managed vector services charge for storage, reads and writes or capacity. Check the vendor's current pricing page, and size your data first using the worked example on this page.
How big does a vector database get?
A rough rule: vectors take dimensions × 4 bytes each when stored as 32-bit floats, plus index overhead and the stored text. One million chunks at 1,536 dimensions is about 6 GB of raw vectors before indexing.
Sources
- Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs (Malkov and Yashunin) , arXiv
- Billion-scale similarity search with GPUs (Johnson, Douze and Jégou) , arXiv
- What is a Vector Database and How Does it Work? , Pinecone
- Running pgvector in production on Amazon Aurora PostgreSQL , Amazon Web Services
- Filtering , Qdrant