Skip to content
NakodaAI

AI EncyclopediaInfrastructure & tooling

Vector Database

A database purpose-built to store embeddings (vectors) and answer 'find me the most similar items' queries fast, using approximate nearest-neighbour search - the storage layer under most RAG systems.

vector databaseembeddingssimilarity searchPineconeWeaviate

In plain English

A normal database finds exact matches - the row where the customer ID equals 4471. A vector database finds nearest neighbours - the documents whose meaning is closest to your question, even if they don't share a single word with it. That is what lets a search understand 'cheap flights to Dubai' and 'affordable Dubai airfare' as basically the same request.

Technical explanation

Text, images or other data are converted into embeddings - numeric vectors, typically produced by a model like OpenAI's text-embedding-3-small or a Sentence-BERT model, positioned in high-dimensional space so that semantically similar inputs land close together. A vector database indexes these vectors for CRUD operations, metadata filtering and horizontal scaling, and answers similarity queries using approximate nearest-neighbour algorithms such as HNSW or IVF, scored by a distance metric like cosine similarity, rather than brute-force comparison against every stored vector.

Why it matters

Vector databases are the retrieval layer under most production RAG deployments and semantic-search products. Choosing one is now a routine infrastructure decision for any team building an AI assistant over its own data.

Real-world example

Pinecone is a fully managed, cloud-only vector database (API key, create index, query); Weaviate is open-source with a built-in vectorizer, hybrid vector-plus-keyword search and a GraphQL API - two different points on the managed-versus-open-source spectrum teams choose between.

Common misunderstanding

That a vector database is a new kind of database engine unrelated to anything before it. Increasingly it is a feature, not a category - PostgreSQL's pgvector extension turns an ordinary relational database into a vector store, and several general-purpose databases now ship native vector-search support.

Something wrong here?

Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.