Skip to content
NakodaAI

AI EncyclopediaCore concepts

Embedding

Also known as vector embedding

A numeric vector that represents the meaning of a piece of data - a word, sentence, image - such that similar meanings sit close together in vector space.

embeddingsvectorsrepresentation learning

In plain English

An embedding turns 'cat' and 'kitten' into two lists of numbers that happen to sit near each other in a giant coordinate space, because the model learned they mean similar things - while 'cat' and 'spreadsheet' end up far apart. It is how a machine represents meaning as geometry.

Technical explanation

Embedding is a representation-learning technique that maps high-dimensional, complex data into a lower-dimensional vector space, learned from data rather than hand-designed (unlike one-hot encoding). Models are trained so that objects with similar context or meaning are placed near each other under a chosen distance metric, enabling downstream tasks - similarity comparison, clustering, classification, retrieval - to operate on that geometry instead of on raw text or pixels.

Why it matters

Embeddings are the shared numeric language that connects language models, vector databases and semantic search. Every RAG pipeline, recommendation system and semantic-search feature depends on an embedding model as its first step.

Real-world example

OpenAI's text-embedding-3-small and open alternatives such as Sentence-BERT are widely used to embed documents before indexing them in a vector database like Pinecone or Weaviate for retrieval.

Common misunderstanding

That embeddings are interchangeable between models. Vectors from two different embedding models are not directly comparable - mixing them (or switching embedding models without re-indexing existing data) silently breaks similarity search.

Something wrong here?

Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.