AI EncyclopediaCore concepts
Embedding
A numeric vector that represents the meaning of a piece of data - a word, sentence, image - such that similar meanings sit close together in vector space.
In plain English
An embedding turns 'cat' and 'kitten' into two lists of numbers that happen to sit near each other in a giant coordinate space, because the model learned they mean similar things - while 'cat' and 'spreadsheet' end up far apart. It is how a machine represents meaning as geometry.
Technical explanation
Embedding is a representation-learning technique that maps high-dimensional, complex data into a lower-dimensional vector space, learned from data rather than hand-designed (unlike one-hot encoding). Models are trained so that objects with similar context or meaning are placed near each other under a chosen distance metric, enabling downstream tasks - similarity comparison, clustering, classification, retrieval - to operate on that geometry instead of on raw text or pixels.
Why it matters
Embeddings are the shared numeric language that connects language models, vector databases and semantic search. Every RAG pipeline, recommendation system and semantic-search feature depends on an embedding model as its first step.
Real-world example
OpenAI's text-embedding-3-small and open alternatives such as Sentence-BERT are widely used to embed documents before indexing them in a vector database like Pinecone or Weaviate for retrieval.
Common misunderstanding
That embeddings are interchangeable between models. Vectors from two different embedding models are not directly comparable - mixing them (or switching embedding models without re-indexing existing data) silently breaks similarity search.
Something wrong here?
Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.

