Computers work with numbers, but people communicate with words, images, sounds, and other rich data. An embedding is a numerical representation that captures useful patterns in that data. For text, an embedding model turns a word, sentence, or document into a list of numbers called a vector. Other embedding models do the same for images, audio, code, or mixed media.
The numbers are not a dictionary definition or a list of hand-written features. They are learned from examples. Classic word embeddings learned which words tend to appear in similar contexts. Modern text embedding models are often trained contrastively: matching pairs, such as a question and its answer, are pulled closer together while non-matching pairs are pushed farther apart.
Meaning as geometry
Imagine each vector as a point in a very large space. The space may have hundreds or thousands of dimensions, so it is impossible to draw directly, but the basic idea is still geometric:
- Items with related meaning tend to be close together.
- Items with different meaning tend to be farther apart.
- Each direction in the space can capture a pattern the model found useful, although people usually cannot name every dimension.
For example, vectors for “puppy” and “dog” may be near each other, while “puppy” and “spreadsheet” are likely farther apart. This does not mean that nearby items are identical. It means the embedding model considers them related in a way that fits its training and design.
From text to a vector
An embedding pipeline often looks like this:
- A tokenizer splits the input into tokens.
- An embedding model processes those tokens and their relationships.
- The model produces one vector for the requested item.
- An application stores the vector or compares it with other vectors.
The output might look like this:
"reset my password" -> [0.18, -0.42, 0.07, 0.91, ...]
Real vectors are much longer than this example. The individual values are usually not meaningful on their own. The useful signal comes from comparing complete vectors.
Embedding models can represent different kinds of data:
- Word embeddings represent individual words or tokens.
- Sentence and document embeddings represent larger pieces of text.
- Image and audio embeddings represent visual or sound content in multimodal systems.
- Multimodal embeddings place different data types in a shared space, such as an image and a caption describing it.
Measuring similarity
To compare two vectors, an application uses a distance or similarity measure. A common choice is cosine similarity, which measures the angle between vectors. Vectors pointing in a similar direction receive a higher similarity score, even if their lengths are different.
Other systems use Euclidean distance or a dot product. The right choice depends on the embedding model and how its vectors were trained. A vector database can use these measures to return the nearest stored vectors.
Consider a support search:
Query: "I cannot sign in because I forgot my password"
Match: "Steps for resetting a forgotten password"
The words are not identical, but their embeddings may be close because they express the same problem. A keyword search might miss this match unless it includes synonyms. An embedding search can find it through meaning.
Try the embeddings playground → Try the vector search playground →
Why embeddings matter in applications
Embeddings turn several useful tasks into comparisons between vectors:
- Semantic search finds information that matches the meaning of a query, not just its exact words.
- Retrieval-augmented generation uses similarity search to select documents for a model’s context.
- Clustering groups related messages, documents, or users without labeling every item first.
- Recommendations find products, articles, or songs that resemble something a person liked.
- Classification uses the position of an item in vector space to help assign a category.
For large collections, a vector database stores the vectors and uses an index to search them efficiently. The database usually stores the original text or metadata alongside each vector so the application can retrieve the source after finding a match.
In production, nearest-neighbor search is usually approximate. The index looks for very close matches quickly instead of comparing the query with every vector one by one. Vector representations covers chunking, indexes, filters, and retrieval evaluation in more detail.
Here is a compact end-to-end example for retrieval-augmented generation:
Index time:
Chunk A: "Reset passwords from Account > Security" -> vector A
Chunk B: "Update billing address from Settings" -> vector B
Chunk C: "Export project data as CSV" -> vector C
Query time:
User asks: "How do I change a forgotten password?"
Query embedding is compared with stored vectors.
Illustrative cosine scores:
A: 0.89
B: 0.28
C: 0.12
Result:
Pass Chunk A to the LLM as context for the answer.
The embedding model does not answer the question by itself. It ranks likely evidence. The application then decides which chunks to send into the model’s context, which is a context engineering choice.
Embeddings are useful, not perfect
An embedding is a model’s learned representation, not an objective measurement of truth. Its behavior depends on the data and training choices behind the embedding model.
- Similarity does not prove that two items are factually equivalent.
- Score thresholds are model-specific;
0.80from one embedding model is not the same promise as0.80from another. - A general-purpose model may perform poorly on specialized technical or business language.
- Language coverage can vary, so quality may differ across languages.
- Embeddings can encode bias from their training data, so similar-looking clusters are not automatically fair or safe.
- Long documents may contain several topics and need to be split into smaller chunks before embedding.
- Changing the embedding model can make old and new vectors incompatible, so collections often need to be re-embedded together.
When evaluating an embedding system, test it with real queries and expected matches. Do not choose a model only because it produces more dimensions or sounds more advanced.
The key idea is simple: embeddings place complex data into a numerical space where relationships can be measured. That shift from exact strings to learned similarity is what makes semantic search, retrieval, clustering, and many recommendation systems possible.