← Back to concepts
8 min read

Embeddings and vector representations

Computers work with numbers, but people communicate with words, images, sounds, and other rich data. An embedding is a numerical representation that captures useful patterns in that data. For text, an embedding model turns a word, sentence, or document into a list of numbers called a vector. Other embedding models do the same for images, audio, code, or mixed media.

The numbers are not a dictionary definition or a list of hand-written features. They are learned from examples. Classic word embeddings learned which words tend to appear in similar contexts. Modern text embedding models are often trained contrastively: matching pairs, such as a question and its answer, are pulled closer together while non-matching pairs are pushed farther apart.

Meaning as geometry

Imagine each vector as a point in a very large space. The space may have hundreds or thousands of dimensions, so it is impossible to draw directly, but the basic idea is still geometric:

  • Items with related meaning tend to be close together.
  • Items with different meaning tend to be farther apart.
  • Each direction in the space can capture a pattern the model found useful, although people usually cannot name every dimension.

For example, vectors for “puppy” and “dog” may be near each other, while “puppy” and “spreadsheet” are likely farther apart. This does not mean that nearby items are identical. It means the embedding model considers them related in a way that fits its training and design.

Related items cluster together in embedding space A simplified two-dimensional map shows two clusters. Puppy, dog, kitten, and cat sit close together in an animals cluster. Spreadsheet, invoice, report, and budget sit together in an office documents cluster on the other side. A short line marks puppy and dog as close, and a long dashed line marks cat and spreadsheet as far apart. Embedding space, flattened to 2D for illustration animals office documents close far: unrelated meaning puppy dog kitten cat spreadsheet invoice report budget Positions are illustrative. Real embeddings use hundreds or thousands of dimensions, not two.
Distance reflects what the model was trained to treat as similar, such as topic, meaning, style, or language.

From text to a vector

An embedding pipeline often looks like this:

  1. A tokenizer splits the input into tokens.
  2. An embedding model processes those tokens and their relationships.
  3. The model produces one vector for the requested item.
  4. An application stores the vector or compares it with other vectors.
The embedding pipeline from text to vector The input text reset my password is split by a tokenizer into tokens. An embedding model reads the tokens in context and outputs one vector of numbers. The application then stores that vector or compares it with other vectors. The number of tokens varies with the input, but for a given embedding model, the output vector has the same length for every input. Input text "reset my password" Tokenizer [reset] [ my] [ password] Embedding model reads tokens in context One vector [0.18, -0.42, 0.07, 0.91, ...] Store or compare vector database similarity search Token count varies with the input Same length for that model (e.g. 768 numbers)
Many tokens go in, but one fixed-length vector comes out for that model, which is what makes any two inputs directly comparable.

The output might look like this:

"reset my password" -> [0.18, -0.42, 0.07, 0.91, ...]

Real vectors are much longer than this example. The individual values are usually not meaningful on their own. The useful signal comes from comparing complete vectors.

Embedding models can represent different kinds of data:

  • Word embeddings represent individual words or tokens.
  • Sentence and document embeddings represent larger pieces of text.
  • Image and audio embeddings represent visual or sound content in multimodal systems.
  • Multimodal embeddings place different data types in a shared space, such as an image and a caption describing it.

Measuring similarity

To compare two vectors, an application uses a distance or similarity measure. A common choice is cosine similarity, which measures the angle between vectors. Vectors pointing in a similar direction receive a higher similarity score, even if their lengths are different.

Other systems use Euclidean distance or a dot product. The right choice depends on the embedding model and how its vectors were trained. A vector database can use these measures to return the nearest stored vectors.

Consider a support search:

Query:   "I cannot sign in because I forgot my password"
Match:   "Steps for resetting a forgotten password"

The words are not identical, but their embeddings may be close because they express the same problem. A keyword search might miss this match unless it includes synonyms. An embedding search can find it through meaning.

Ranking support articles by cosine similarity Three vectors start from a shared origin. The query vector and the password reset article vector are 25 degrees apart, giving an illustrative cosine similarity of 0.91. The billing address article vector is 83 degrees away from the query, giving 0.12. The password reset article ranks first even though its vector is shorter, because cosine similarity compares direction rather than length. Query: "I cannot sign in because I forgot my password" Query Reset article Billing article Ranked results (illustrative scores) 1 Steps for resetting a forgotten password small angle: 25 deg -> cosine 0.91 2 Update your billing address large angle: 83 deg -> cosine 0.12 The reset article's vector is shorter, but cosine ignores length.
Search ranks stored items by how closely their vectors point in the same direction as the query vector, not by shared keywords.
**See it for yourself:** [Explore embeddings](/playgrounds/embeddings/) to watch related words cluster together, or [explore vector search](/playgrounds/vector-search/) to see a paraphrased query retrieve the right match without sharing a single keyword.

Try the embeddings playground → Try the vector search playground →

Why embeddings matter in applications

Embeddings turn several useful tasks into comparisons between vectors:

  • Semantic search finds information that matches the meaning of a query, not just its exact words.
  • Retrieval-augmented generation uses similarity search to select documents for a model’s context.
  • Clustering groups related messages, documents, or users without labeling every item first.
  • Recommendations find products, articles, or songs that resemble something a person liked.
  • Classification uses the position of an item in vector space to help assign a category.

For large collections, a vector database stores the vectors and uses an index to search them efficiently. The database usually stores the original text or metadata alongside each vector so the application can retrieve the source after finding a match.

Indexing and querying with a vector database In the indexing phase, done ahead of time, documents are split into chunks, embedded, and stored in a vector database along with the chunk text and source metadata. In the query phase, a user query is embedded by the same embedding model and used to search the database. The nearest chunks and their source text are returned for search results or as context for retrieval-augmented generation. 1. INDEXING - ahead of time Documents help articles, FAQs Split into chunks one topic each Embedding model chunk -> vector store must be the same model 2. QUERY - at request time User query "I forgot my password" Same embedding model query -> vector search Vector database [0.12, ..] chunk text source [-0.3, ..] chunk text source [0.71, ..] chunk text source Index finds nearest vectors fast Nearest chunks + source text -> search results or RAG context
Documents are embedded once ahead of time; each query is embedded with the same model and matched against the stored vectors.

In production, nearest-neighbor search is usually approximate. The index looks for very close matches quickly instead of comparing the query with every vector one by one. Vector representations covers chunking, indexes, filters, and retrieval evaluation in more detail.

Here is a compact end-to-end example for retrieval-augmented generation:

Index time:
  Chunk A: "Reset passwords from Account > Security" -> vector A
  Chunk B: "Update billing address from Settings"    -> vector B
  Chunk C: "Export project data as CSV"              -> vector C

Query time:
  User asks: "How do I change a forgotten password?"
  Query embedding is compared with stored vectors.

Illustrative cosine scores:
  A: 0.89
  B: 0.28
  C: 0.12

Result:
  Pass Chunk A to the LLM as context for the answer.

The embedding model does not answer the question by itself. It ranks likely evidence. The application then decides which chunks to send into the model’s context, which is a context engineering choice.

Embeddings are useful, not perfect

An embedding is a model’s learned representation, not an objective measurement of truth. Its behavior depends on the data and training choices behind the embedding model.

  • Similarity does not prove that two items are factually equivalent.
  • Score thresholds are model-specific; 0.80 from one embedding model is not the same promise as 0.80 from another.
  • A general-purpose model may perform poorly on specialized technical or business language.
  • Language coverage can vary, so quality may differ across languages.
  • Embeddings can encode bias from their training data, so similar-looking clusters are not automatically fair or safe.
  • Long documents may contain several topics and need to be split into smaller chunks before embedding.
  • Changing the embedding model can make old and new vectors incompatible, so collections often need to be re-embedded together.

When evaluating an embedding system, test it with real queries and expected matches. Do not choose a model only because it produces more dimensions or sounds more advanced.

The key idea is simple: embeddings place complex data into a numerical space where relationships can be measured. That shift from exact strings to learned similarity is what makes semantic search, retrieval, clustering, and many recommendation systems possible.