Embeddings are useful because they turn messy inputs — sentences, documents, images, audio, or code — into vector representations. A vector representation is a fixed list of numbers that a system can compare, store, and search.
The important idea is not that the numbers are meaningful by themselves. The meaning comes from their position relative to other vectors produced by the same model.
This article focuses on retrieval-system engineering: how chunks become vectors, how indexes find candidates, how metadata filters narrow the search, and how ranking choices affect what reaches the model.
From content to coordinates
An embedding model maps an input into a point in a high-dimensional space. Instead of two or three coordinates, a production embedding might have hundreds or thousands of dimensions.
You do not inspect each dimension directly. The vector is an internal representation learned from data. What matters is the pattern across all dimensions:
- Similar inputs should produce nearby vectors.
- Unrelated inputs should produce distant vectors.
- The same model should represent all items in a comparable way.
For example, a question about “resetting a password” should land closer to a help article about account recovery than to a pricing page, even if the exact words differ.
Similarity is a distance calculation
Once content is represented as vectors, search becomes a comparison problem. A system embeds the user’s query, compares that query vector to stored document vectors, and returns the closest matches.
Common similarity measures include:
- Cosine similarity, which compares direction and is common for text embeddings.
- Dot product, which can be efficient when vectors are unit-normalized or model-specific scoring expects it.
- Euclidean distance, which compares straight-line distance between points.
The metric must match the embedding model and vector database configuration. A strong embedding model can perform poorly if vectors are indexed with the wrong similarity measure.
Dot product and cosine similarity give the same ordering only when every stored vector and the query vector are normalized to length 1. If lengths carry signal, the two can rank results differently. The broader metric tradeoffs are covered in vector distance metrics.
Retrieval is a pipeline
A production retrieval system rarely does “embed query, return nearest vector” and stop. It usually combines several steps:
- Turn documents into chunks with metadata.
- Embed each chunk and store it in an index.
- At query time, retrieve candidates with vector search, keyword search, or both.
- Filter candidates by metadata such as tenant, language, product, or date.
- Rerank the best candidates with a stronger model or business rule.
- Send the top chunks into the LLM as context.
Why chunks matter
Most AI applications do not embed an entire knowledge base as one vector. They split content into chunks, embed each chunk, and search across those chunks.
Chunking controls what the vector represents:
- A chunk that is too small may lose context.
- A chunk that is too large may mix unrelated ideas.
- A chunk without useful metadata may be hard to explain or cite.
- Overlap can keep an answer that spans a boundary from being split apart, but too much overlap wastes storage and tokens.
- Headings, document titles, URLs, and timestamps often belong in metadata or in the chunk text so retrieved evidence is easy to explain.
Good chunks are usually centered on one coherent idea: a section, policy, answer, API endpoint, or troubleshooting step. The embedding should represent something specific enough to retrieve and broad enough to be useful.
Indexes trade speed for recall
A small system can compare the query vector with every stored vector. This is exact search, but it gets slow as the collection grows.
Most production systems use an approximate nearest neighbor index, or ANN index. “Approximate” means the index is designed to find very close candidates quickly, not prove it checked every possible vector. A common idea is an HNSW graph: each vector is connected to nearby vectors, and search walks the graph toward better matches instead of scanning the whole collection.
The tradeoff is speed versus recall. Higher recall means the index is more likely to include the truly relevant chunk in its candidate set. Faster settings may miss some good matches. Retrieval systems usually tune this with real queries, not just benchmark numbers.
Vectors are model-specific
Vectors from different embedding models should not be mixed in the same index unless the system explicitly supports that. Each model learns its own coordinate system, so vectors from different models live in different spaces and cannot be compared directly.
Changing the embedding model usually means re-embedding stored content. Otherwise, new query vectors may be compared against old document vectors that were created in a different space.
This is why production systems track:
- Which embedding model created each vector.
- The vector dimension count.
- The chunking strategy used.
- When the source content and vector were last updated.
Metadata filters, hybrid search, and reranking
Vector representation choices affect quality, cost, and reliability:
- Dimension size influences storage, memory, and search speed.
- Domain fit determines whether the model understands your vocabulary.
- Refresh strategy decides how quickly changed content appears in search.
- Filtering metadata helps combine semantic similarity with business rules such as tenant, language, date, or product area.
- Evaluation data shows whether retrieval improves for real user questions, not just demo prompts.
Metadata filters are useful, but they can also remove the answer. A filter like language = en is safe if the metadata is complete. A filter like product = billing can fail if the answer lives in a shared account article labeled differently.
Hybrid search combines keyword search with vector search. Keyword search is good at exact names, error codes, IDs, and rare terms. Vector search is good at paraphrases and related meaning. A reranker can then read the top candidates and reorder them with a more expensive model before the final chunks are sent to the LLM.
Treat embeddings as a search index with model behavior, not as a magic understanding layer. The vectors are powerful, but they only help if the content, chunking, metadata, model, and ranking strategy work together.
Evaluating retrieval
A simple starting metric is recall@k: for each test query, did the correct chunk appear in the top k retrieved results?
Three labeled test queries, k = 3:
Q1 correct chunk appears at rank 1 -> hit
Q2 correct chunk appears at rank 4 -> miss for recall@3
Q3 correct chunk appears at rank 2 -> hit
recall@3 = 2 hits / 3 queries = 0.67
Recall does not measure whether the final answer is well written. It measures whether the retriever gave the generator a fair chance. Pair it with answer-quality checks, citation checks, and latency and cost measurements.
Common failure modes
- Bad chunk boundaries: the answer is split across chunks, so no single vector represents it well.
- Stale index: source content changed, but old vectors are still being searched.
- Mixed embedding versions: queries and documents were embedded with different models or dimensions.
- Filters remove the answer: metadata rules are too strict or source metadata is incomplete.
- Keyword-only blind spots: exact search misses paraphrases.
- Vector-only blind spots: semantic search misses exact IDs, names, or error codes.
The key idea
Vector representations make semantic comparison practical. They let an AI system ask “which items mean something similar to this request?” instead of only “which strings contain the same words?”
That shift powers retrieval-augmented generation, recommendations, duplicate detection, clustering, and many other AI product features.