What is a vector?
- A vector, in this context, is a numeric representation expressed as an ordered list (an array) of numbers—typically floating-point values.
- Vectors (often called embeddings) encode properties of data such as the semantic meaning of text, visual features of an image, or other learned attributes produced by machine learning models.
- Dimensionality is the number of elements in the vector. Higher dimensionality can capture more nuance, but it increases storage, compute, and indexing complexity.
What is a database?
- A database is a system to store, organize, and manage data.
- A vector database specifically stores vectors (and usually associated metadata or payload) and provides efficient mechanisms to index, search, and retrieve them based on similarity rather than exact matches.

Why vectors matter in ML and AI
Vectors are the common numeric format that models use to represent data—text, images, audio, or structured records. Once data is converted into vectors (embeddings), you can:- Measure similarity between items using distance metrics (cosine similarity, Euclidean distance, etc.)
- Retrieve nearest neighbors for search and recommendation
- Cluster related items or perform semantic grouping
- Power retrieval-augmented generation (RAG) and contextual retrieval for LLMs
Embeddings are the numeric outputs of models that capture semantic or feature-level information. You typically do not compare raw text or images directly; you compare their vector embeddings to find related content.
Typical operations and examples
Common indexing algorithms (ANN)
- HNSW (Hierarchical Navigable Small World): high recall and fast queries for many real-world workloads.
- IVF (Inverted File) + PQ (Product Quantization): good for very large collections to reduce memory usage.
- Annoy (Approximate Nearest Neighbors Oh Yeah): disk-backed, memory-efficient, optimized for read-heavy workloads.
- Faiss (Facebook AI Similarity Search): a library that implements many ANN techniques and optimizations.
Where vector databases are used
- Semantic search: return results based on meaning rather than keyword matches.
- Recommendation engines: find similar items or users in embedding space.
- Retrieval-Augmented Generation (RAG): fetch context for language models to ground outputs.
- Deduplication & clustering: detect near-duplicates or group similar records.
- Multimodal retrieval: search across text, images, and other modalities using unified embeddings.
High-dimensional vectors can be expensive to store and index. Choosing the right index type and dimensionality (embedding size) is critical for balancing recall, latency, and cost.
Pros and cons (summary)
Helpful links and references
- Kubernetes Basics (general infrastructure)
- Faiss: https://github.com/facebookresearch/faiss
- HNSW paper & implementations: https://arxiv.org/abs/1603.09320
- Annoy: https://github.com/spotify/annoy