From tokens to one vector: the bi-encoder
A BGE model is a bi-encoder (also called a dual encoder). It runs the transformer once over a piece of text and collapses the resulting sequence of token vectors into a single embedding. Concretely, input text is tokenized into n tokens, embedded, and passed through the encoder to produce a matrix H: [n, d] of contextual hidden states, where d is the hidden size (768 for BGE-base, 1024 for BGE-large).
The crucial move is pooling: reducing H: [n, d] to a single v: [d]. The word ‘bi’ matters because a query and a document are encoded independently — each becomes its own vector with no cross-attention between them. That independence is what makes BGE usable at scale: embed a million documents once, store the vectors, and at query time only encode the query and compare. A cross-encoder, which feeds query and document together, cannot be precomputed — a distinction we return to with rerankers.
Pooling: CLS versus mean
There are two standard ways to pool H: [n, d] into one vector. CLS pooling takes the hidden state of the special [CLS] token prepended to every input: v = H[0]. The model is trained so this one position aggregates the whole sequence’s meaning. BGE uses CLS pooling.
Mean pooling averages the token states, usually masking out padding: v = (Σ_i m_i · H[i]) / Σ_i m_i, where m_i ∈ {0,1} is the attention mask. Neither is universally best; what matters is that training and inference use the same pooling. A mismatch — training with CLS but mean-pooling at inference — reads vectors from a space the loss never shaped, and retrieval quality collapses. Always pool a BGE model the way its model card specifies.
Normalization and cosine similarity
BGE embeddings are compared by cosine similarity, the cosine of the angle between two vectors: cos(a, b) = (a · b) / (‖a‖ ‖b‖). Cosine ignores magnitude and measures only direction, which is what we want — a long document and a short query about the same topic should score high regardless of vector length.
The standard trick is to L2-normalize every embedding to unit length: v̂ = v / ‖v‖. After normalization ‖v̂‖ = 1, so cosine similarity reduces to a plain dot product: cos(â, b̂) = â · b̂. This is not cosmetic. It means a vector database can use fast inner-product search over normalized vectors and get cosine ranking for free, since for unit vectors ‖â − b̂‖² = 2 − 2(â · b̂) — smaller distance means larger cosine. Normalize once at index time and once per query; everything downstream is a dot product.
InfoNCE: the contrastive loss, written out
Here is the heart of it. BGE is trained on pairs: a query q and a positive passage p+ that matches it. The goal is to arrange the space so sim(q, p+) is high while sim(q, p−) for every negative is low. This is contrastive: the model learns no absolute score, only that the positive should beat the negatives — which reframes embedding as picking the right passage from a set, letting us reuse softmax cross-entropy.
The objective is InfoNCE. For a query q with positive p+ and negatives {p−_1 … p−_k}, scale each pair’s similarity by a temperature τ; the loss is the negative log-probability of picking the positive:
s_j = sim(q, p_j) / τ # one logit per candidate
L = −log( exp(s_+) / Σ_j exp(s_j) )
= −s_+ + log Σ_j exp(s_j) # softmax cross-entropy, label = the positiveThe denominator runs over the positive and all negatives. Minimizing L pushes s_+ up and every other s_j down — exactly the ‘pull together, push apart’ behavior we wanted, expressed as one differentiable quantity whose gradients flow back through the encoder to reshape the space.
Temperature: sharpening the contrast
The temperature τ (BGE uses small values, around 0.01–0.05) scales the logits before the softmax and controls how peaky the distribution is. Dividing by a small τ spreads the similarities into a wide range, so the softmax becomes sharp: it heavily rewards getting the single hardest negative right and penalizes any near-miss.
Think of it as a magnifying glass on the gap between the positive and the best negative. A large τ flattens the softmax and yields fuzzy embeddings; a tiny τ forces crisp separation but can destabilize training if negatives are noisy, since a mislabeled ‘negative’ that is actually relevant now dominates the loss. Sharp softmax plus clean, hard negatives is what produces the tight, discriminative spaces BGE is known for.
Negatives: in-batch and hard
Where do the negatives come from? The cheapest source is in-batch negatives. In a batch of B query–positive pairs (q_i, p_i), the other B−1 passages serve as negatives for query i at no extra encoding cost. Encode queries into Q: [B, d] and passages into P: [B, d]; one matrix multiply S = Q Pᵀ gives a [B, B] similarity matrix whose diagonal S[i, i] holds the positives, so InfoNCE is just cross-entropy with labels [0, 1, …, B−1]. Hence larger batches mean more negatives per step — which is why embedding training scales batch size into the thousands.
But in-batch negatives are usually easy: a random passage is obviously unrelated, so the model beats it early. To draw fine distinctions you add hard negatives — passages that look relevant but are wrong — mined by retrieving top candidates with an existing model and taking high-ranked non-answers. Their gradients carry the most information because they are exactly where the model is currently wrong. The risk is false negatives: a mined ‘negative’ that is actually a valid answer teaches the model to push apart things that should be close, so careful mining is much of what separates a strong embedding model from a mediocre one.
BGE-M3: three retrieval modes at once
BGE-M3 extends the family so one forward pass yields three kinds of representation. Dense is the classic single CLS-pooled vector scored by dot product — everything above. Sparse (lexical) assigns a learned weight to each vocabulary token; its score is the sum of matched-term weights, like a learned BM25, recovering the exact-keyword matching that dense vectors blur away.
Multi-vector (ColBERT-style) keeps one vector per token rather than pooling, and scores a pair by Σ_i max_j (q_i · d_j) — the ‘late interaction’ MaxSim: for each query token, take its best-matching document token and sum. This is more expressive than a single dot product but far cheaper than a full cross-encoder. In practice M3 lets you retrieve with dense + sparse for recall and rerank with multi-vector from one model, trained on a mix of the three objectives.
Rerankers: when a bi-encoder is not enough
BGE also ships rerankers, and these are cross-encoders — a different computation. A reranker takes the concatenation [query, document] as one input, runs the transformer over both together so every query token attends to every document token, and outputs a single relevance score. Modeling that full interaction makes it more accurate than the bi-encoder’s dot product of two independent vectors, but nothing can be precomputed: scoring N documents costs N full passes — hopeless for a large corpus. The standard architecture is two stages: the fast bi-encoder retrieves the top 100 candidates from millions, then the slow, accurate reranker scores just those 100. Recall from the cheap model, precision from the expensive one.