All posts

#ai#embeddings#contrastive learning

The Search Box That Can't Find a "Phone Cover" When You Sell "iPhone Cases": Semantic Search With Contrastive Embeddings

Your customer types "phone cover" and your site says "no results", because the product is called a "case". How embedding models learn meaning through contrastive training, and semantic search with PostgreSQL + pgvector, step by step.

The Search Box That Can't Find a "Phone Cover" When You Sell "iPhone Cases": Semantic Search With Contrastive Embeddings
Contents 9

A customer types "phone cover" into an e-commerce site's search box. The site says "no results". Yet the warehouse holds hundreds of phone cases; none of them just happens to have "cover" in its name. The customer closes the tab and goes to a competitor.

Classic search engines match words. They don't know that "cover" and "case", "raincoat" and "waterproof jacket", or "a light dinner" and "salad" mean the same thing. Semantic search looks at meaning rather than words. What makes it possible is embedding models trained with contrastive learning.

In our contrastive language models post we covered the general idea, and in our CLIP post we looked at how images and text are matched. This post focuses on text-to-text search: how does a model learn "meaning", and how do you use it on your own site?

In short:

  • An embedding model turns each text into a list of numbers (a vector); texts with similar meaning end up with nearby vectors.
  • That closeness is learned through contrastive training: the model is forced to pick a query's correct answer out of a batch of wrong candidates.
  • What matters most for training quality is "hard negatives": wrong examples that look very much like the right answer.
  • With PostgreSQL's pgvector extension you can do semantic search without a separate database; the best results come from a hybrid that combines keyword and semantic search.

What is an embedding?

An embedding compresses the meaning of a text into a vector of hundreds of numbers. In the model we'll use here, each text becomes 768 numbers.

The individual numbers mean nothing on their own. What matters is where vectors sit relative to each other. In a well-trained model:

  • "phone cover" and "iPhone case" are very close,
  • "phone cover" and "phone charger" are at a middle distance,
  • "phone cover" and "food processor" are far apart.

Closeness is usually measured with cosine similarity: vectors pointing in the same direction have a similarity near 1.

How does a model learn "meaning"?

Nobody hands the model a thesaurus to teach it that "cover" and "case" are the same thing. It learns through a contrastive game.

Positive pairs

The raw material is pairs of related texts:

  • A question and its answer (e.g. a forum question and the accepted answer),
  • A search query and the page the user clicked,
  • A title and the body of the article,
  • Two phrasings of the same sentence.

These pairs are positive examples: the model should pull their vectors together.

In-batch negatives

So how do we tell the model what to push apart? Here's the elegant trick of contrastive learning: the other answers in the same training batch are natural negative examples for a query.

Picture a batch with 64 questions and 64 answers. For question 1, answer 1 is correct and the other 63 are wrong. For each question, the model tries to pick the correct answer among 64 candidates. It's like a multiple-choice exam, and the model is scored with a loss function called InfoNCE.

InfoNCE without the formula

Intuitively, InfoNCE does this:

  1. Compute the similarity between the query and every candidate.
  2. Turn those similarities into probabilities that sum to 1 (softmax).
  3. The lower the probability given to the correct answer, the bigger the penalty.

To reduce the penalty, the model learns to do two things at once: pull the correct answer toward the query and push the wrong ones away.

There's one more setting: temperature. Similarities are divided by it before being turned into probabilities. A low temperature makes the model more sensitive to small differences between the right answer and the closest wrong one, which makes temperature one of the most critical settings in contrastive training.

Hard negatives: where the real learning happens

Randomly chosen wrong answers are usually far too easy. Ruling out a "food processor" description for the query "iPhone 15 case" teaches the model nothing. It really learns from distinctions like these:

  • "iPhone 15 case" vs "iPhone 15 screen protector"
  • "iPhone 15 case" vs "iPhone 15 Pro case"
  • "women's running shoes" vs "men's running shoes"

Wrong examples that closely resemble the right answer are called hard negatives. Strong embedding models are trained with deliberately collected hard negatives, for example results a search engine ranked highly that weren't actually correct. When fine-tuning a model on your own data, well-chosen hard negatives usually bring the biggest gain.

Modern embedding models like E5 follow this pattern: they are first pre-trained contrastively on a huge amount of weakly labeled text pairs from the web, then fine-tuned on smaller, high-quality data with hard negatives. The multilingual-e5-base model we'll use belongs to this family and supports about 100 languages.

Hands-on: turn your products into vectors

First, install the library:

pip install sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("intfloat/multilingual-e5-base")

products = [
    "Silicone iPhone 15 case, shockproof, transparent",
    "Tempered glass iPhone 15 screen protector",
    "65W fast charging adapter, USB-C",
    "Waterproof women's raincoat with hood",
    "Stainless steel teapot, 2 liters",
]

# E5 models expect "passage: " before documents and "query: " before queries
product_vecs = model.encode(["passage: " + p for p in products], normalize_embeddings=True)

def search(query, k=3):
    q = model.encode(["query: " + query], normalize_embeddings=True)
    score = (q @ product_vecs.T)[0]
    for i in score.argsort()[::-1][:k]:
        print(f"{score[i]:.3f}  {products[i]}")

search("phone cover")
search("something to wear in the rain")

Two details to note:

  • Prefixes: the E5 family was trained with "query: " and "passage: " prefixes to tell queries and documents apart. Forgetting them noticeably hurts results. Always check the model card of the model you use for rules like this.
  • Normalization: once vectors are normalized, their product is the cosine similarity directly.

The query "phone cover" brings the case to the top even though its name has no "cover"; "something to wear in the rain" finds the raincoat. No word match, but a meaning match.

Scaling with PostgreSQL and pgvector

For a few hundred products, the code above is enough. For thousands or millions of records, you need to store vectors in a database. If you already run PostgreSQL, you don't need a separate vector database; the pgvector extension does the job:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE products (
    id          bigserial PRIMARY KEY,
    name        text NOT NULL,
    description text,
    embedding   vector(768)
);

-- HNSW index for approximate nearest neighbor search
CREATE INDEX ON products USING hnsw (embedding vector_cosine_ops);

Searching is just passing the query vector as a parameter and asking for the nearest rows:

SELECT name, 1 - (embedding <=> $1) AS similarity
FROM products
ORDER BY embedding <=> $1
LIMIT 10;

The <=> operator returns the cosine distance; subtracting it from 1 gives the similarity. The HNSW index keeps search in the millisecond range even across millions of vectors. The trade-off is that results are "approximate": occasionally one of the nearest rows can be missed. For most search scenarios that's more than acceptable. For how indexes work in general, see our database index post.

Hybrid search: the best of both worlds

Semantic search doesn't solve everything. Embedding models capture meaning well but can be weak on exact matches:

  • Product codes and SKUs ("SKU-48213"),
  • Model numbers ("WH-1000XM5"),
  • Rare brand names and proper nouns.

When a user types an exact model number, the product containing it must come first. Classic keyword search does that flawlessly.

The best results usually come from hybrid search: the same query runs through keyword search (such as PostgreSQL's full-text search) and semantic search, and the two result lists are merged. A simple, robust way to merge them is Reciprocal Rank Fusion (RRF): each result gets a score like 1 / (60 + rank) for its position in each list, and the scores are summed. Results ranked high in both lists rise to the top.

Fine-tuning on your own data

Ready-made multilingual models are enough for many jobs. But if your domain is very specific (legal texts, medical products, technical spare parts), fine-tuning on your own data can noticeably improve results.

The good news: the contrastive training described above ships as a ready-made loss in the sentence-transformers library: MultipleNegativesRankingLoss. It does InfoNCE-style training with in-batch negatives. All it needs from you is (query, correct result) pairs, which you can often derive from data you already have:

  • Site search logs: the query a user typed and the product they then bought,
  • Support tickets: the customer's question and the help article that solved it,
  • FAQ pages: question and answer.

Add hard negatives to those pairs (results your current search ranks high that are wrong), and the model learns exactly the distinctions your customers struggle with.

How do you measure results?

Don't ship semantic search because "I tried it and it seems fine". Build a simple evaluation set:

  1. Pick 50-100 queries from your real search logs.
  2. Mark the correct results for each query by hand.
  3. Run your old search and the new one on the same queries.
  4. Compare how often the correct result appears in the top 10 (recall@10).

This small set shows, by measurement rather than guesswork, which model, which prefixes and which hybrid setting really work better on your data.

Frequently asked questions

Does semantic search use an AI like ChatGPT?

No, it uses a much smaller and cheaper model. Embedding models don't generate text; they only turn text into vectors. A model like multilingual-e5-base can run on your own server, in many cases even without a GPU, and you don't need a paid API call for every search.

How does semantic search relate to RAG?

RAG (retrieval-augmented generation) grounds a language model's answer in your documents. The "retrieval" step usually uses exactly the semantic search described here: the document chunks closest to the user's question are found and handed to the model. It's one of the most effective ways to reduce AI hallucinations.

What do I do when product descriptions change?

Recompute and update the changed row's vector. If you switch models, though, you must regenerate all vectors; vectors from different models can't be compared with each other.

Which model should I choose for my language?

Multilingual models (e.g. the multilingual-e5 family) are a good starting point for most languages. The best way to choose is to compare a few models on your own data with a small evaluation set like the one above.

Semantic search is the most direct commercial payoff of contrastive learning: it closes the gap between your customers' words and your product names. If you want semantic search or a RAG-based assistant for your site or business documents, reach us through our enterprise software development page.

Sources

ShareLinkedInXWhatsApp
Need help with this?

If you would like to apply what this post covers to your own project, let’s look at it together.

Write to us
YE

Founder of EngerekTech. Builds web, mobile and enterprise software for businesses with Angular, Spring Boot and Flutter, and made the KPSS Düello and Kelime Kavanozu apps. On the blog he covers AI tools and software development as he uses them in his own projects.