The Model That Thinks an Apple Is an iPod: Build Your Own Image Search With CLIP
CLIP is a contrastive model that lets you search photos with words, yet a handwritten note on an apple can fool it. How CLIP learns, how to search your photo archive in 30 lines of Python, and where it fails.

Contents 10
In 2021, researchers at OpenAI ran an interesting experiment. They showed a photo of a Granny Smith apple to a model called CLIP, and it guessed correctly: apple. Then they stuck a piece of paper reading "iPod" on the apple and asked again. Now the model was sure: it's an iPod.
The mistake looks funny, but it actually shows how powerful the model is. CLIP has learned to place images and text in the same "meaning space". Because it can read the text in a photo, it trusts that text too much. The same ability lets you type "a dog running on the beach" into your photo archive and find the right picture, without ever entering a tag.
In our contrastive language models post we covered this family in general. Here we go inside CLIP: how it learns, how to build your own image search engine, and where it goes wrong.
In short:
- CLIP was trained on 400 million image-text pairs from the internet by trying to work out which text belongs to which image.
- Because it places images and text in the same vector space, searching photos with a sentence comes down to measuring the similarity of two vectors.
- It can classify without ever seeing a labeled example (zero-shot); on ImageNet it matched a classic ResNet-50.
- Its weak spots are instructive too: it over-trusts text in a photo, carries the biases of internet data and is weaker in languages other than English.
What does CLIP learn?
Classic image classifiers are trained on a fixed list of labels: thousands of categories like "cat", "dog", "car". The model only recognizes those categories. To teach it something not on the list, you collect new labeled data and retrain.
CLIP (Contrastive Language-Image Pre-training) turned this around. Instead of a label list, it used the text that already sits next to images on the internet: captions, alt text, titles. The training data was 400 million image-text pairs.
The model has two parts:
- Image encoder: turns a photo into a list of numbers (a vector).
- Text encoder: turns a sentence into a vector of the same size.
The goal is to place the vectors of an image and its text close together, and those that don't belong together far apart.
Contrastive training: a giant matching game
Think of CLIP's training as a matching game. At each step the model gets a batch of images and their captions, but shuffled. Its task is to find each image's own caption.
Technically, it works like this:
- A batch holds N images and their N captions.
- The model computes a vector for every image and every text.
- It computes the similarity between every image and every text, producing an N×N table.
- The diagonal (each image paired with its own caption) holds the correct matches. The other N² − N cells are wrong matches.
- The model is updated to raise the similarity of correct matches and lower that of the wrong ones.
The subtle part: the model doesn't need separately collected "wrong" examples. The other captions in the same batch are natural wrong examples. The bigger the batch, the harder the game and the finer the distinctions the model learns. CLIP was trained with a batch size of 32,768, so each image had to find its own caption among more than 32,000 candidates.
Zero-shot: recognizing categories it never saw
CLIP's most surprising ability is zero-shot classification. You give the model a photo and a few candidate sentences, and it picks the one most similar to the photo.
To tell whether a photo shows a cat, a dog or a bird, you don't need a single labeled example. Give it "a photo of a cat", "a photo of a dog" and "a photo of a bird" and see which is closest. According to OpenAI, CLIP matched a ResNet-50 trained specifically on ImageNet without using any of ImageNet's 1.28 million labeled training examples.
An interesting detail: giving candidates as a sentence like "a photo of a cat" instead of a single word ("cat") improves results, because the model learned from captions where words usually appear inside sentences. According to the CLIP paper, this simple template alone raised ImageNet accuracy by about 1.3 points. Prompting mattered even before large language models.
Your own image search engine in 30 lines
Enough theory; let's write a small tool that searches a folder of photos with words. We'll use Hugging Face's transformers library:
pip install torch transformers pillow
from pathlib import Path
import torch
from PIL import Image
from transformers import CLIPModel, CLIPProcessor
MODEL = "openai/clip-vit-base-patch32"
model = CLIPModel.from_pretrained(MODEL)
processor = CLIPProcessor.from_pretrained(MODEL)
# 1. Turn the photos into vectors once
paths = sorted(Path("photos").glob("*.jpg"))
images = [Image.open(p).convert("RGB") for p in paths]
with torch.no_grad():
img = model.get_image_features(**processor(images=images, return_tensors="pt"))
img = img / img.norm(dim=-1, keepdim=True)
# 2. On each search, only embed the query and compare
def search(query, k=5):
with torch.no_grad():
txt = model.get_text_features(**processor(text=[query], return_tensors="pt", padding=True))
txt = txt / txt.norm(dim=-1, keepdim=True)
score = (img @ txt.T).squeeze(1)
for i in score.topk(min(k, len(paths))).indices:
print(f"{score[i]:.3f} {paths[i].name}")
search("a dog running on the beach")
search("a birthday cake with candles")
The code follows CLIP's training exactly:
- Photos are turned into vectors once. That's the expensive part, and for a large archive you do it once and store the result.
- On each search, only the query sentence is embedded. That's very fast.
- Since the vectors are normalized, their product gives the cosine similarity directly. The highest-scoring photos are the results.
With thousands of photos, process images in batches of 32 or 64 instead of all at once, and save the vectors to a file or a vector database. For searching millions of records, vector indexes (e.g. PostgreSQL's pgvector extension or FAISS) are used.
A zero-shot classification example
The same model can sort an unlabeled product photo into categories:
labels = ["a photo of a sneaker", "a photo of a handbag", "a photo of a wristwatch"]
image = Image.open("product.jpg").convert("RGB")
inputs = processor(text=labels, images=image, return_tensors="pt", padding=True)
with torch.no_grad():
probs = model(**inputs).logits_per_image.softmax(dim=1)[0]
for label, p in zip(labels, probs):
print(f"{p:.1%} {label}")
To change the category list you don't retrain anything; you just add a sentence.
Why did the apple become an iPod? Typographic attacks
Back to the opening experiment. When OpenAI researchers looked inside CLIP in 2021, they found "multimodal neurons" that respond to the same concept whether it appears as a photo or as text. One neuron, for example, responded to a photo of a spider, the word "spider" and drawings of Spider-Man.
That power of abstraction also created an attack surface. The researchers showed that a neuron responding to "finance" fired for both piggy banks and the string "$$$". Adding "$$$" to a photo of a dog could make the model classify the dog as a piggy bank. In OpenAI's words, the attack requires no more technology than pen and paper. These became known as typographic attacks.
The finding carries an important warning for real products: a CLIP-based system may trust text in a photo more than the visual content. A brand name written over a product photo, or a word added to fool a content moderation system, can change the result.
CLIP's limits
- Biases: the model carries the biases of internet data. OpenAI's own analysis found neurons associating certain regions or groups with negative concepts. Be careful in applications that classify people.
- Counting and fine distinctions: it's weak at telling "three apples" from "four apples", identifying specific car models or estimating distance.
- Language: the original CLIP was trained mostly on English text, and results drop for other languages. One fix is a multilingual text encoder aligned with CLIP's image encoder (e.g. sentence-transformers'
clip-ViT-B-32-multilingual-v1), which lets you search in Turkish, German and many other languages. - License and intended use: OpenAI's model card says the released CLIP models are intended for research. Check the model's license and terms before using it in a commercial product; open-source OpenCLIP models trained on LAION data are an alternative.
Where can businesses use CLIP?
- Visual search in e-commerce: when a customer types "red high-heeled shoes", the right products are found from their images even if those words aren't in the descriptions.
- Photo and media archives: agencies, newsrooms and real estate companies can make thousands of photos searchable without tagging them by hand.
- Duplicate and near-duplicate detection: finding very similar listing photos or copied product images.
- Automatic tagging: sorting incoming content into predefined categories zero-shot, then passing it to human review.
Frequently asked questions
Can CLIP generate images?
No. CLIP measures how well an image and a text match; it doesn't generate images. That measuring ability did, however, play an important role in building many image generation systems.
Do I need to train CLIP on my own data?
Usually not. Zero-shot search and classification with the ready-made model is enough for many jobs. If you need better results in a very specific domain (e.g. medical images or a particular product catalog), you can fine-tune with your own image-text pairs.
Do I need a GPU to run CLIP?
Not for small archives; small models like clip-vit-base-patch32 run on an ordinary computer's CPU. A GPU greatly speeds up embedding hundreds of thousands of photos, but that's done once; searches are fast either way.
How do I protect against typographic attacks?
Don't leave critical decisions to the CLIP score alone. Reading text in the photo separately (OCR), combining several models' results and adding human review for doubtful cases all reduce the risk.
CLIP is one of the most concrete examples of how powerful contrastive learning can be: just by asking "which text belongs to which image?" millions of times, it produced a model that joins seeing and reading in one space. If you want visual search and AI solutions for your products or archive, reach us through our web application development page.


