Interactive Demo

Contrastive Learning

Explore how neural networks learn by comparing positive and negative pairs in embedding space.

What is augmentation?

Augmentation creates two different views of the same image. The model learns that both views should have similar embeddings — forcing it to understand the image's true content, not surface details.

Original Image
View 1
Augmented View
No augmentation

Choose Augmentations

Key Insight

Both views come from the same image → they form a positive pair. The model is trained to produce similar embedding vectors for both, no matter how different they look visually.

What is the Encoder?

The encoder is a neural network (like ResNet) that converts a raw image into a small list of numbers called an embedding vector. Similar images get similar vectors. Different images get different vectors.

🖼️
Raw Image
224×224×3
150,528 numbers
→
🧠
Encoder f(·)
ResNet-50
learns features
→
📊
Embedding h
2048 numbers
compact description
Interactive — Watch pixels compress into an embedding

The encoder takes raw pixel values and compresses them into a small meaningful vector. Click Randomize to see different images.

Input Pixels (12 sample values)
↓   Encoder f(·)   ↓
Embedding h (6 values)
What the encoder learns — layer by layer
Early Layers
Edges & Corners
Detects basic shapes — lines, curves, colour boundaries. Universal low-level features.
Middle Layers
Shapes & Parts
Combines edges into parts — eyes, ears, wheels, legs. Recognises components.
Deep Layers
Semantic Concepts
Assembles parts into whole concepts — "this is a dog", "this is a car".
Output h
2048-dim Vector
Rich description used for all tasks — classification, retrieval, clustering.
Key Insight

The encoder has shared weights — the same network processes both augmented views of an image. This forces it to map the same object to the same region of space, no matter what augmentation was applied.

What is Similarity?

After the encoder produces embedding vectors, we need to measure how similar two vectors are. We use cosine similarity — it measures the angle between two vectors. The closer the angle to 0°, the more similar.

Cosine Similarity Formula
sim(h₁, h₂) = (h₁ · h₂) / (‖h₁‖ × ‖h₂‖)
h₁ · h₂ = dot product (multiply each pair of numbers and add up)  |  ‖h‖ = length of the vector  |  Result is always between −1 and +1
Similarity Scale
+1.0
Identical
Same dog, two augmentations. Vectors point in exactly the same direction.
0.0
Unrelated
Completely different concepts. Vectors are perpendicular — 90° apart.
−1.0
Opposite
Maximally different. Vectors point in completely opposite directions.
Interactive Similarity Calculator

Enter two simple 3-dimensional vectors and see the cosine similarity computed live.

VECTOR h₁
VECTOR h₂
Similarity
Why cosine similarity and not just subtraction?

Subtracting two vectors gives a different answer depending on how long the vectors are. A dog described in "loud" numbers vs "quiet" numbers would seem different even if it's the same dog. Cosine similarity only looks at the direction, not the magnitude — so the same meaning always gives the same score.

Contrastive Loss

The loss function rewards the model for pulling positive pairs close and pushing negative pairs apart. Select a pair type below to see how loss behaves.

🐕🐕
Same dog, two augmentations
Positive Pair ✅
🐕🐈
Dog vs Cat
Negative Pair ✅
🐕🐕
Two different dogs
False Negative ⚠️
Similarity
0.92
Distance
0.08
Loss
0.04
✅ Good pair! High similarity → low loss. The model is rewarded for correctly recognizing these as the same image.