Explore how neural networks learn by comparing positive and negative pairs in embedding space.
Augmentation creates two different views of the same image. The model learns that both views should have similar embeddings — forcing it to understand the image's true content, not surface details.
Both views come from the same image → they form a positive pair. The model is trained to produce similar embedding vectors for both, no matter how different they look visually.
The encoder is a neural network (like ResNet) that converts a raw image into a small list of numbers called an embedding vector. Similar images get similar vectors. Different images get different vectors.
The encoder takes raw pixel values and compresses them into a small meaningful vector. Click Randomize to see different images.
The encoder has shared weights — the same network processes both augmented views of an image. This forces it to map the same object to the same region of space, no matter what augmentation was applied.
After the encoder produces embedding vectors, we need to measure how similar two vectors are. We use cosine similarity — it measures the angle between two vectors. The closer the angle to 0°, the more similar.
Enter two simple 3-dimensional vectors and see the cosine similarity computed live.
Subtracting two vectors gives a different answer depending on how long the vectors are. A dog described in "loud" numbers vs "quiet" numbers would seem different even if it's the same dog. Cosine similarity only looks at the direction, not the magnitude — so the same meaning always gives the same score.
The loss function rewards the model for pulling positive pairs close and pushing negative pairs apart. Select a pair type below to see how loss behaves.