CLIP (Contrastive Language-Image Pre-training)

/klɪp/

A multimodal model trained to understand relationships between images and text.

In italiano: CLIP (Contrastive Language-Image Pre-training)

CLIP learns joint embeddings of images and text through contrastive learning on 400M image-text pairs. It enables zero-shot image classification and powers many vision-language applications.

Examples

  • Zero-shot image classification
  • Image search by text
  • Text-to-image generation conditioning