CLIP (Contrastive Language-Image Pre-training)
/klɪp/
A multimodal model trained to understand relationships between images and text.
In italiano: CLIP (Contrastive Language-Image Pre-training)CLIP learns joint embeddings of images and text through contrastive learning on 400M image-text pairs. It enables zero-shot image classification and powers many vision-language applications.
Examples
- Zero-shot image classification
- Image search by text
- Text-to-image generation conditioning