零样图识
原名:clip
零样本图像分类和图文搜索。
- 分类
- 开发提效
- 版本
- v1.0.0
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 2
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-29
- TRACE 评分
- 3.8 / 5
内容概览
OpenAI's model that understands images from natural language. Use when: - Zero-shot image classification (no training data needed) - Image-text similarity/matching - Semantic image search - Content moderation (detect NSFW, violence) - Visual question answering - Cross-modal retrieval (image→text, text→image) Metrics : - 25,300+ GitHub stars - Trained on 400M image-text pairs - Matches ResNet-50 on ImageNet (zero-shot) - MIT License Use alternatives instead : - BLIP-2 : Better captioning - LLaVA : Vision-language chat - Segment Anything : Image segmentation Model Parameters Speed Quality ------- ------------ ------- --------- RN50 102M Fast Good ViT-B/32 151M Medium Better ViT-L/14 428M Slow Best 1. Use ViT-B/32 for most cases - Good balance 2. Normalize embeddings - Required for cosine similarity 3. Batch processing - More efficient 4. Cache embeddings - Expensive to recompute 5. Use des…