零样图识

原名:clip

零样本图像分类和图文搜索。

分类
开发提效
版本
v1.0.0
作者
弈韬(@ra1nzzz)
下载
2
收藏
0
发布
2026-08-25
更新
2026-09-29
TRACE 评分
3.8 / 5

内容概览

OpenAI's model that understands images from natural language. Use when: - Zero-shot image classification (no training data needed) - Image-text similarity/matching - Semantic image search - Content moderation (detect NSFW, violence) - Visual question answering - Cross-modal retrieval (image→text, text→image) Metrics : - 25,300+ GitHub stars - Trained on 400M image-text pairs - Matches ResNet-50 on ImageNet (zero-shot) - MIT License Use alternatives instead : - BLIP-2 : Better captioning - LLaVA : Vision-language chat - Segment Anything : Image segmentation Model Parameters Speed Quality ------- ------------ ------- --------- RN50 102M Fast Good ViT-B/32 151M Medium Better ViT-L/14 428M Slow Best 1. Use ViT-B/32 for most cases - Good balance 2. Normalize embeddings - Required for cosine similarity 3. Batch processing - More efficient 4. Cache embeddings - Expensive to recompute 5. Use des…

查看 SKILL 详情