慧眼识言

原名:llava

视觉语言对话:视觉问答、图像描述和图像对话。

分类
开发提效
版本
v1.0.0
作者
弈韬(@ra1nzzz)
下载
1
收藏
0
发布
2026-08-25
更新
2026-10-02
TRACE 评分
3.8 / 5

内容概览

Open-source vision-language model for conversational image understanding. Use when: - Building vision-language chatbots - Visual question answering (VQA) - Image description and captioning - Multi-turn image conversations - Visual instruction following - Document understanding with images Metrics : - 23,000+ GitHub stars - GPT-4V level capabilities (targeted) - Apache 2.0 License - Multiple model sizes (7B-34B params) Use alternatives instead : - GPT-4V : Highest quality, API-based - CLIP : Simple zero-shot classification - BLIP-2 : Better for captioning only - Flamingo : Research, not open-source Model Parameters VRAM Quality ------- ------------ ------ --------- LLaVA-v1.5-7B 7B 14 GB Good LLaVA-v1.5-13B 13B 28 GB Better LLaVA-v1.6-34B 34B 70 GB Best 1. Start with 7B model - Good quality, manageable VRAM 2. Use 4-bit quantization - Reduces VRAM significantly 3. GPU required - CPU infer…

查看 SKILL 详情