慧眼识言
原名:llava
视觉语言对话:视觉问答、图像描述和图像对话。
- 分类
- 开发提效
- 版本
- v1.0.0
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-10-02
- TRACE 评分
- 3.8 / 5
内容概览
Open-source vision-language model for conversational image understanding. Use when: - Building vision-language chatbots - Visual question answering (VQA) - Image description and captioning - Multi-turn image conversations - Visual instruction following - Document understanding with images Metrics : - 23,000+ GitHub stars - GPT-4V level capabilities (targeted) - Apache 2.0 License - Multiple model sizes (7B-34B params) Use alternatives instead : - GPT-4V : Highest quality, API-based - CLIP : Simple zero-shot classification - BLIP-2 : Better for captioning only - Flamingo : Research, not open-source Model Parameters VRAM Quality ------- ------------ ------ --------- LLaVA-v1.5-7B 7B 14 GB Good LLaVA-v1.5-13B 13B 28 GB Better LLaVA-v1.6-34B 34B 70 GB Best 1. Start with 7B model - Good quality, manageable VRAM 2. Use 4-bit quantization - Reduces VRAM significantly 3. GPU required - CPU infer…