Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.
.claude/skills/openlair-llava/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 218% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 68% | 0% |
Open-source vision-language model for conversational image understanding.
Use when:
Metrics:
Use alternatives instead:
bash# Clone repository git clone https://github.com/haotian-liu/LLaVA cd LLaVA # Install pip install -e .
pythonfrom llava.model.builder import load_pretrained_model from llava.mm_utils import get_model_name_from_path, process_images, tokenizer_image_token from llava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN from llava.conversation import conv_templates from PIL import Image import torch # Load model model_path = "liuhaotian/llava-v1.5-7b" tokenizer, model, image_processor, context_len = load_pretrained_model( model_path=model_path, model_base=None, model_name=get_model_name_from_path(model_path) ) # Load image image = Image.open("image.jpg") image_tensor = process_images([image], image_processor, model.config) image_tensor = image_tensor.to(model.device, dtype=torch.float16) # Create conversation conv = conv_templates["llava_v1"].copy() conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?") conv.append_message(conv.roles[1], None) prompt = conv.get_prompt() # Generate response input_ids = tokenizer_image_token(prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors='pt').unsqueeze(0).to(model.device) with torch.inference_mode(): output_ids = model.generate( input_ids, images=image_tensor, do_sample=True, temperature=0.2, max_new_tokens=512 ) response = tokenizer.decode(output_ids[0], skip_special_tokens=True).strip() print(response)
| Model | Parameters | VRAM | Quality | |-------|------------|------|---------| | LLaVA-v1.5-7B | 7B | ~14 GB | Good | | LLaVA-v1.5-13B | 13B | ~28 GB | Better | | LLaVA-v1.6-34B | 34B | ~70 GB | Best |
python# Load different models model_7b = "liuhaotian/llava-v1.5-7b" model_13b = "liuhaotian/llava-v1.5-13b" model_34b = "liuhaotian/llava-v1.6-34b" # 4-bit quantization for lower VRAM load_4bit = True # Reduces VRAM by ~4×
bash# Single image query python -m llava.serve.cli \ --model-path liuhaotian/llava-v1.5-7b \ --image-file image.jpg \ --query "What is in this image?" # Multi-turn conversation python -m llava.serve.cli \ --model-path liuhaotian/llava-v1.5-7b \ --image-file image.jpg # Then type questions interactively
bash# Launch Gradio interface python -m llava.serve.gradio_web_server \ --model-path liuhaotian/llava-v1.5-7b \ --load-4bit # Optional: reduce VRAM # Access at http://localhost:7860
python# Initialize conversation conv = conv_templates["llava_v1"].copy() # Turn 1 conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?") conv.append_message(conv.roles[1], None) response1 = generate(conv, model, image) # "A dog playing in a park" # Turn 2 conv.messages[-1][1] = response1 # Add previous response conv.append_message(conv.roles[0], "What breed is the dog?") conv.append_message(conv.roles[1], None) response2 = generate(conv, model, image) # "Golden Retriever" # Turn 3 conv.messages[-1][1] = response2 conv.append_message(conv.roles[0], "What time of day is it?") conv.append_message(conv.roles[1], None) response3 = generate(conv, model, image)
pythonquestion = "Describe this image in detail." response = ask(model, image, question)
pythonquestion = "How many people are in the image?" response = ask(model, image, question)
pythonquestion = "List all the objects you can see in this image." response = ask(model, image, question)
pythonquestion = "What is happening in this scene?" response = ask(model, image, question)
pythonquestion = "What is the main topic of this document?" response = ask(model, document_image, question)
bash# Stage 1: Feature alignment (558K image-caption pairs) bash scripts/v1_5/pretrain.sh # Stage 2: Visual instruction tuning (150K instruction data) bash scripts/v1_5/finetune.sh
python# 4-bit quantization tokenizer, model, image_processor, context_len = load_pretrained_model( model_path="liuhaotian/llava-v1.5-13b", model_base=None, model_name=get_model_name_from_path("liuhaotian/llava-v1.5-13b"), load_4bit=True # Reduces VRAM ~4× ) # 8-bit quantization load_8bit=True # Reduces VRAM ~2×
| Model | VRAM (FP16) | VRAM (4-bit) | Speed (tokens/s) | |-------|-------------|--------------|------------------| | 7B | ~14 GB | ~4 GB | ~20 | | 13B | ~28 GB | ~8 GB | ~12 | | 34B | ~70 GB | ~18 GB | ~5 |
On A100 GPU
LLaVA achieves competitive scores on:
pythonfrom langchain.llms.base import LLM class LLaVALLM(LLM): def _call(self, prompt, stop=None): # Custom LLaVA inference return response llm = LLaVALLM()
pythonimport gradio as gr def chat(image, text, history): response = ask_llava(model, image, text) return response demo = gr.ChatInterface( chat, additional_inputs=[gr.Image(type="pil")], title="LLaVA Chat" ) demo.launch()
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 16,603 | 13,379 | -19% | 1 | 1 | 0% | 3,532 | 5,228 | +48% | 0 | 0 | — |
case-11 | fail→fail | 17,227 | 11,982 | -30% | 1 | 1 | 0% | 3,721 | 5,125 | +38% | 0 | 0 | — |
case-01 | fail→pass | 10,962 | 9,304 | -15% | 1 | 1 | 0% | 2,352 | 4,589 | +95% | 0 | 0 | — |
case-03 | pass→pass | 9,884 | 2,668 | -73% | 1 | 1 | 0% | 1,848 | 2,889 | +56% | 0 | 0 | — |
case-04 | fail→pass | 27,396 | 10,281 | -62% | 1 | 1 | 0% | 1,340 | 4,258 | +218% | 0 | 0 | — |
case-05 | fail→pass | 12,383 | 6,445 | -48% | 1 | 1 | 0% | 2,112 | 3,353 | +59% | 0 | 0 | — |
case-06 | fail→pass | 12,280 | 7,203 | -41% | 1 | 1 | 0% | 2,195 | 3,685 | +68% | 0 | 0 | — |
case-07 | pass→pass | 5,616 | 1,906 | -66% | 1 | 1 | 0% | 1,068 | 2,658 | +149% | 0 | 0 | — |
case-08 | pass→pass | 3,032 | 2,661 | -12% | 1 | 1 | 0% | 549 | 2,837 | +417% | 0 | 0 | — |
case-09 | pass→pass | 4,999 | 2,099 | -58% | 1 | 1 | 0% | 952 | 2,622 | +175% | 0 | 0 | — |
case-10 | pass→pass | 4,264 | 1,463 | -66% | 1 | 1 | 0% | 804 | 2,553 | +218% | 0 | 0 | — |
case-12 | pass→pass | 10,177 | 1,225 | -88% | 1 | 1 | 0% | 1,688 | 2,489 | +47% | 0 | 0 | — |
case-13 | pass→pass | 6,849 | 2,113 | -69% | 1 | 1 | 0% | 1,217 | 2,694 | +121% | 0 | 0 | — |
case-14 | fail→pass | 14,215 | 7,099 | -50% | 1 | 1 | 0% | 2,891 | 3,867 | +34% | 0 | 0 | — |
case-15 | pass→pass | 11,397 | 1,595 | -86% | 1 | 1 | 0% | 2,008 | 2,564 | +28% | 0 | 0 | — |
case-16 | fail→pass | 11,299 | 2,305 | -80% | 1 | 1 | 0% | 1,966 | 2,673 | +36% | 0 | 0 | — |
case-17 | pass→pass | 16,786 | 2,109 | -87% | 1 | 1 | 0% | 3,616 | 2,646 | -27% | 0 | 0 | — |
case-18 | pass→pass | 5,201 | 2,073 | -60% | 1 | 1 | 0% | 982 | 2,771 | +182% | 0 | 0 | — |
case-19 | pass→pass | 7,504 | 2,728 | -64% | 1 | 1 | 0% | 1,524 | 2,872 | +88% | 0 | 0 | — |
case-20 | pass→pass | 5,333 | 1,276 | -76% | 1 | 1 | 0% | 988 | 2,501 | +153% | 0 | 0 | — |
case-21 | pass→pass | 6,067 | 2,459 | -59% | 1 | 1 | 0% | 1,001 | 2,716 | +171% | 0 | 0 | — |
case-22 | pass→pass | 2,647 | 1,369 | -48% | 1 | 1 | 0% | 423 | 2,573 | +508% | 0 | 0 | — |
case-23 | pass→pass | 4,850 | 1,837 | -62% | 1 | 1 | 0% | 1,012 | 2,677 | +165% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +30 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.