How It Works
When you select text containing image references, Local GPT:- Detects image syntax in your selection
- Extracts the images and converts them to base64
- Sends both text and images to a vision-capable model
- Receives AI-generated analysis or descriptions
Vision support automatically activates when your selection contains images and you have a vision provider configured.
Supported Image Formats
PNG
.png filesJPEG
.jpg and .jpeg filesSetup
1. Install a Vision Model
For Ollama users, install a vision-capable model:- bakllava (Recommended)
- llava
2. Configure Vision Provider
- Open Local GPT Settings
- Find Vision Provider
- Select your vision model provider from the dropdown
Image Reference Syntax
Use Obsidian’s standard image embedding syntax:Wiki-Style Embed
Multiple Images
Example Use Cases
Describe a Screenshot
Describe a Screenshot
Analyze a Diagram
Analyze a Diagram
Compare Images
Compare Images
Extract Text from Images
Extract Text from Images
Identify Objects
Identify Objects
How Images Are Processed
Local GPT converts images to base64-encoded data URLs for transmission:Provider Selection Logic
When images are detected, Local GPT automatically switches to your vision provider:If images are present in your selection, the vision provider takes precedence over your main provider.
Performance Considerations
Combining Vision with RAG
You can combine vision support with Enhanced Actions (RAG):- Process the image with the vision model
- Retrieve context from “Project Context” using RAG
- Generate a response informed by both the visual and textual context
Troubleshooting
Images not being processed
Images not being processed
- Verify your vision provider is configured in settings
- Check that images use the correct syntax:
![[image.png]] - Ensure image files exist in your vault
- Confirm image format is PNG or JPEG
Slow processing
Slow processing
- Vision models require more compute resources
- Consider using smaller/optimized models
- Reduce image file sizes
- Process fewer images at once
Provider errors
Provider errors
- Ensure your vision model is properly installed
- Check that the provider service is running
- Verify the model supports image inputs
Next Steps
Community Actions
Browse and install community-contributed actions
Enhanced Actions
Learn about RAG for context-aware responses