Models · reviewed · reviewed Oct 5, 2026 · 4 min
How does an LLM use images?
A multimodal system encodes image regions into learned representations and connects them to a language model alongside text. The model generates an answer from the supplied representations; cropping, resolution, encoding, and interpretation can all lose or distort the evidence.
Pixels need a numerical interface too
A text tokenizer maps strings to vocabulary IDs. Images enter through a different encoding path. A vision encoder transforms pixels or image regions into learned vectors; a connector lets a language model use those representations together with text.
A common vision-transformer design divides an image into patches, projects them into vectors, and mixes information across the sequence. A patch is a region of pixels, not a recognized object or a word. Useful visual features emerge from learned transformations and context.
flowchart LR I[Image and chosen preprocessing] --> V[Vision encoder] V --> C[Learned connection to language model] T[Text question] --> L[Language model] C --> L L --> A[Generated answer]
The original LLaVA architecture is one concrete example: a pretrained visual encoder supplies features, and a trainable projection connects them to the language model's embedding space. Other architectures use different connectors and attention paths. “Multimodal” describes supported input types, not one universal implementation.
Follow one receipt through a missing region
A fictional receipt shows quantity 2, unit price 20, and shipping 5. The question asks for the total including shipping. All three values matter: 2 × 20 + 5 = 45.
Crop away the shipping region or cover it while preserving the rest of the image. An optional 3×3 grid illustrates patch boundaries. The observation panel contains authored field visibility, not the output of OCR or a vision model. The arithmetic uses only the available observations.
Explore the mechanism
Change the image, change the evidence
Question: What is the total including shipping?
- Illustrated patches
- 9
- Authored observations
- Quantity: 2
Unit price: 20
Shipping: 5
All required fields are present in this authored observation.
Native illustration of a synthetic receipt; field visibility is authored from the chosen scenario. No image encoder, OCR, or LLM runs. The 3×3 grid and patch counts illustrate preprocessing only; they are not real visual-token counts.
Cropping leaves fewer illustrated regions. Covering retains the image's illustrated patch count while removing useful information. Neither operation supplies the missing shipping amount. A hypothetical zero-shipping total is 40, but it does not answer the original question with evidence.
Real systems can also misread a value that remains visible. The experiment isolates availability, not measured visual recognition. Its nine-region grid is not a provider's tokenization scheme or a visual-token billing estimate.
Alignment is learned, not a dictionary lookup
An image feature is not automatically a sentence. The system needs training that makes visual representations useful to its language behaviour.
CLIP learns image and text representations using paired examples and a contrastive objective: corresponding images and descriptions should align better than mismatched pairs. That supports representation learning and comparison. CLIP alone is not a conversational assistant.
A vision-language assistant adds a connection to a language model and training for its intended interaction. It can answer questions, describe a scene, or propose an action, but those outputs remain generated claims. A feature vector does not carry a guaranteed label such as “the shipping cost is exactly five.”
Availability and interpretation fail separately
Start diagnosis at the first lost boundary:
| Boundary | Example failure | Useful evidence |
|---|---|---|
| Preprocessing | A crop excludes the relevant footer | Exact supplied image and crop |
| Representation | Small print or spatial detail is poorly preserved | Resolution and encoder configuration |
| Interpretation | A visible digit or relationship is read incorrectly | Independent inspection of the region |
| Answer or action | Correct observations become wrong arithmetic or an unsafe click | Computation, target state, and policy trace |
A longer text prompt cannot restore pixels that were never supplied. A sharper image may improve available evidence, but it does not establish correct counting or exact interpretation. If the question needs a hidden region, obtain another image or ask for the missing fact.
Seeing a control does not authorize its use
A screenshot may contain private information, misleading instructions, or a visible destructive button. Image content is evidence with provenance; it does not become trusted policy because the model can read it.
For browser work, combine the representation suited to the task: accessibility trees expose roles and states, screenshots expose appearance and spatial context, and application state establishes effects. A model's proposed click still crosses the harness permission boundary.
For exact document work, prefer authoritative structured data where available. Keep original images when authorized, link extracted fields to their regions, check calculations deterministically, and preserve missing or ambiguous values rather than forcing a plausible answer.
Keep a visual request reproducible
Record the original authorized image, preprocessing, crop and resize settings, image ordering, text question, model/encoder revision where available, and the resulting observations. A thumbnail or a later screenshot can change the input being compared.
Use frozen fixtures with missing regions, occluded fields, small text, conflicting labels, and questions that require spatial relationships. Compare actual task outcomes and unsupported claims. A correct-looking explanation is insufficient evidence that the model used the relevant pixels.
Sources
Sources and further reading
- 01An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleDosovitskiy et al. · research · published Oct 22, 2020 · source checked Oct 5, 2026
Primary description of patch projection, sequence representations, and positional information in a vision Transformer; patches are not recognized objects or words.
- 02Learning Transferable Visual Models From Natural Language SupervisionRadford et al. · research · published Feb 26, 2021 · source checked Oct 5, 2026
Primary evidence for contrastive image/text representation learning; useful for separating learned alignment from a complete conversational vision-language assistant.
- 03Visual Instruction TuningLiu et al. · research · published Apr 17, 2023 · source checked Oct 5, 2026
Primary LLaVA architecture and training account connecting a visual encoder through a learned projection to an LLM; this is one design, not a universal multimodal interface.
