Awesome Testing

Models · reviewed · reviewed Oct 5, 2026 · 4 min

How does an LLM use images?

A multimodal system encodes image regions into learned representations and connects them to a language model alongside text. The model generates an answer from the supplied representations; cropping, resolution, encoding, and interpretation can all lose or distort the evidence.

Pixels need a numerical interface too

A text tokenizer maps strings to vocabulary IDs. Images enter through a different encoding path. A vision encoder transforms pixels or image regions into learned vectors; a connector lets a language model use those representations together with text.

A common vision-transformer design divides an image into patches, projects them into vectors, and mixes information across the sequence. A patch is a region of pixels, not a recognized object or a word. Useful visual features emerge from learned transformations and context.

flowchart LR
  I[Image and chosen preprocessing] --> V[Vision encoder]
  V --> C[Learned connection to language model]
  T[Text question] --> L[Language model]
  C --> L
  L --> A[Generated answer]

The original LLaVA architecture is one concrete example: a pretrained visual encoder supplies features, and a trainable projection connects them to the language model's embedding space. Other architectures use different connectors and attention paths. “Multimodal” describes supported input types, not one universal implementation.

Follow one receipt through a missing region

A fictional receipt shows quantity 2, unit price 20, and shipping 5. The question asks for the total including shipping. All three values matter: 2 × 20 + 5 = 45.

Crop away the shipping region or cover it while preserving the rest of the image. An optional 3×3 grid illustrates patch boundaries. The observation panel contains authored field visibility, not the output of OCR or a vision model. The arithmetic uses only the available observations.

Explore the mechanism

Change the image, change the evidence

Question: What is the total including shipping?

ORDER 17 · EXAMPLEQuantity2Unit price20Shipping5
Grid borders show regions, not recognized objects or text tokens.
Illustrated patches
9
Authored observations
Quantity: 2
Unit price: 20
Shipping: 5
RegionsVisual featuresLanguage model input
Supported fixture total: 45

All required fields are present in this authored observation.

Native illustration of a synthetic receipt; field visibility is authored from the chosen scenario. No image encoder, OCR, or LLM runs. The 3×3 grid and patch counts illustrate preprocessing only; they are not real visual-token counts.

Cropping leaves fewer illustrated regions. Covering retains the image's illustrated patch count while removing useful information. Neither operation supplies the missing shipping amount. A hypothetical zero-shipping total is 40, but it does not answer the original question with evidence.

Real systems can also misread a value that remains visible. The experiment isolates availability, not measured visual recognition. Its nine-region grid is not a provider's tokenization scheme or a visual-token billing estimate.

Alignment is learned, not a dictionary lookup

An image feature is not automatically a sentence. The system needs training that makes visual representations useful to its language behaviour.

CLIP learns image and text representations using paired examples and a contrastive objective: corresponding images and descriptions should align better than mismatched pairs. That supports representation learning and comparison. CLIP alone is not a conversational assistant.

A vision-language assistant adds a connection to a language model and training for its intended interaction. It can answer questions, describe a scene, or propose an action, but those outputs remain generated claims. A feature vector does not carry a guaranteed label such as “the shipping cost is exactly five.”

Availability and interpretation fail separately

Start diagnosis at the first lost boundary:

BoundaryExample failureUseful evidence
PreprocessingA crop excludes the relevant footerExact supplied image and crop
RepresentationSmall print or spatial detail is poorly preservedResolution and encoder configuration
InterpretationA visible digit or relationship is read incorrectlyIndependent inspection of the region
Answer or actionCorrect observations become wrong arithmetic or an unsafe clickComputation, target state, and policy trace

A longer text prompt cannot restore pixels that were never supplied. A sharper image may improve available evidence, but it does not establish correct counting or exact interpretation. If the question needs a hidden region, obtain another image or ask for the missing fact.

Seeing a control does not authorize its use

A screenshot may contain private information, misleading instructions, or a visible destructive button. Image content is evidence with provenance; it does not become trusted policy because the model can read it.

For browser work, combine the representation suited to the task: accessibility trees expose roles and states, screenshots expose appearance and spatial context, and application state establishes effects. A model's proposed click still crosses the harness permission boundary.

For exact document work, prefer authoritative structured data where available. Keep original images when authorized, link extracted fields to their regions, check calculations deterministically, and preserve missing or ambiguous values rather than forcing a plausible answer.

Keep a visual request reproducible

Record the original authorized image, preprocessing, crop and resize settings, image ordering, text question, model/encoder revision where available, and the resulting observations. A thumbnail or a later screenshot can change the input being compared.

Use frozen fixtures with missing regions, occluded fields, small text, conflicting labels, and questions that require spatial relationships. Compare actual task outcomes and unsupported claims. A correct-looking explanation is insufficient evidence that the model used the relevant pixels.

Sources and further reading

  1. 01
    An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleDosovitskiy et al. · research · published Oct 22, 2020 · source checked Oct 5, 2026

    Primary description of patch projection, sequence representations, and positional information in a vision Transformer; patches are not recognized objects or words.

  2. 02
    Learning Transferable Visual Models From Natural Language SupervisionRadford et al. · research · published Feb 26, 2021 · source checked Oct 5, 2026

    Primary evidence for contrastive image/text representation learning; useful for separating learned alignment from a complete conversational vision-language assistant.

  3. 03
    Visual Instruction TuningLiu et al. · research · published Apr 17, 2023 · source checked Oct 5, 2026

    Primary LLaVA architecture and training account connecting a visual encoder through a learned projection to an LLM; this is one design, not a universal multimodal interface.