Awesome Testing

Models · reviewed · reviewed Oct 6, 2026 · 4 min

How do LoRA and adapters change a model?

LoRA freezes selected base weight matrices and learns a correction through two smaller factors. The layer adds that correction to the base transformation. This reduces the trainable parameter budget when the chosen rank is small; the base model still has to be stored and executed.

Keep the base, learn a correction

A neural-network layer can transform an input vector by multiplying it by a weight matrix. Full fine-tuning can update that matrix directly. Low-Rank Adaptation, or LoRA, instead freezes the base matrix and learns an additional transformation through two smaller matrices.

Both paths receive the same input. Their outputs are added: base transformation + learned correction. The surrounding model continues with that combined representation. The adapter is therefore part of the computation, not a document retrieved into the prompt or a separate assistant consulted afterward.

“Adapter” is a broader term for a trainable addition to a model. Here it means the LoRA factors. Other adapter designs insert different modules; the LoRA multiplication rule does not describe every method called an adapter.

Two small factors can change a larger matrix

Let the base matrix be four by four. A rank-one LoRA path maps four input coordinates into one intermediate coordinate, then maps that coordinate into four output coordinates. Its two factors contain eight values in total. Their product describes a four-by-four correction with a restricted structure.

Rank two supplies two intermediate coordinates and sixteen factor values. It can express additional directions of change, but in this tiny example its factor count already equals the sixteen values of a full matrix update. The useful saving depends on dimensions and rank, not the name LoRA.

The experiment uses authored factors instead of training them. Change the rank, input, and correction gain. Watch the correction matrix and the combined output; predict what setting the gain to zero does to the base matrix.

Explore the mechanism

Change the correction, keep the base

Frozen base W₀
1.000.200.000.00
0.001.000.200.00
0.000.001.000.20
0.200.000.001.00
Correction: gain × BA
0.400.00-0.400.00
0.200.00-0.200.00
-0.400.000.400.00
0.100.00-0.100.00

Input vector: [1.00, 0.00, 0.00, 1.00]. Both paths receive this same input.

OutputBaseCorrectionCombined
11.000.401.40
20.000.200.20
30.20-0.40-0.20
41.200.101.30
Base parameters (frozen)
16
Adapter parameters
8
Full-matrix update budget
16

The output changes through the correction; the base matrix stays fixed.

Inspect factors and the merged matrix
A: 1 × 4
1.000.00-1.000.00
B: 4 × 1
0.40
0.20
-0.40
0.10
Merged W₀ + gain × BA
1.400.20-0.400.00
0.201.000.000.00
-0.400.001.400.20
0.300.00-0.101.00

Merged-path output: [1.40, 0.20, -0.20, 1.30]. It equals the separate-path sum in this numerical example.

The base matrix, factors, and inputs are authored teaching data. Products, sums, and parameter counts are computed locally. Gain changes a forward calculation; no optimizer fits these factors, no LLM runs, and the counts do not measure full-model memory or quality.

Zero gain removes the correction’s contribution. The base output remains intact. A different input can reveal a different effect from the same correction because matrix multiplication depends on the input vector as well as the weights.

The numbers illustrate a forward computation. Moving the gain slider is not gradient descent, and increasing rank here is not evidence of improved task quality. Real training learns the factors from examples and a loss; the widget uses fixed values so the relationship remains inspectable.

What the parameter saving actually buys

For an output-by-input matrix with dimensions d × k, a rank-r LoRA pair contains r(d + k) parameters, compared with dk entries in a full update. At sufficiently small rank, fewer parameters need gradients and optimizer state. A task-specific adapter can also be stored separately from a shared base.

The base weights still exist. The forward calculation still uses them, and training may still need activation storage and gradients through the frozen network to reach the adapter. A small adapter file is not the memory footprint of the complete system.

QLoRA combines trainable low-rank adapters with a frozen quantized base. It distinguishes the base’s storage representation from the precision used for computation. This connects adaptation with quantization without making them interchangeable techniques.

Merging changes the release artifact

For the simple linear path, the correction can be added to the base matrix ahead of inference. The optional inspector compares the merged matrix output with the separate-path sum: they match in this numerical example.

That does not make an arbitrary adapter compatible with an arbitrary checkpoint. Matrix dimensions, targeted layers, base identity, scaling, tokenizer, and serving configuration belong to the artifact’s specification. Matching shapes alone does not establish matching learned meaning.

Keep the original base identifiable when producing a merged artifact. Numerical precision, quantization, and runtime support can affect the practical result, so verify the deployed representation rather than assuming that a symbolic equality resolves every compatibility issue.

Rank is a budget choice, not a quality dial

The rank bounds the structure of each learned correction. It does not prescribe which layers need adaptation or how much budget every layer deserves. AdaLoRA studies adaptive allocation across weight matrices rather than assigning one identical budget everywhere.

Imagine adapting a support model for one product’s terminology. A compact correction may learn the common phrasing while still failing an exceptional rule. A larger rank can represent more directions of change, but the examples, objective, targeted layers, and independent evaluation determine whether those directions are useful.

LoRA reduces the scope of parameter updates. It does not update frequently changing facts automatically, validate an answer, or grant permission to execute an action.

Inspect the factorization used here

The frozen base is W₀. Factor A has shape r × 4; factor B has shape 4 × r. The authored first direction uses A [1, 0, −1, 0] and B column [0.4, 0.2, −0.4, 0.1]. Rank two adds A [0, 1, 0, −1] and B column [0.1, −0.3, 0.2, 0.4].

The widget computes output = W₀x + gain × BAx. The gain is a directly controlled scale, not a fitted parameter or learning rate. Real LoRA recipes often use a scale such as α/r; this example isolates the effective scale so changing rank does not silently change the gain as well.

Input A is [1, 0, 0, 1]; input B is [0, 1, 1, 0]. The complete matrices and both output paths are available in the experiment’s inspector. Ordinary JavaScript numbers carry the calculation; no low-bit checkpoint is stored.

Sources and further reading

  1. 01
    LoRA: Low-Rank Adaptation of Large Language ModelsHu et al. · research · published Jun 17, 2021 · source checked Oct 6, 2026

    A primary parameter-efficient fine-tuning method that freezes base weights and trains low-rank adapter matrices.

  2. 02
    QLoRA: Efficient Finetuning of Quantized LLMsDettmers et al. · research · published May 23, 2023 · source checked Oct 6, 2026

    Frozen quantized base weights with trainable low-rank adapters separate storage precision from adaptation and computation precision.

  3. 03
    AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-TuningZhang et al. · research · published Mar 18, 2023 · source checked Oct 6, 2026

    Adaptive rank allocation across weight matrices shows why a single uniform adapter budget is not a universal quality prescription.