← Back to concepts
9 min read

Parameters and weights in AI models

A machine learning model is a large mathematical function. A large language model is one example. It takes an input, performs many calculations, and produces an output. The values that control those calculations are called parameters. In neural networks, the largest group of parameters is usually weights. A bias is a learned offset added before an activation, which lets a layer shift its output up or down.

You can think of weights as adjustable settings inside the model. They are not rules that a programmer writes one by one. Training changes them so the model becomes better at the task it is learning.

How weights are learned

Training usually starts with weights that contain random or unhelpful values. During pre-training or fine-tuning, the model repeats a learning loop:

  1. The model receives an example, such as a sequence of tokens.
  2. It makes a prediction using its current weights.
  3. A loss function measures how far the prediction is from the expected result.
  4. Backpropagation computes gradients: signals that show which direction to nudge each weight to reduce the loss.
  5. An optimizer applies tiny updates using a step size called the learning rate.
The training loop that adjusts weights A cycle surrounds the model's weights. First the model receives an example with its expected result. Second it makes a prediction using the current weights. Third a loss function measures how wrong the prediction was. Fourth backpropagation computes gradients and an optimizer nudges the weights to reduce the loss. The loop then repeats with the next example. 1. Example input tokens + expected result 2. Prediction computed with current weights 3. Loss how wrong was the prediction? 4. Gradient update backprop computes direction optimizer nudges weights repeat with the next example Weights the model's parameters used by changes Repeated across many examples, small updates add up to learned patterns.
Training is a feedback loop: each prediction error is turned into a small change to the weights.

This process is repeated across many examples. Backpropagation does not try weights one by one. It computes gradients for all weights efficiently, then the optimizer applies small steps. Over time, those small steps add up to patterns that help the model predict, classify, or transform new inputs.

For a language model, the examples may teach patterns such as which tokens tend to occur together, how code is structured, or how an answer usually follows an instruction. The model does not store these patterns as a simple list of facts. They are distributed across many parameters and interact during each calculation.

A small example

Imagine a simple model that predicts whether a message is about a password problem. It might learn that words such as forgot, reset, and login are useful signals.

At first, the weights connected to those words may be poor, so the model makes many wrong predictions. After seeing labeled examples, training increases or decreases those weights as needed:

Message: "I forgot my password"
Expected label: password problem
Weights before and after training in a tiny password classifier A tiny model scores the message I forgot my password by adding the weights of feature words that appear in it. Only forgot appears. Before training the weights are small and arbitrary: forgot 0.10, reset minus 0.20, login 0.05, pricing 0.30. The score of 0.10 is below the 0.5 threshold, so the model wrongly predicts not a password problem. After training, forgot is 1.60, reset 1.40, login 0.90, and pricing minus 1.20. The score of 1.60 is above the threshold, so the model correctly predicts a password problem. Message: "I forgot my password" (of these feature words, only "forgot" appears) Before training forgot in message reset login pricing 0.10 -0.20 0.05 0.30 Score = weight of "forgot" = 0.10 Below 0.5: "not password" (wrong) After training on labeled examples forgot in message reset login pricing 1.60 1.40 0.90 -1.20 Score = weight of "forgot" = 1.60 Above 0.5: "password problem" (correct) Illustrative weights for a tiny linear model. Bars right of the line are positive; bars left are negative.
Training does not write a rule; it shifts numbers until useful signals such as "forgot" push the score in the right direction.

The real weights in a modern AI model are far more numerous and the relationships are much more complex. The example is only meant to show that training changes numerical values based on errors; it does not manually write a rule such as “if the message contains forgot, return password problem.”

Parameters, weights, and layers

In everyday discussions, parameters and weights are often used almost interchangeably. More precisely, parameters include every learned value in the model, while weights are the learned values that multiply or connect signals in its layers.

Neural networks are organized into layers. Each layer transforms the values from the previous layer, and its weights determine how strongly different signals influence that transformation. A model with billions of parameters has billions of learned values spread across these layers.

Counting the parameters in a tiny neural network A network has an input layer with 3 nodes, a hidden layer with 4 nodes, and an output layer with 2 nodes. Every input connects to every hidden node, giving 12 weights, and every hidden node connects to every output node, giving 8 weights. Each hidden and output node also has one bias, giving 6 biases. In total the network has 26 parameters. Thicker lines represent larger illustrative weights. INPUT LAYER HIDDEN LAYER OUTPUT LAYER +b +b +b +b +b +b 3 x 4 = 12 weights 4 x 2 = 8 weights 20 weights + 6 biases = 26 parameters Each line is one weight (thicker = larger, illustrative). Each +b is one bias.
Every connection and bias is one learned number; a production model repeats this pattern until the count reaches billions.

The number of parameters is a rough measure of model capacity, not a complete quality score. In a mixture-of-experts model, there may also be a difference between total parameters and active parameters: the model stores many expert weights, but only routes each token through some of them. Two models with similar parameter counts can behave very differently because of differences in:

  • Training data and data quality.
  • Model architecture and tokenizer.
  • Training objective and optimization process.
  • Fine-tuning and alignment.
  • The task and domain where the model is used.

Precision, memory, and quantization

A parameter count also helps estimate memory. The rough rule is:

model memory ~= number of parameters x bytes per parameter

For an illustrative 7B-parameter model, 16-bit weights use about 14 GB just for the weights. The same weights in 8-bit form use about 7 GB. In 4-bit form, they use about 3.5 GB. Real serving also needs extra memory for activations, runtime overhead, and the KV cache used during inference, so hardware planning needs a safety margin.

Quantization means storing weights with fewer bits. It can make an open-weight model cheaper or easier to run, and it can improve speed on hardware that supports that format. The tradeoff is that very aggressive quantization can reduce quality, especially on tasks that need careful reasoning or exact formatting.

Illustrative weight memory for a 7B parameter model A bar chart compares memory for only the weights of an illustrative seven billion parameter model. Sixteen-bit weights use about fourteen gigabytes, eight-bit weights use about seven gigabytes, and four-bit weights use about three point five gigabytes. A note says serving needs extra memory for activations, runtime overhead, and KV cache. Illustrative 7B model: weight memory only Extra memory is still needed for activations, runtime overhead, and KV cache. 16-bit weights ~14 GB 7B x 2 bytes = about 14 GB 8-bit weights ~7 GB 7B x 1 byte = about 7 GB 4-bit weights ~3.5 GB 7B x 0.5 bytes = about 3.5 GB
Lower precision reduces weight memory, but the full serving footprint is larger than the weights alone.

Parameters are not the same as context

When you call a trained model, its parameters are normally fixed. Your prompt does not rewrite the weights. Instead, the prompt supplies temporary context that changes the output for that request.

For example, a model may know a general pattern for writing a summary. A prompt can ask it to summarize a particular document, but that document is not permanently added to the model’s parameters.

This distinction helps explain several common techniques:

  • Prompting supplies instructions and examples for one request.
  • Retrieval-augmented generation supplies relevant documents in the context without changing the model.
  • Fine-tuning performs additional training that changes all weights or trains small add-on adapters, depending on the method.
Temporary context versus permanent parameter changes Two separate requests send their own context into the same model. Request 1 carries a prompt plus retrieved documents, and request 2 carries a different prompt. The model parameters stay fixed while serving both requests, and each request's context is not written into the weights. Apps or serving systems may keep history, memory, or caches separately. Separately, fine-tuning is offline additional training that writes new weights or adapters, which is a lasting change. Request 1 context prompt + retrieved docs (RAG) Request 2 context a different prompt Model parameters fixed while serving same for every request Answer 1 not saved in weights Answer 2 not saved in weights Fine-tuning (offline) additional training writes weights/adapters TEMPORARY context for one request LASTING lasting model change
Prompts and retrieved documents shape an answer but are not written into the weights; apps may store history outside the model.

What parameter count means in practice

More parameters can give a model room to represent more complex patterns, but bigger is not automatically better. Larger models often require more memory, compute, and time to run. They may also be more expensive to train and serve.

When choosing a model, compare quality on your actual task with:

  • Latency: how quickly it responds.
  • Cost: how much inference and storage use.
  • Memory: whether it fits your hardware or hosting limits.
  • Reliability: how consistently it handles real inputs.
  • Maintenance: how easy it is to update, evaluate, and operate.

Some techniques reduce the resources needed to store or update parameters. Quantization stores weights with fewer bits, which can reduce memory use. Parameter-efficient fine-tuning updates a smaller set of added parameters instead of changing the whole base model. These methods can improve practical deployment, but they do not remove the need for evaluation.

The key idea

Parameters are the learned numerical values inside a model. Training adjusts them from examples so the model becomes better at its task. A prompt can influence a model’s behavior temporarily, while fine-tuning changes learned weights or adds trained adapters. Understanding that difference helps you choose between prompting, retrieval, fine-tuning, and a different model.