Dev & EngARTICLE

Behind Every Prompt There Is More Math Than Magic

Behind Every Prompt There Is More Math Than Magic
Image: Cezar Taurion

I've seen a lot of people wanting to work with LLMs starting directly with prompts, RAG, agents, frameworks and APIs. All of this is useful. But if the intention is to understand technically what happens inside a model, there's a layer before that that's hard to avoid: math.

This doesn't mean you need to be a mathematician. The level needed depends on what you intend to do. But some concepts help a lot to stop seeing the LLM as a black box.

I would start with linear algebra, which I consider the most important foundation. A token is initially a discrete unit associated with an index. This index is used to select an embedding vector, which provides an initial representation of that token. As it passes through the Transformer's layers, this representation is continuously transformed and becomes dependent on context.

It's important to understand what "dimension" means here. We're not necessarily talking about directly interpretable features like "color", "size" or "sentiment". They are coordinates learned by the model in high-dimensional spaces.

That's why vectors and matrices appear everywhere. Much of the internal representations can be described by vectors, while many of the learned transformations involve multiplications by matrices.

The dot product is particularly important because it produces a number from two vectors and appears in the calculation of attention. In the attention mechanism, it helps produce a learned compatibility or relevance score between representations. Norms measure magnitude, and concepts from high-dimensional geometry help explain why relationships between vectors are so important.

This appears clearly in the expression Attention(Q,K,V) = softmax(QKᵀ/√dₖ)V. Q stands for Query, K for Key, and V for Value. They are obtained through learned projections of the representations arriving at the layer. A useful, though simplified, interpretation is to think of Q as what a position is searching for, K as the features used to determine the relevance of the other positions, and V as the information that will be weighted and combined.

The product between Q and K produces the attention scores. Dividing by √dₖ helps control their scale. Softmax turns the scores into normalized weights, and these weights determine how the V values will be combined. This is only part of the full attention mechanism, but it already shows how much linear algebra exists inside a single operation.

The second foundation is probability and statistics. An autoregressive LLM receives a context and produces a distribution over the possible next tokens. If I write "The capital of France is", the model doesn't need to retrieve a stored sentence. It calculates scores, called logits, for the possible next tokens.

Softmax turns these logits into a probability distribution. Depending on the tokenization, "Paris" may correspond to one or more tokens, but the corresponding continuation tends to receive a high probability in a model that has properly learned this regularity.

This is where conditional probability starts to make sense. The model is essentially estimating something like P(next token | previous tokens). Once a token is selected, it becomes part of the context and the process repeats. This way, a sequence can be produced token by token.

This also helps explain temperature, top-k and top-p. Temperature modifies the distribution used in decoding. Top-k restricts the candidates to the k highest-scoring tokens. Top-p selects the smallest set of the most probable tokens whose accumulated probability mass reaches or exceeds a given threshold p. These are decoding mechanisms, not different forms of the model's knowledge.

Exponentials and logarithms also appear here. Softmax uses exponentials to turn logits into normalized positive values. Log-probabilities are particularly useful when dealing with sequences, since they turn products of probabilities into sums and help avoid numerical problems associated with multiplying many small values.

The third foundation is differential calculus. During training, we need to determine how changes in the parameters would affect the loss function. That's where derivatives and gradients come in.

A derivative answers, intuitively: if I change this value a little, how much does the result change? Since the loss function depends on an enormous number of parameters, its partial derivatives indicate its local sensitivity with respect to each of them. The set of these derivatives forms the gradient. The chain rule allows these relationships to be propagated through the network's successive operations. Backpropagation is the efficient procedure used to calculate these gradients across the computational graph.

This brings us to the fourth foundation: optimization. Backpropagation calculates the gradients. The optimizer determines how to use them to update the parameters. Gradient descent seeks to reduce the loss function by making updates in the direction opposite to the gradient. The learning rate controls the size of these updates. Too small can make learning slow; too large can produce instability.

Instead of calculating the gradient over the entire training set before each update, we normally use mini-batches. The gradient calculated over each mini-batch works as an estimate of the gradient of the objective defined over the training distribution.

Techniques like momentum incorporate information from previous updates. Optimizers like Adam maintain estimates of the moments of the gradients and use them to adapt the parameter updates.

The well-known metaphor of a ball rolling down a mountain helps, but it has limits. An LLM with billions of parameters creates an optimization problem in an extremely high-dimensional parameter space, with flat regions, different curvatures and saddle points, points where curvature can have different signs depending on the direction considered.

And the goal isn't simply to get the lowest possible error on the training data. We also want generalization: for the model to perform adequately when faced with combinations and situations that didn't appear exactly in that form during training.

The fifth foundation is information theory. Entropy can be understood, simplifying, as a measure of uncertainty associated with a distribution. If practically all the probability is concentrated on one possibility, there is little uncertainty. If it's spread across many possibilities, there is more.

Cross-entropy is central to autoregressive pretraining. Instead of talking about the "correct token", it's more rigorous to talk about the token observed in the training data.

For that token, the contribution to the loss is essentially Loss = −log P(observed token | context). If the model assigns a high probability to the observed token, the loss is small. If it assigns a very low probability, the penalty is large.

Training can therefore be seen as the process of adjusting the parameters to increase, on average, the probability assigned to the tokens observed in the data. And one important clarification is worth making: cross-entropy is not, technically, a distance metric.

Finally, we enter an area that already brings math closer to engineering: numerical methods and computational representation. All this math needs to run on GPUs and other accelerators, and numbers need to be represented with a finite number of bits.

FP32, FP16, BF16 and FP8 use different formats and offer different trade-offs between precision, numerical range, memory and performance. Fewer bits doesn't simply mean "lower quality". BF16, for example, has different precision and numerical range characteristics than FP16.

Quantization goes further, representing weights and, depending on the technique, activations or other values with fewer bits. This can significantly reduce memory consumption and, under certain conditions, speed up inference. But lower numerical precision doesn't automatically guarantee lower latency: the result also depends on hardware, kernels, architecture and implementation.

Overflow, underflow, numerical stability and mixed precision matter because a mathematically correct equation can run into problems when executed with finite numerical representations.

This is where math meets engineering. At bottom, these areas tell different parts of the same story. Linear algebra helps explain how representations are built and transformed. Probability describes how the model distributes its predictions. Calculus shows how we measure the sensitivity of the loss to the parameters. Optimization determines how these parameters are updated. Information theory helps formulate training objectives. Numerical methods determine how all this math can be executed efficiently on real hardware.

But there's still an important caveat. Understanding this math doesn't mean fully understanding why certain capabilities appear in the representations learned by the model. Knowing how to calculate attention or the gradient doesn't mean knowing how to causally interpret everything a Transformer has learned.

That's exactly where areas like mechanistic interpretability, activation analysis, circuits and causal interventions come in.

We don't need to master all of this to use an LLM, build a RAG, or create an agent. But there's a huge difference between knowing how to operate the technology and understanding the principles that make its functioning possible.

The more we want to move from the first dimension to the second, the more math stops being just complementary knowledge and starts being part of the language necessary to understand what really happens inside these models.

Translated from the Brazilian Portuguese original · Read the original

More from Cezar Taurion
View profile →
↳

Threads