Language becomes a prediction problem
A large language model processes text as a sequence of tokens and predicts possible continuations. Tokens may correspond to words, pieces of words, punctuation, or other units. A tokenizer turns text into integer identifiers from a fixed vocabulary. The same visible word can use different numbers of tokens across tokenizers and languages.
For a decoder-only model, the basic training task is next-token prediction. Given “The capital of Nepal is,” the model learns to assign probabilities to possible next tokens. Repeated over vast text collections, this task can produce useful language and problem-solving behavior. The training objective is still predictive; it does not require every statement to be true.
Embeddings turn identifiers into vectors
Token identifiers are mapped to vectors called embeddings. These vectors are learned parameters, and subsequent layers transform them into representations that depend on surrounding context.
Position information is also needed because word order matters. Transformer architectures use positional encodings or related mechanisms such as rotary position embeddings. Without a way to represent order, attention alone would not distinguish many sequences that contain the same tokens in different positions.
A rough pipeline is:
Text → Tokens → Embeddings + position information
→ Transformer layers → Output scores
→ Probability distribution → Next token
Attention connects relevant positions
Self-attention lets a token's representation draw information from other token positions. Consider “Mina put the book on the table because it was heavy.” To interpret “it,” a model may need information from the earlier noun and the surrounding sentence.
Attention computes query, key, and value vectors from the current representations. Similarity between a query and keys determines how strongly to weight the corresponding values. A common scaled dot-product formulation is:
attention(Q, K, V) = softmax(QKᵀ / √d) V
Here d is the key dimension. Multiple attention heads provide several learned projections. The model is not literally consulting a human-readable index of concepts; it performs numerical operations whose useful patterns emerge through training.
In a causal decoder, a mask prevents a position from looking at future tokens during training. Otherwise, the model could trivially see the target it is meant to predict. Encoder-style models may use different masking patterns and support different tasks.
A transformer layer contains more than attention
After attention, a feed-forward network transforms each position's representation. Residual connections and normalization help information flow through deep networks and make training more stable. Many layers stack these operations, gradually building context-sensitive representations.
At the output, a projection produces a score for each vocabulary token. A softmax converts scores into a probability distribution. During training, loss compares that distribution with the actual next token, and backpropagation adjusts the parameters.
Modern models vary in architecture, activation functions, normalization, sparsity, and positional methods. “Transformer” describes a family of designs rather than one unchanging recipe.
Pretraining is only part of assistant behavior
Pretraining learns broad patterns from data. Supervised instruction tuning can teach the model to follow examples of useful responses. Preference optimization can further shape which responses it tends to produce. These stages can improve helpfulness and behavior, while also introducing tradeoffs that must be evaluated.
Tool use adds another system layer. A model may emit a structured request to a search tool, calculator, or database. The application executes that request, returns results, and lets the model continue. The tool's correctness and authorization are responsibilities of the application, not properties guaranteed by the model.
How generation works
The model predicts a next token, selects one according to a decoding rule, appends it to the context, and repeats. Greedy decoding chooses a highest-probability token. Sampling can introduce variation; temperature changes the sharpness of the probability distribution, and top-p sampling limits choices to a probability mass.
A low temperature does not establish truth. It can make a wrong response more consistent. Likewise, a detailed explanation can sound authoritative even when the underlying continuation is unsupported.
Generation can reuse cached keys and values from earlier positions rather than recomputing all previous attention work. This KV cache reduces repeated computation but consumes memory. Longer contexts and larger batches increase memory pressure. Traditional full attention also has costs that grow quadratically with sequence length, although optimized implementations and alternative architectures change practical behavior.
Limits that matter in an application
A context window limits what the model can consider in a request. Having information somewhere within that window does not guarantee it will be used correctly. Prompts, relevant sources, conversation history, and requested output all compete for space.
Models can hallucinate facts, struggle with precise counting, reproduce biases, or become vulnerable to malicious instructions in external content. Evaluate representative tasks rather than relying only on benchmark scores or a few impressive examples. Use retrieval for current evidence, deterministic tools for exact calculations, structured output validation, and human escalation where error costs are high.
A transformer is a powerful learned sequence processor. A dependable assistant requires a surrounding design that grounds claims, limits privileges, measures failure, and makes uncertainty visible when it affects a user's decision.