By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres walks non-specialists through how large language models like ChatGPT, Claude, and Gemini turn a prompt into a response, without assuming any machine learning background. She organizes the explanation around input, the black box, and output: the input is a system prompt plus user prompt plus history, converted into tokens and then embedding vectors placed in a high-dimensional space. Inside the black box she traces neural networks, the limits of processing words in isolation, and the attention mechanism with query, key, and value vectors that gives each token context, combined into transformer blocks. The output stage predicts the next token from the last token's enriched embedding and repeats. The piece matters because a working mental model of these mechanics helps product people reason about strengths, weaknesses, and behaviors like context rot.
01Key takeaways
- The input you see is only part of what the model receives; system prompts and conversation history are always prepended, which explains output differences across tools.
- Tokens let a limited vocabulary handle any word, misspelling, or new term by combining subword pieces.
- Context windows are measured in tokens, so they set how much input a model can process at once.
- Attention lets each token absorb context from earlier tokens, which is why the same word can mean different things depending on surrounding words.
- Generation is iterative: each predicted token is added to the input and the full process runs again, so understanding this helps explain model behavior and limits.
02Key sections
- The Input
- The prompt sent to a model is actually a system prompt plus conversation history plus the user's message, which explains why the same prompt can yield different outputs across tools. Text is then split into tokens and converted into embedding vectors.
- Tokens and Embeddings
- Models read subword tokens rather than whole words, which lets a modest vocabulary represent nearly any text. Each token maps to a point in a huge space where similar tokens sit closer together, learned through training.
- Neural Networks and Training
- A neural network multiplies inputs by weights, sums them, and applies an activation function, with weights learning meaning by repeatedly being nudged toward less wrong answers. Its limitation is that each input is treated in isolation, so 'great' and 'not great' look the same.
- Attention and Transformer Blocks
- The attention layer uses query, key, and value vectors to weight how relevant each preceding token is, reshaping each token's embedding to carry its context. Attention layers paired with neural network layers form transformer blocks, stacked dozens of times in modern models.
- The Output
- Only the last token's enriched embedding is used to predict the next token, by checking which vocabulary embeddings point in a similar direction. The predicted token is appended to the input and the whole process repeats for each generated token.
03From the post
“Everybody is buzzing about generative AI and there's a lot of jargon that comes along with it: tokens, embeddings, attention, neural networks, transformers. Whether we are talking about ChatGPT, Claude, or Gemini, the large language models all work the same way. But understanding how they work can be a”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.