4 October 2026
Transformer (Architecture)
The neural network architecture, introduced in 2017, that processes a whole sequence at once using self-attention instead of reading it word by word - and that underlies almost every modern LLM.
In plain English
Before transformers, models read text roughly the way a person does - left to right, one word affecting the next. Self-attention lets a model instead look at every word in a sentence at the same time and weigh how much each one matters to every other one. That parallel, whole-sentence view is what made today's very large, very capable language models practical to train.
Technical explanation
Introduced in 'Attention Is All You Need' (Vaswani et al., Google, NeurIPS 2017), the transformer dispenses with recurrence and convolution entirely in favour of stacked self-attention and feed-forward layers, with positional encodings substituting for the sequence-order information recurrence used to provide. Because attention over a sequence is highly parallelisable on GPU/TPU hardware (unlike the sequential dependency of RNNs), transformers made training at the parameter and data scale LLMs now require computationally feasible.
Why today
Every major model family discussed in this encyclopedia - GPT, Claude, Gemini, Llama, Falcon - is transformer-based. It is the single architectural idea that made the current generation of generative AI possible, and 'Attention Is All You Need' is among the most cited papers in machine learning history.
Real-world example
BERT (2018) and the GPT series (2019 onward) were both built on the transformer architecture the 2017 paper introduced, establishing it as the default for large-scale language modelling within two years of publication.

