Skip to content
NakodaAI

AI EncyclopediaArchitectures & models

Transformer (Architecture)

Also known as self-attention, attention mechanism

The neural network architecture, introduced in 2017, that processes a whole sequence at once using self-attention instead of reading it word by word - and that underlies almost every modern LLM.

architectureattentionneural network

In plain English

Before transformers, models read text roughly the way a person does - left to right, one word affecting the next. Self-attention lets a model instead look at every word in a sentence at the same time and weigh how much each one matters to every other one. That parallel, whole-sentence view is what made today's very large, very capable language models practical to train.

Technical explanation

Introduced in 'Attention Is All You Need' (Vaswani et al., Google, NeurIPS 2017), the transformer dispenses with recurrence and convolution entirely in favour of stacked self-attention and feed-forward layers, with positional encodings substituting for the sequence-order information recurrence used to provide. Because attention over a sequence is highly parallelisable on GPU/TPU hardware (unlike the sequential dependency of RNNs), transformers made training at the parameter and data scale LLMs now require computationally feasible.

Why it matters

Every major model family discussed in this encyclopedia - GPT, Claude, Gemini, Llama, Falcon - is transformer-based. It is the single architectural idea that made the current generation of generative AI possible, and 'Attention Is All You Need' is among the most cited papers in machine learning history.

Real-world example

BERT (2018) and the GPT series (2019 onward) were both built on the transformer architecture the 2017 paper introduced, establishing it as the default for large-scale language modelling within two years of publication.

Common misunderstanding

That 'transformer' names a specific product or company technology (confusion with the unrelated toy franchise is common in casual conversation). It is an open architecture, described in a public paper, that any lab can and does implement.

Something wrong here?

Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.