Skip to content
NakodaAI
← Research & Resources

Research paperGoogle Brain / Google Research2017-06-12

Attention Is All You Need

By Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin

The paper that introduced the Transformer architecture - a sequence model built entirely on attention mechanisms, with no recurrence or convolution - and became the architectural basis for essentially every large language model that followed it.

Why it matters

Every model discussed elsewhere in this library - GPT-4, Falcon, Jais, Claude - is a descendant of the architecture this paper introduced. It is the single most consequential piece of reading for understanding why current AI looks the way it does.

Key takeaways

  • Proposes the Transformer, dispensing entirely with the recurrent and convolutional layers that dominated sequence modelling before it.
  • Self-attention lets the model relate any two positions in a sequence directly, independent of distance - the property that made the architecture parallelisable and hugely trainable at scale.
  • Achieved a new state-of-the-art 28.4 BLEU on WMT 2014 English-to-German translation, and a new single-model best of 41.8 BLEU on English-to-French after only 3.5 days of training on eight GPUs.
  • Its parallelisability, more than its accuracy alone, is what made training the very large models of the years that followed computationally feasible.

Part of these reading paths

Related

Also worth reading

Academic resourceGaoling School of Artificial Intelligence, Renmin University of China (with Université de Montréal)2023-03-31

A Survey of Large Language Models

A continuously updated academic survey covering the background, key findings and mainstream techniques of large language models, organised around pre-training, adaptation tuning, utilisation and capability evaluation.

  • Research Foundations

Research paperTechnology Innovation Institute (TII), Abu Dhabi2023-11-28

The Falcon Series of Open Language Models

The technical paper behind Falcon-7B, -40B and -180B - causal decoder-only language models built by Abu Dhabi's Technology Innovation Institute and, at 180B parameters and 3.5 trillion training tokens, among the largest openly documented pretraining runs at time of release.

  • Research Foundations
  • UAE & MENA AI

Research paperInception (G42) / Cerebras Systems / Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)2023-08-30

Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

The paper introducing Jais, at the time the largest Arabic-centric open large language model, built with a bilingual (Arabic/English) training and vocabulary approach rather than adapting an English-first model after the fact.

  • Research Foundations
  • UAE & MENA AI