Skip to content
NakodaAI
← Research & Resources

Academic resourceGaoling School of Artificial Intelligence, Renmin University of China (with Université de Montréal)2023-03-31

A Survey of Large Language Models

By Wayne Xin Zhao, Kun Zhou, Junyi Li, and 19 further co-authors

A continuously updated academic survey covering the background, key findings and mainstream techniques of large language models, organised around pre-training, adaptation tuning, utilisation and capability evaluation.

Why it matters

It is the single most useful entry point for someone who wants the technical map of how LLMs are actually built - rather than a product announcement or a single model's paper - and it is kept current, with the authors publishing new revisions as the field moves.

Key takeaways

  • Structures the LLM lifecycle into four stages: pre-training, adaptation tuning (instruction and alignment tuning), utilisation (prompting and in-context learning), and capacity evaluation.
  • Maintained as a living document with an accompanying GitHub repository (RUCAIBox/LLMSurvey) tracking new models and techniques as they are published.
  • Surveys evaluation methodology as its own topic, not an afterthought - covering benchmark design and the limits of current evaluation practice.
  • Written as a reference text rather than a single research contribution, making it well suited to readers building foundational understanding rather than following one narrow subfield.

Part of these reading paths

Related

Also worth reading

Research paperGoogle Brain / Google Research2017-06-12

Attention Is All You Need

The paper that introduced the Transformer architecture - a sequence model built entirely on attention mechanisms, with no recurrence or convolution - and became the architectural basis for essentially every large language model that followed it.

  • Research Foundations

Research paperTechnology Innovation Institute (TII), Abu Dhabi2023-11-28

The Falcon Series of Open Language Models

The technical paper behind Falcon-7B, -40B and -180B - causal decoder-only language models built by Abu Dhabi's Technology Innovation Institute and, at 180B parameters and 3.5 trillion training tokens, among the largest openly documented pretraining runs at time of release.

  • Research Foundations
  • UAE & MENA AI

Research paperInception (G42) / Cerebras Systems / Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)2023-08-30

Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

The paper introducing Jais, at the time the largest Arabic-centric open large language model, built with a bilingual (Arabic/English) training and vocabulary approach rather than adapting an English-first model after the fact.

  • Research Foundations
  • UAE & MENA AI