Skip to content
NakodaAI
← Research & Resources

Research paperOpenAI2023-03-15

GPT-4 Technical Report

By OpenAI

OpenAI's technical report on GPT-4, a large multimodal model accepting image and text input, covering its benchmark performance and a substantial section on the safety evaluation and mitigation work done before release.

Why it matters

It is deliberately light on architectural detail - OpenAI states this explicitly, citing competitive and safety considerations - which makes it a useful case study in how much a frontier lab will and will not publish, as much as a technical document about the model itself.

Key takeaways

  • GPT-4 accepts both image and text inputs and produces text outputs, and is reported to perform at or near human level on several professional and academic benchmarks, including a simulated bar exam.
  • The report explicitly withholds architecture, training compute and dataset details, citing the competitive landscape and safety implications of large-scale models.
  • Devotes a substantial section to safety: adversarial testing ("red teaming"), a model-generated risk card system, and reinforcement learning from human feedback used to improve factuality and reduce disallowed content.
  • Reports quantitative improvements in factuality and rule-adherence from post-training alignment relative to the pre-trained base model.

Part of these reading paths

Related

Also worth reading

Research paperUniversity of Washington / Black in AI / Independent2021-03-01

On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜

A widely cited critique of the trend toward ever-larger language models, examining environmental cost, documentation practices, bias amplification and the risk of models producing fluent but meaningless ("stochastic parrot") text mistaken for understanding.

  • AI Safety & Alignment
  • Research Foundations

Research paperTechnology Innovation Institute (TII), Abu Dhabi2023-11-28

The Falcon Series of Open Language Models

The technical paper behind Falcon-7B, -40B and -180B - causal decoder-only language models built by Abu Dhabi's Technology Innovation Institute and, at 180B parameters and 3.5 trillion training tokens, among the largest openly documented pretraining runs at time of release.

  • Research Foundations
  • UAE & MENA AI