AI EncyclopediaUAE & regional
Falcon (LLM family, TII)
A family of open-weight large language models developed by Abu Dhabi's Technology Innovation Institute (TII), released from 2023 onward, including one of the largest openly available LLMs at the time of its release.
In plain English
Falcon is the UAE's own entry in the race to build large language models - and, unlike many competitors, TII released the model's weights openly, meaning anyone can download, inspect, run and build on it rather than only accessing it through a paid API.
Technical explanation
TII made Falcon 40B - described as the UAE's first large-scale AI model, with 40 billion parameters trained on one trillion tokens - open source for research and commercial use on 25 May 2023. TII followed in September 2023 with Falcon 180B, 180 billion parameters trained on 3.5 trillion tokens (four times the training data of Llama 2), reported at release as the largest openly available LLM and shown to outperform GPT-3.5 on the MMLU benchmark. Subsequent releases include Falcon 2 11B (a smaller, more efficient model with a vision-language variant) and Falcon Mamba 7B, an alternative state-space architecture TII describes as the top-performing open-source model of that architecture type as independently verified by Hugging Face.
Why it matters
Falcon is concrete evidence that frontier-scale, openly licensed model development is not exclusive to US Big Tech - it is a UAE government research institute's own contribution to the global open-model ecosystem, and it directly ties TII, and by extension Abu Dhabi's Advanced Technology Research Council, into the same conversation as Meta's Llama and Mistral AI.
Real-world example
Falcon 180B's September 2023 release was reported by multiple outlets as the largest openly available language model at that time, trained on 3.5 trillion tokens - a scale claim that was independently covered rather than only self-reported by TII.
Common misunderstanding
That 'open source' here means the same thing it does for ordinary software. Falcon's releases are open-weight - the trained model parameters are published for research and commercial use - which is a meaningfully different (and narrower) kind of openness than publishing full training code, data and methodology; the distinction matters when evaluating any 'open' model's actual reproducibility.
Something wrong here?
Every entry is hand-researched and hand-written by Nakoda. If a fact is stale, a source has changed or a definition needs sharpening, tell us and we will check it.

