Skip to content
Friday, September 18, 2026GULF & MENA BUSINESS NEWS
Dijla News

The Arabic large language models: Falcon, Jais, ALLaM

Falcon from Abu Dhabi's TII, Jais from G42 and partners, and Saudi Arabia's ALLaM anchor a growing family of Arabic-first models, now distributed on global clouds.

Arabic and English text on screens in an AI research lab
The Arabic large language models: Falcon, Jais, ALLaM

Three Arabic large language models anchor the Gulf's homegrown AI stack: Falcon, developed by Abu Dhabi's Technology Innovation Institute, whose open-source Falcon 40B topped download charts in 2023; Jais, the bilingual Arabic-English model built by Inception with Core42, Cerebras and Abu Dhabi's AI university in 2023; and ALLaM, the Saudi state's model family trained by the national data authority SDAIA, which by 2026 was distributed on Amazon's Bedrock cloud. Around them sits a growing family of tuned, smaller and domain-specific Arabic models.

The premise behind all of them is a measurable gap: frontier models trained overwhelmingly on English and Chinese text handle Arabic's morphology, dialects and right-to-left script measurably worse, and the fix requires intention, curated Arabic corpora and evaluation in the language itself. For the Gulf's AI sector, Arabic models are both a public-service necessity and a credibility test of sovereign AI claims.

The three flagships

ModelDeveloperSignificance
FalconTechnology Innovation Institute, Abu DhabiOpen-source releases from 2023; Falcon 180B among the largest open models of its year; Falcon Arabic released 2025
JaisInception with Core42, Cerebras, MBZUAIFirst open bilingual Arabic-English LLM family, 13B at launch in 2023, trained partly on Condor Galaxy supercomputers
ALLaMSDAIA, Saudi ArabiaState-trained Arabic-English models in multiple sizes; national deployment and cloud distribution

Why open weights mattered

Falcon and Jais shipped with open weights at exactly the moment developers worldwide were searching for usable non-GPT options, and the download numbers, Falcon 40B briefly led Hugging Face's charts, bought the region something Gulf technology had never had: developer mindshare. Openness is also a sovereign choice, the models can be hosted on national compute, audited, and tuned for government use without external API dependence, which fits both states' data-residency rules.

The Saudi route: state data, state model

ALLaM's differentiator is data access. SDAIA sits atop the kingdom's data programmes, and ALLaM was trained with Arabic-English corpora curated through state institutions, released in successive sizes and integrated into Saudi government platforms. Its 2026 availability on AWS Bedrock, announced with the US partnership wave, put a Saudi-trained model on a US hyperscaler's catalogue, a small bilateral marker in both directions.

What the models are actually used for

  • Government services: Arabic citizen-facing chat and document processing, the anchor use case in both states.
  • Enterprise: customer service and compliance tasks where dialect and terminology accuracy matter.
  • Media and religion-adjacent content: classical Arabic handling for publishing and reference applications.
  • Research: Arabic NLP benchmarks and university collaborations, now a functioning field rather than a workshop topic.

The honest scorecard

Gaps remain. Arabic evaluation benchmarks are thinner than English ones, so model comparisons carry uncertainty; dialect coverage, Gulf, Egyptian, Levantine, Maghrebi, is uneven; and the frontier-English models' Arabic keeps improving, narrowing the gap the local models exist to fill. Investment also flows overwhelmingly to compute infrastructure, with model teams comparatively small beside the megawatt budgets.

The counterweight is Jais's own lineage: it was trained on the Cerebras Condor Galaxy machines in Abu Dhabi, the same national-compute logic that drives G42's infrastructure build-out, and its successor generations tie the model story to the compute story, which is where Gulf funding actually concentrates.

The sovereignty argument, in plain terms

Why do states fund language models? The governance answer is the one officials give: public services in Arabic should not depend on inference APIs whose behaviour, pricing and data handling a foreign board controls, and a state that runs its citizen interfaces on rented intelligence has outsourced a civic function. The economic answer is adjacent: Arabic-language AI is a product category across twenty-plus markets, and owning credible entries in it positions Gulf firms as suppliers to a language community rather than customers of someone else's. Both arguments reduce to the same sentence, the language layer is too consequential to rent, and the region's three flagship models are the policy's working proof.

Building an Arabic model, technically

What distinguishes Arabic model-building from English is the data problem. Arabic text exists in vast quantity, classical corpora, news archives, social media across dozens of dialects, but it is fragmented, inconsistently normalised and skewed toward a few dialects and registers, so the industry's harder work has been corpus curation: cleaning, licensing and balancing data so a model learns the language's full range rather than its noisiest slice. Tokenisation is the second frontier, Arabic's rich morphology packs meaning into affixes that standard tokenisers shred inefficiently, and dedicated Arabic tokenisers measurably improve both performance and compute cost. Evaluation is the third, the community has built Arabic benchmark suites for reasoning, knowledge and safety, but they remain younger and thinner than English equivalents, which is why model cards carry more caveats here than elsewhere.

ChallengeWhat builders do
Corpus qualityCuration, licensing, dialect balancing
MorphologyArabic-aware tokenisation
EvaluationArabic benchmark suites, still maturing
Dialect rangeMulti-dialect training data

The three flagships solved these problems with different emphases, Falcon's open releases optimised for global re-use, Jais for bilingual balance from its first release, ALLaM for institutional Arabic of the kind governments actually use, and the differences show in deployments: developers tune Falcon worldwide, UAE platforms run Jais descendants, Saudi government systems run ALLaM. A model family for every register of a language is what a mature ecosystem looks like, and Arabic is closer to it than any language outside the Sino-English duopoly.

Why it matters

Language models are infrastructure for a language's digital future: search, education, government, publishing. Arabic is spoken by more than 400 million people with a thin open-model inventory, and the Gulf states have chosen to own that layer rather than rent it, for cultural reasons as much as commercial ones. Falcon, Jais and ALLaM are the proof the choice was executable, and their download counts, deployments and successors are the region's quietest credible AI metric.

Frequently Asked Questions

What are the main Arabic large language models?
Falcon from Abu Dhabi's Technology Innovation Institute, Jais from Inception with Core42, Cerebras and MBZUAI, and Saudi Arabia's state-trained ALLaM family from SDAIA.
Is Falcon an open-source model?
Yes. Falcon's releases from Falcon 40B in 2023 onward shipped with open weights, and Falcon 40B briefly topped global open-model download charts; a dedicated Falcon Arabic model followed in 2025.
Why do Arabic-language models matter?
Frontier models trained mostly on English handle Arabic morphology, dialects and script measurably worse. Arabic LLMs let the region's governments and enterprises run accurate language services on sovereign infrastructure.

Sources

  1. Technology Innovation Institute
  2. Saudi Data and AI Authority (SDAIA)