Evidence at a glance
The Gap Between Modern Standard Arabic and Spoken Dialects
Arabic functions as a family of languages where Modern Standard Arabic (MSA) dominates textbooks and news while rarely reflecting everyday street speech. In the UAE, daily conversations, humor, poetry, and negotiations happen in Emirati Arabic, embedded with distinct rhythms and cultural nuances that literal translations miss. A model trained only on MSA can translate word-for-word yet still misunderstand the underlying meaning.
Falcon-Emirati-7B, developed by the Technology Innovation Institute (TII) in Abu Dhabi, aims to close this exact gap. Rather than training from scratch, it adapts the existing Falcon-H1-Arabic foundation model family specifically for the Emirati dialect. By targeting local vocabulary, tone, and cultural context, the model attempts to emulate native-level pragmatic competence.
Dialect Specialization Built on a Hybrid Architecture Foundation
Falcon-Emirati-7B builds directly upon the 7B variant of Falcon-H1-Arabic, leveraging a hybrid architecture that runs Mamba state space models and Transformer attention in parallel within every block. This design merges Mamba's linear-time efficiency on long sequences with Transformer's precision on long-range dependencies, which is critical for a morphologically rich language.
The 7B size represents a sweet spot between retaining the nuance required for dialect adaptation and maintaining practical training and inference costs. While a 34B model increases serving overhead, a 3B model leaves insufficient headroom for deep cultural understanding. Building on this base, the adaptation effort focused squarely on gathering and aligning spoken Emirati dialect data.
Composite Data Engineering Against Spoken Text Scarcity
Transforming a general Arabic model into an Emirati specialist presents a core challenge: Emirati is primarily a spoken dialect, appearing far less in online writing than MSA. With no established public training recipe, the team relied on trial and error to build a specialized pipeline consisting of three complementary data sources.
The first source comprises native dialect content crawled from Emirati websites and forums written directly by locals, establishing a baseline for everyday phrasing. The second source includes MSA materials detailing Emirati culture, history, and social norms to anchor background knowledge. The third relies on synthetic data governed by dialect glossaries and style rules, expanding vocabulary and style coverage where authentic text falls short.
The Discrepancy Between Multiple-Choice Scores and Generation Fidelity
To evaluate performance, the team tested the model on the Alyah benchmark, which contains 1,173 native-collected multiple-choice questions. Falcon-Emirati-7B scored 84.83% accuracy, outperforming ALLaM-7B-Instruct-preview at 77.24% and Jais-2-8B-Chat at 78.09%. It also reached 85.57% accuracy across 283 scenarios in the ArabCulture-Dialogue UAE subset, demonstrating solid cultural knowledge.
However, open-ended generation evaluations reveal a critical nuance: when judged by Gemini 3.7 Flash on open responses, the model achieved a dialect fidelity score of 52.1% against a content correctness score of 48.8%. Comparison models like ALLaM scored only 5.3% in fidelity. This gap demonstrates that answering multiple-choice cultural questions correctly does not automatically translate to naturally generating authentic dialect text.
Boundaries and Actionable Takeaways for Specialized Dialect Adaptation
Falcon-Emirati-7B demonstrates that specialized fine-tuning for regional dialects can outperform larger general models at a modest parameter scale, though it comes with distinct boundaries and transparency limits. While the data pipeline and ablation strategies are outlined, the lack of published training hyperparameters and synthetic data ratios poses replication hurdles, and open-ended evaluations rely heavily on external LLM grading.
For technical leaders, this offers an actionable deployment takeaway: when handling local customer service, social content, or cultural contexts in the UAE, specialized dialect models serve as valuable complements to general Arabic models. However, evaluations must not rely solely on multiple-choice benchmark scores; teams should independently verify colloquial fidelity and idiom appropriateness before production use.