Falcon-Emirati-7B:让 LLM 学会阿联酋方言、文化与语义细微差别
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Falcon-Emirati-7B 发布,这是一个在 Falcon-H1-Arabic 基础上专攻阿联酋方言的 7B 模型,目标是像母语者一样理解和生成 Emirati Arabic 的词汇、语气与文化语境。
Arabic is really a family of languages living under one name. Modern Standard Arabic is what you read in the news or a textbook, but it's rarely how people actually talk to each other. In the UAE, day-to-day conversation, humor, negotiation, and storytelling happen in Emirati Arabic, a Gulf dialect with its own vocabulary, its own rhythm, and a culture wrapped tightly around it. Emirati poetry, especially nabati poetry, along with proverbs and short anecdotes, carries meaning that doesn't survive a literal, word-for-word reading. A model that only knows MSA can translate every word of an Emirati sentence and still miss what it actually means.
That's the gap Falcon-Emirati-7B is built to close. It's a dialect-specialized model on top of Falcon-H1-Arabic, aimed at understanding and generating Emirati Arabic the way a native speaker would: the vocabulary, the tone, and the cultural context behind it.
Built on Falcon-H1-Arabic
We didn't start from scratch. Falcon-Emirati-7B is built on Falcon-H1-Arabic, our Arabic model family that already set new benchmarks for the language earlier this year. Falcon-H1-Arabic uses the Falcon-H1 hybrid architecture: State Space Models (Mamba) and Transformer attention running in parallel inside every block, with their outputs fused before each block's projection. That combination gives the linear-time efficiency of Mamba on long sequences while keeping the precision of attention for long-range dependencies, which matters for a morphologically rich language like Arabic. The family spans three scales (3B, 7B, and 34B parameters) with context windows up to 128K and 256K tokens, and it was already trained on a broad mix of MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data.
That gave us a strong starting point: a model that already understood Arabic broadly, handled long context well, and had some dialectal exposure baked in. Falcon-Emirati-7B takes that foundation and pushes it specifically toward the Emirati dialect, the vocabulary, the grammar, and the cultural knowledge that a general Arabic model, however capable, doesn't pick up on its own.
We built Falcon-Emirati-7B on the 7B variant specifically. It's the sweet spot in the family: large enough to hold onto the nuance that dialect adaptation needs, but small enough that both training and inference stay practical. The 34B model would likely push quality a bit further, but at a training and serving cost that doesn't make sense for a dialect-specialized chat model, and the 3B model doesn't leave enough headroom for the depth of cultural and linguistic understanding we were after. 7B gave us the best balance of quality against training and inference cost.
Why Dialect Adaptation Is Hard
Turning a general Arabic model into an Emirati-dialect specialist sounds like a smaller job than building the base model in the first place. It isn't. A few things make it genuinely difficult:
- Emirati is mostly a spoken dialect. It shows up far less in writing online than MSA, or even other Gulf and Levantine dialects, so there just isn't as much raw text to learn from.
- Meaning is often non-literal. Idioms, proverbs, and poetic references lean on shared cultural context, not surface vocabulary.
- There's no established playbook. There isn't a well-documented recipe for how much dialectal data is enough, how to mix it with MSA and general Arabic, or which training stage (continued pre-training, SFT, or preference optimization) matters most for picking up a dialect.
That last point shaped how we worked. A lot of building Falcon-Emirati-7B came down to trial and error: testing different data mixes, training stages, and supervision strategies, and using both human judgment and benchmark scores to figure out what actually moved the needle.
Our Approach to Data
We built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic's pretraining, drawing on three complementary sources.
1. Authentic Emirati-Dialect Web Data
We crawled and curated content from Emirati websites and forums written natively in the dialect, not translated or transliterated from MSA. This is where we got our ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and MSA that shows up in real usage.
2. MSA Data About Emirati Culture and Identity
Alongside the dialectal text, we pulled in MSA-language material specifically about Emirati culture, heritage, and language: articles and references on local customs, values, history, and social norms, including how Emiratis are perceived and stereotyped. This doesn't teach the model to write in dialect, but it teaches the model what it's talking about when Emirati topics come up, things like heritage, etiquette, and the context a native speaker just knows.
3. Synthetic Data, Guided by Glossaries and Style Rules
Authentic dialectal text alone wasn't enough to cover the range of topics a chat model actually needs to handle day to day. So we generated a large amount of synthetic Emirati-dialect data to fill the gaps. We didn't just let a generator model improvise in "Gulf-ish" Arabic. We constrained it with strict rules and glossaries and dictionaries built specifically for Emirati vocabulary and grammar. Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that's grammatically fine but sounds off to anyone who actually speaks the dialect.
Finding the Right Adaptation Recipe
Since there's no standard recipe for MSA-to-dialect adaptation, we treated the training strategy itself as something to figure out experimentally. We ran ablations on how much dialectal data to inject and at which stage of training, how to balance authentic crawled data against synthetic data without the model overfitting to synthetic patterns, and how much MSA cultural context was actually needed to keep it culturally grounded rather than just fluent on the surface. At each step we leaned on a mix of automatic scoring and native-speaker review, since automatic metrics alone don't capture naturalness, tone, or cultural fit well enough to trust on their own.
Evaluation Methodology
We tracked progress throughout training with two complementary approaches:
Manual Evaluation by Native Speakers
Emirati native speakers reviewed model outputs directly, judging not just whether an answer was correct but whether it sounded right: naturalness, tone, cultural appropriateness. These are the things a benchmark score won't tell you but a native ear catches immediately.
Automatic Evaluation on Alyah
For quantitative tracking, we used Alyah (الياه, "North Star"), a benchmark we and the community released specifically to evaluate Emirati-dialect capability in Arabic LLMs. Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers and spanning categories from everyday greetings and etiquette to figurative language, heritage knowledge, and Emirati poetry: the categories where dialect and culture matter most and where generic Arabic models tend to struggle. Full details on Alyah's construction and category breakdown are available in our benchmark blog post, and background on the base model family is available in the Falcon-H1-Arabic announcement.
Results
Falcon-Emirati-7B scores 84.83% on Alyah, ahead of every other Arabic and multilingual model we compared it against, including several models many times its size. The chart below shows where it lands next to a representative set of leading instruction-tuned models on the Alyah leaderboard.
Alyah accuracy (%), instruction-tuned models. Falcon-H1-Arabic family models are excluded from this comparison since Falcon-Emirati-7B is built on top of them.
What the Results Tell Us
A couple of things jump out from this comparison. Size alone doesn't buy you dialect competence. Some of the largest multilingual models here score well below smaller, more dialect-aware ones, which tells you Emirati proficiency has to be trained for on purpose, not picked up as a side effect of scale. The models that do best also tend to be Arabic-native or Arabic-focused to begin with, which lines up with what we saw during our own ablations: general Arabic and dialect coverage is a necessary starting point, but it still takes targeted, dialect-specific work to close the rest of the gap, particularly on the hardest parts of Alyah, like poetry, heritage knowledge, and the language-and-dialect category itself.
This also matches what came out of the Alyah benchmark release more broadly: even strong models show real degradation once you move into genuinely dialectal, culturally embedded content. That gap doesn't close on its own with bigger models. It takes data and evaluation built specifically for the dialect.
Beyond Multiple Choice: LLM-as-Judge Evaluation
Multiple-choice accuracy tells you whether a model can recognize the right answer among four options. It doesn't tell you whether the model will actually produce Emirati Arabic on its own when someone just talks to it. So alongside Alyah, we ran a second evaluation: open-ended generation on the same 1,173 Alyah questions, scored by an LLM judge (Gemini 3.7 Flash) against five models, Falcon-Emirati-7B, ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat, and Fanar-2-27B-Instruct, chosen as the strongest competing models from the Alyah leaderboard.
The judge scored each answer on two separate dimensions: whether the content was correct, and, independently, whether the answer actually came back in Emirati dialect rather than MSA. We report both a partial-credit score (the judge's graded assessment) and a stricter pass/fail version, plus how often each model abstained instead of answering.
LLM-judged correctness on the 1,173 Alyah questions, open-ended generation, Gemini 3.7 as judge.
LLM-judged dialect fidelity on the same questions: does the answer actually come back in Emirati, or does the model default to MSA?
Falcon-Emirati-7B leads on correctness, but the real gap is in the second chart. On dialect fidelity, Falcon-Emirati-7B scores 0.52 (partial credit) against 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and effectively 0.00 for Fanar-2-27B-Instruct. That's not a small edge, it's close to two orders of magnitude at the low end. In practice, this means the other models often know the right answer but say it in Modern Standard Arabic by default, even when asked directly in Emirati. Falcon-Emirati-7B is the only one of the five that reliably answers back in the dialect it was asked in.
Fanar-2-27B-Instruct stands out for a second reason too: it abstains far more than any other model, declining to answer 26.2% of the time, versus under 5% for every other model in the comparison. Combined with its correctness score of 0.27 (partial credit), the lowest of the five, it suggests a model that is both less willing and less able to engage with Emirati-specific content, rather than one that's just answering in the wrong register.
Dialect fidelity by Alyah category, partial credit. Falcon-Emirati-7B is the only model that consistently switches into Emirati; the others stay in MSA across nearly every category.
Breaking dialect fidelity down by category makes the pattern even clearer. It holds across every single category in Alyah, from everyday greetings to poetry, which suggests this isn't a narrow trick learned for a handful of question types. It's a general shift in what register the model defaults to when Emirati is the expected register. The one place competing models do relatively better, Greetings & Daily Expressions, is also the category where Emirati and MSA overlap the most, so it's the easiest place for a generic Arabic model to accidentally sound right.
Pairwise Comparison, Category by Category
As a third lens on the same question, we ran head-to-head pairwise judging: for every Alyah question, the judge (Gemini 3.7 Flash) was shown Falcon-Emirati-7B's answer next to a competing model's answer, blind to which was which, and asked to pick the better one. The radar charts below show the resulting win rate by category against three competing models: Jais-2-8B-Chat, ALLaM-7B-Instruct-preview, and Fanar-2-27B-Instruct.
Pairwise win rate by category, Falcon-Emirati-7B vs. Jais-2-8B-Chat, ALLaM-7B-Instruct-preview and Fanar-2-27B-Instruct judged head-to-head by Gemini 3.7.
Falcon-Emirati-7B wins the majority of categories against all three competitors, and by a wide margin in the categories that depend most on dialect and cultural fluency. Against Jais-2-8B-Chat, the gap is largest in Poetry & Creative Expression (0.69 vs. 0.31) and Language & Dialect (0.62 vs. 0.38). Against ALLaM-7B-Instruct-preview, the same two categories again show the widest margins: Poetry & Creative Expression (0.66 vs. 0.34) and Language & Dialect (0.58 vs. 0.42). Against Fanar-2-27B-Instruct, the margins are the widest of all three matchups, and Falcon-Emirati-7B wins every single category, topping out at Poetry & Creative Expression (0.88 vs. 0.12) and Religious & Social Sensitivity (0.80 vs. 0.20).
The one category where competing models hold their own is Greetings & Daily Expressions: Falcon-Emirati-7B narrowly loses it to Jais-2-8B-Chat (0.46 vs. 0.54) and ties ALLaM-7B-Instruct-preview at 0.50, though it still wins clearly against Fanar-2-27B-Instruct (0.70 vs. 0.30). That's broadly consistent with what we saw in the dialect-fidelity breakdown above, greetings are the category where Emirati and MSA overlap most, so it's the easiest place for a generic Arabic model to sound native even without dedicated dialect training. Everywhere the dialect is more distinctive, poetry, figurative language, heritage knowledge, Falcon-Emirati-7B's advantage holds up clearly across every competing model we've tested it against.
Understanding Emirati culture
Language is also about understanding the culture behind a conversation. We evaluated Falcon-Emirati-7B on the UAE portion of ArabCulture-Dialogue, a benchmark for cultural understanding in Arabic. In its multiple-choice task, models choose the most culturally appropriate reply from three options. Read more in the research paper.
We tested all four models on the same 283 UAE scenarios, in both Emirati Arabic and Modern Standard Arabic, with varying amounts of location information.
Overall accuracy on the UAE multiple-choice task, averaged across both language varieties and all location settings. These are our evaluation results.
Falcon-Emirati-7B scored 85.57%, the highest among the four models tested, ahead of ALLaM-7B (83.39%), Jais-2-8B (73.79%), and Fanar-2-27B (71.50%). These results highlight its ability to recognize culturally appropriate responses in Emirati conversations.
See It In Action
Numbers only tell part of the story, so below is a live, interactive comparison: five real Emirati prompts, run through Falcon-Emirati-7B side by side with Fanar-2-27B-Instruct and ALLaM-7B-Instruct-preview. Swipe or use the arrows to move between prompts, and expand any response to read it in full.
Interactive comparison: Falcon-Emirati-7B vs. Fanar-2-27B-Instruct vs. ALLaM-7B-Instruct-preview on five Emirati prompts, covering Eid greetings, local history, poetry, and heritage knowledge. Outputs are shown as generated and may contain errors; highlighting identifies our model, not a factual rating.
Try Falcon-Emirati-7B
Falcon-Emirati-7B is available on our chat platform. 🚀 Try it now: https://chat.falconllm.tii.ae/?model=Falcon-Emirati-7B
Responsible AI and Limitations
Like any language model, Falcon-Emirati-7B can reflect biases in its training data and will sometimes get things wrong, especially on rare expressions, highly localized references, or edge cases we didn't have much data for. Dialect and cultural nuance are subjective in places, and even native speakers won't always agree on the "right" answer. We'd recommend evaluating the model for your specific use case before relying on it for anything sensitive, official, or high-stakes, and we'd genuinely welcome feedback from the Emirati community to help us improve both the model and Alyah going forward.
Aknowledgment
We would like to extend our sincere thanks to Mikhail Lubinets and Matthieu Berjon for their continuous support with the compute infrastructure, and to Jatin Mittal for his help in making the model accessible through the Falcon Chat app.
Citation
@misc{falcon_emirati_7b_2026,
title = {Falcon-Emirati-7B: A Dialect-Specialized Arabic LLM for the Emirati Dialect},
author = {Shaikha Alsuwaidi, Omar Alkaabi, Maitha Alhammadi, Hamza Alobeidli, Ahmed Alzubaidi, Mohammed Alyafeai, Leen AlQadi, Basma Boussaha, Hakim Hacid},
organization = {Technology Innovation Institute},
year = {2026}
}
来源:Hugging Face:Blog · huggingface.co






