AI News Feed
Market watch
Large Language Models

TII Introduces Falcon-ASR, a 1.6B-Parameter Arabic Speech Recognition Model

TII has introduced Falcon-ASR, a 1.6-billion-parameter speech recognition model focused on Arabic and the Emirati dialect. It supports five languages and reports lower error rates than compared systems on Arabic and internal Emirati evaluations.

TII evaluated Falcon-ASR on the six test sets of the Open Universal Arabic ASR Leaderboard, which is maintained by the ELM Research Center and ranks systems by equal-weight average word error rate and character error rate. Using the leaderboard's pinned manifests, TII said Falcon-ASR recorded 20.92% average WER and 8.79% average CER. Audar-ASR-V1-Turbo, with 2.35 billion parameters, recorded 23.17% WER and 9.23% CER; Cohere Transcribe Arabic (07-2026), with 2.0 billion parameters, recorded 25.87% WER and 11.80% CER; and omniASR LLM 7B, with 7.0 billion parameters, recorded 28.32% WER and 12.52% CER. TII said competitor figures were published leaderboard averages checked on September 30, 2026, and that Falcon-ASR's average WER was 2.25 percentage points better than the best published result in that snapshot.

For Emirati speech, TII said public evaluation data already includes Emirati through the UAE subset of Casablanca. It added an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts to assess accuracy beyond the public UAE subset. In that internal evaluation, Falcon-ASR recorded 22.73% WER and 10.19% CER. Qwen3-Omni-30B-A3B-Instruct recorded 26.80% WER and 12.72% CER; Audar-ASR-V1-Turbo recorded 27.89% WER and 13.75% CER; Cohere Transcribe Arabic (07-2026) recorded 31.05% WER and 18.07% CER; Qwen3-ASR-1.7B-hf recorded 31.52% WER and 13.35% CER; and Audar-ASR-V1-Flash recorded 32.87% WER and 15.36% CER. Falcon-ASR had the lowest WER and CER among the systems compared, with its WER 4.07 percentage points below Qwen3-Omni, the next best result.

TII said it trained Falcon-ASR on Emirati, Modern Standard Arabic, other Gulf and Arabic dialects, and English. The institute included background noise, overlapping speech, music, room reverberation and telephony effects, as well as variations in speed and pitch. It applied the same treatment to Emirati recordings, exposing the model to conditions it may encounter in meetings, calls and other everyday recordings. The model also supports word-level timestamps, linking each transcribed word to its position in the audio. TII said dialectal Arabic has fewer transcribed resources than Modern Standard Arabic, which makes training and evaluation harder, and that the aim is to transcribe the words people use in everyday speech, including dialectal forms and changes between languages.

Falcon-ASR transcribes English with the same model weights. On the seven public English test sets used by the Hugging Face Open ASR Leaderboard, TII reported a mean WER of 5.74%. The individual results were 1.75% on LibriSpeech clean, 4.21% on LibriSpeech other, 2.02% on SPGISpeech, 3.87% on VoxPopuli, 8.15% on GigaSpeech, 8.33% on AMI and 11.86% on Earnings-22. The model also supports French, Spanish and Portuguese, and all five languages use the same weights without requiring a language flag. The output is a transcript in the language spoken.

TII said Falcon-ASR builds on its Falcon3-Audio work. The architecture and training approach for Falcon3-Audio are described in the paper Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data. Falcon-ASR can be tried in a Hugging Face Demo Space, and TII said API access and native applications are planned.