NVIDIA Taught Its Speech Model Najdi and Hijazi With SDAIA's SADA Data: Errors Fell From 55 to 30 Words in 100

NVIDIA Taught Its Speech Model Najdi and Hijazi With SDAIA's SADA Data: Errors Fell From 55 to 30 Words in 100
The Saudi Press Agency reported today, Saturday 3 October 2026, that NVIDIA used SADA, the Saudi audio dataset released by the Saudi Data and AI Authority (SDAIA) with the Saudi Broadcasting Authority, to train its automatic speech recognition model, Nemotron 3.5 ASR. According to the study NVIDIA published on its developer blog on 30 September, a model that used to get about 55 words in every 100 wrong in Najdi and Hijazi speech now gets about 30 wrong.
That matters to any business with an archive of customer calls, a customer-service line, or recorded meetings. But a headline number is not a decision. This article explains what this level of accuracy is good for and what it is not, the licensing condition the news does not mention, and where your customers' voices should be processed.
What happened, in numbers
- The model: Nemotron 3.5 ASR, a 600-million-parameter streaming transcription model covering 40 language-locales, 32 of which transcribe out of the box, with Arabic in its transcription-ready tier. Its Hugging Face page says it is ready for commercial use, under the OpenMDW-1.1 licence.
- The data: 133.7 hours of Najdi and Hijazi speech from SADA. NVIDIA's team mixed it 90% Saudi speech with 10% public FLEURS data (7% English, 3% Arabic) so the model would not forget what it already did well.
- Word error rate on Najdi and Hijazi: from 55.05% to 29.96%. Across the whole SADA set and all its dialects: from 58.84% to 35.61%.
- Character error rate on Najdi and Hijazi: from 31.63% to 12.18%.
- No loss on the two checks NVIDIA ran: English improved slightly (11.04% to 10.42%), and so did the FLEURS Arabic test (12.67% to 11.41%).
- Training cost: 12,000 steps in about four and a half hours on two RTX PRO 6000 Blackwell workstation GPUs.
- Speed: the model runs with latency starting at 80 milliseconds, and its page lists five streaming settings from 80 to 1,120 milliseconds. The error rates above, however, were measured at the 320-millisecond default; accuracy improves with larger chunks, so the faster settings trade some accuracy for speed.
As for SADA itself, the agency says it holds about 667 hours of transcribed audio, mostly in Saudi dialects, covering more than ten dialects across more than 125,000 clips labelled by the speaker's dialect. SDAIA published it on Kaggle for researchers and developers.
30 wrong words in 100: what it is good for, and what it is not
Word error rate counts every word that was substituted, dropped or added. At 30%, the transcript of a 100-word stretch of call may contain around 30 errors. That is a big improvement on 55, but it does not make the transcript a verbatim record you can rely on.
The gap between the word rate (about 30%) and the character rate (about 12%) has a practical meaning. A word with a single wrong letter counts as a whole wrong word, so part of the errors are a letter or two inside a word, not a word lost entirely. That is why a transcript can be more useful for search and tagging than the 30% figure suggests, provided your search tolerates near-miss spellings rather than requiring an exact match.
- Good for: searching your call archive for a product name or a recurring complaint, tagging calls by reason, summaries a staff member reviews, and picking calls for quality review instead of listening at random.
- Not good enough on its own for: a record you would rely on in a dispute, executing a financial instruction from a call transcript without review, or an automated agent that takes the customer's words literally and acts on them.
Note too that SADA comes from television: more than 600 hours from more than 57 programmes and series supplied by the Saudi Broadcasting Authority. Your customer's call is phone audio, possibly from a car or a busy shop, sometimes with two people talking over each other. NVIDIA itself wrote in the study that this workflow is not evidence for every Arabic dialect or every deployment environment. The number that matters to you is the number on your own calls.
The condition the news leaves out: SADA's licence is non-commercial
The base model is licensed for commercial use. The SADA dataset, however, is published on Kaggle under the Creative Commons CC BY-NC-SA 4.0 licence: attribution, non-commercial, share-alike. If you want to train a model for your business on SADA, check the licence terms or get the data owner's permission first, and do not assume that being open to researchers means being open to your product.
NVIDIA published the experiment's steps and code in the post itself, and linked a general fine-tuning notebook whose sample data is an English dataset, not SADA. But the study does not say it released the model as fine-tuned on SADA. What is available to you today is the base model and the method. The data you actually have the right to use in your business is, in most cases, your own recordings.
Your calls are the data: what to prepare before any trial
- A real sample: pick calls from your actual line, in your customers' dialects, at the quality your system records, including hard ones with noise, crosstalk, numbers and product names.
- A reference transcript: have a staff member transcribe them by hand once. That transcript is the ruler you measure every model and every vendor against, not the demo the seller brings.
- Measure what matters to you: did the model catch the order number, the product name and the reason for the call? Those fields matter more to your business than the overall error rate.
- Notice and purpose: the Personal Data Protection Law defines personal data as any data, whatever its source or form, that can identify an individual, including contact numbers, and it treats recording, storing, using and transmitting as forms of processing. A recorded customer call is usually personal data, and transcribing it is processing. Make sure you have a lawful basis for each purpose, and that the privacy notice your customer hears or reads covers transcription and analysis, especially if you plan to use recordings to train a model. We covered this in detail in our article on customer conversations and the PDPL.
Where your customers' voices are processed
This is where the two paths really differ:
- An open model on your own server or in an in-Kingdom cloud: the audio never leaves the Kingdom, and you decide how long it is kept and who sees it. It runs best on a server with an NVIDIA GPU (the model page lists NVIDIA architectures from Volta to Blackwell, on Linux), and the page also points to NeMo-Speech.cpp, a lightweight NVIDIA runtime that can run it on an ordinary CPU. Fine-tuning it on your calls does need GPUs, and either way someone has to run and update it. We walked through the hardware maths in our article on running models on your own servers.
- A cloud transcription service outside the Kingdom: easier to start, but your customers' recordings travel abroad, and transferring personal data outside the Kingdom is governed by the Regulation on Personal Data Transfer outside the Kingdom, one of the implementing regulations of the PDPL. Ask your provider where audio is processed and stored before you ask about price.
Questions for your call-centre or voice-bot vendor
If you already have a call-centre vendor, or are considering a voice agent that answers your customers, these questions reveal a lot in a single meeting:
- Which dialects does your system actually support, and on what data did you measure its accuracy?
- Will you measure accuracy on a sample of our calls, against a reference transcript we produce?
- Where is the audio processed and where is the text stored, inside the Kingdom or outside?
- Do you use our recordings to train your models, and how do we switch that off?
- How long do you keep recordings and transcripts, and can we export and delete them?
The Origami view
Voice is the most wasted data in a Saudi business. A call with a complaint, an order, or a promise from a staff member ends, and then nobody searches it and it never links to the customer's record. What happened with SADA shows that the gap between global models and our dialects narrows with well-labelled Saudi data, even on a small 600-million-parameter model.
Our reading for business owners: do not start with an automated agent answering customers. Start with an archive you can search, tagging that links every call to its reason and to the customer's record in your system, and a person who reviews. Then measure accuracy on your own calls before scaling. The first decision is not which model to pick; it is where your customers' voices are processed and who owns them.
Conclusion
- NVIDIA cut its model's errors in Najdi and Hijazi from about 55 to about 30 words in every 100, training on 133.7 hours of SADA data.
- A 30% error rate is enough for search, tagging and summaries with human review, and not enough for a verbatim record or an automated decision.
- The base model is licensed for commercial use, but SADA is licensed for non-commercial use only, so check the licence before building on it.
- Measure any model or vendor on a sample of your own calls, against a reference transcript you produce.
- Know where the audio is processed, and make sure your privacy notice covers transcription and analysis.
Sources
- Saudi Press Agency: NVIDIA Leverages SDAIA's SADA Dataset to Enhance Nemotron 3.5 ASR Model (3 October 2026)
- NVIDIA Developer Blog: Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects (30 September 2026), including the error-rate tables and training details
- The model's Hugging Face page: size, licence, languages and hardware
- The SADA dataset on Kaggle: description and CC BY-NC-SA 4.0 licence
- SDAIA: Guide to the Saudi Personal Data Protection Law for controllers and processors, including the definitions of personal data and processing in Article 1 and the reference to the Regulation on Personal Data Transfer outside the Kingdom
Frequently asked questions
Can my business use NVIDIA's SADA-trained model today?+
The study published the training method and code, and linked a general fine-tuning notebook, but it does not say it released the version fine-tuned on SADA. What is available is the base Nemotron 3.5 ASR model, whose page says it is ready for commercial use, and you can fine-tune it the same way on data you have the right to use. SADA itself is licensed CC BY-NC-SA 4.0, which is non-commercial, so review that licence before building a product on it.
What does a 30% word error rate mean in practice?+
Roughly 30 words substituted, dropped or added in every 100. A transcript at that accuracy works for searching a call archive, tagging calls and summarising them with a staff member reviewing, but not on its own as a verbatim record or as the basis for an automated decision. The character error rate was about 12%, so part of the errors are a letter or two inside a word.
My customers are not from Najd or Hijaz. Do these results apply to them?+
The training targeted Najdi and Hijazi. Across the full multi-dialect SADA set, the error rate fell to 35.61%, not 30%. NVIDIA itself wrote that the result is not evidence for every Arabic dialect or deployment environment. The only way to know is to measure on a sample of your own customers' calls.
Does recording and transcribing customer calls fall under the Personal Data Protection Law?+
Usually yes, because a call typically carries the customer's number, name or order details. The law defines personal data as any data, in any form, that can identify an individual, and counts recording, storing, using and transmitting as processing. Make sure you have a lawful basis for each purpose and that your privacy notice covers transcription and analysis, and remember that sending recordings to a service outside the Kingdom is governed by the Regulation on Personal Data Transfer outside the Kingdom.
Follow Origami in Google
Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Related articles
- Artificial IntelligenceGoogle Hands Gemini 4 Argon to Cyber Defenders First, With a 1M-Token Output LimitGoogle's Gemini 4 Argon finds and patches vulnerabilities on its own and outputs up to 1M tokens at $2/$10 per million. Who gets it first, and how to prepare.
- Artificial IntelligenceClaude Marketplace Puts 2,000 Connectors One Click Away From Your Company DataAnthropic's Claude Marketplace launched with 2,000+ connectors, plugins and agents. What to check before linking Claude to your systems, and when to build your own.
- Artificial IntelligenceOpenAI's Dots Work While You Sleep: What Do They Do Without Asking You?OpenAI's dots are always-on AI agents with their own cloud computer and 4,000+ app connections. What they do alone, what needs your approval, and how to set them up.
- Artificial IntelligenceEleven v4 Speaks Arabic in 100 Milliseconds. The Hard Part Is What It Tells Your CustomerElevenLabs' Eleven v4 and v4 Turbo bring ~100 ms voice generation and 90+ languages including Arabic. Prices, limits, and what to fix before a voice agent answers customers.
- Artificial IntelligenceXiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 CentsXiaomi released MiMo-V2.6 open source under MIT: one model for text, audio, images and video, with a 1M-token context and prices from $0.14 per million tokens.
- Artificial IntelligenceAn Open Agent Built for Days of Work, Not Minutes: Atria Dawn and Its Real Running BillShanghai AI Lab released Atria Dawn under MIT: 744 billion parameters aimed at long multi-step tasks. What it is actually good for, and what running it in-house costs.
Have a project in mind?
We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.
