Eleven v4 Speaks Arabic in 100 Milliseconds. The Hard Part Is What It Tells Your Customer

ElevenLabs Eleven v4: A Faster, More Expressive Machine Voice in Arabic
On September 28, 2026, ElevenLabs launched two new text-to-speech models: Eleven v4, its highest-quality and most expressive model, and Eleven v4 Turbo, built for real-time conversation with a median generation latency of about 100 milliseconds. Both support more than 90 languages, including Arabic, and their API prices are 72% off until October 12. The takeaway for business owners: the voice itself is no longer what stands between you and an automated receptionist that answers the phone. The real obstacle is now the information it relies on, the rules that govern what it says, and when it hands the call to a person.
What Exactly Was Released?
According to the models page in the company's official documentation, the release is a family of two models:
- Eleven v4: the top-quality model, aimed at voiceovers, multi-speaker dialogue, audiobooks and multilingual projects. A single request accepts up to 10,000 characters, roughly 10 minutes of audio, compared with 5,000 characters in the previous generation, v3.
- Eleven v4 Turbo: a real-time version with a median generation latency of about 100 milliseconds, with customer-support agents and AI assistants at the top of the company's listed uses.
Language coverage rose from 70+ in v3 to 90+, and both models support audio tags that control delivery and tone. TechCrunch reported that the model can clone a voice from a 10-second sample, and that the biggest quality jumps came in Japanese, Brazilian Portuguese, Mandarin and Cantonese, meaning Arabic was not one of the languages the announcement highlighted.
What Does It Actually Cost?
ElevenLabs prices its API in US dollars per 1,000 characters. On the official pricing page, Eleven v4 costs $0.022 instead of $0.08, and Turbo costs $0.011 instead of $0.04, with the discount running until October 12. Since the documentation estimates 10,000 characters at about 10 minutes of audio, a rough rule is that each minute of speech needs about 1,000 characters.
A practical example: a customer-service line where the voice agent speaks 1,000 minutes a month uses about one million characters. Once the discount ends, the voice costs about $40 a month with Turbo and about $80 with v4. That is the voice alone; the language model that decides the reply, the transcription of what the customer says, the phone line and the integration with your systems are separate costs, and usually the larger part of the bill.
Where Can It Help Your Business?
- Answering calls after hours: booking a clinic appointment, checking an order's status, or asking about opening hours and location, without long press-1, press-2 phone menus.
- Voiceovers for content: product videos, internal training and system walkthroughs for new staff in Arabic and English, with one consistent voice across every video and no studio booking each time.
- Reaching more people: reading written content aloud for those who prefer to listen, or serving expatriate workers in their own languages such as Urdu, Hindi and Filipino, all of which are on the supported list.
Before a Machine Voice Answers Your Customers
Better voice quality makes mistakes more convincing, not less. An agent that speaks in a confident, human tone and then gives a customer an appointment that does not exist is worse than a boring phone menu. That is why we start with these questions with our clients before choosing any provider:
- Where do its answers come from? The agent should read live from your booking, inventory or order system, not from a fixed script written once. We explained this in building a company knowledge base for AI.
- When does it hand over to a person? Complaints, large amounts and angry customers go to a human immediately, along with a summary of the conversation. Details in handing a conversation from the assistant to your staff.
- Does the accent suit your customers? The documentation lists Arabic without detailing dialects, so test it with scripts from your real calls: neighborhood names, your product names, numbers and Hijri dates.
- What about data and consent? Call recordings and transcripts are personal data governed by the Personal Data Protection Law, and cloning an employee's voice requires their explicit consent. See customer conversations and AI under the PDPL.
- Does the customer know they are talking to a machine? Disclosing it in the first sentence protects trust, especially since cloned voices have become a known fraud tool, as we covered in protecting your business from deepfake fraud.
Should You Switch Now?
If you already have a system running on v3 or Multilingual v2 that does the job, there is no rush. The documentation says Eleven v4 is available through the Text to Dialogue API and Turbo through its websocket connection, so confirm your current integration supports that before switching. The better move is to use the discount window to run a side-by-side test on a sample of your real scripts, then decide based on the results and the full price, not the temporary one.
How We Build It at Origami
We treat the voice engine as a replaceable part, not the foundation of the system. The logic that decides what gets said, the integration with booking, invoicing and inventory systems, the call log and the handover rules are all built on your side and under your control. That way you can try Eleven v4 today and move to another provider tomorrow if the price changes or a model better suited to your customers' dialect appears, without rebuilding everything. If you are starting from scratch, our article on voice AI in customer service lays out the full picture.
Our practical advice: pick one narrow path, such as confirming appointments or checking order status, run it for two weeks on a share of your calls, then measure how many calls ended without a handover, the error rate, and customer satisfaction. Real numbers from your own calls are more honest than any ranking in a comparison table.
Sources
Frequently asked questions
What is Eleven v4?+
A text-to-speech model ElevenLabs launched on September 28, 2026, alongside a Turbo version for real-time conversation with a median generation latency of about 100 milliseconds. It supports more than 90 languages, including Arabic, and a single request accepts up to 10,000 characters.
Does Eleven v4 support the Saudi dialect?+
The official documentation lists Arabic among supported languages without detailing dialects. The best way to judge is to test it on scripts from your real calls, including neighborhood names, product names, numbers and dates.
How much does Eleven v4 cost through the API?+
Until October 12, 2026: $0.022 per 1,000 characters for v4 and $0.011 for Turbo. After that, $0.08 and $0.04. About 1,000 characters equals roughly one minute of speech; the language model, phone line and integration are separate costs.
Can I use it to answer my customers' calls?+
Yes, the Turbo version is designed for customer-service agents. But success depends on connecting it to your systems so its answers are correct, clear rules for handing over to a person, telling customers they are talking to an automated assistant, and complying with the Personal Data Protection Law.
Follow Origami in Google
Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Related articles
- Artificial IntelligenceXiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 CentsXiaomi released MiMo-V2.6 open source under MIT: one model for text, audio, images and video, with a 1M-token context and prices from $0.14 per million tokens.
- Artificial IntelligenceAn Open Agent Built for Days of Work, Not Minutes: Atria Dawn and Its Real Running BillShanghai AI Lab released Atria Dawn under MIT: 744 billion parameters aimed at long multi-step tasks. What it is actually good for, and what running it in-house costs.
- Artificial IntelligenceThe Model Race Taps the Brakes: Outside Evaluators Get a Badge and a Desk Inside AnthropicOn September 12, 2026 Anthropic's CEO called for slowing AI capability gains and committed to letting independent evaluators inside. Here is what actually changes for your tech plan.
- Artificial IntelligenceSWE-2 Matches Frontier Coding at a Third of the Cost, Then Fails One TestCognition released SWE-2 on September 10, 2026: 50.0% on FrontierCode versus Fable 5.1's 50.9% at 64% lower cost, yet 28 points behind on Terminal-Bench 4.
- Artificial IntelligenceOpenAI's Two New Image Models Edit One Part of Your Product Photo Without a ReshootOpenAI shipped GPT Image 2.5 in two versions on September 8, 2026: Sunburst for editing precision, Flare for speed. What changed for your store, what it costs, where it helps.
- Artificial IntelligenceDeepSeek V4.1 Flash: 77% Cheaper, and Peak Hours Hit Your MorningDeepSeek ships V4.1 Flash on 10 September and routes V4 Pro requests to it at the cheaper rate: 77% off input, 70% off output. Its peak hours sit inside your working morning. The numbers, and the largest saving nobody notices.
Have a project in mind?
We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.
