top of page

Sarvam AI Saaras V4 Raises the Bar for Indic Speech, but Benchmarks Are Only the Start

Sep 27
12 min read

Sarvam AI launched Saaras V4 with support for 22 Indian languages, global English, and five ways to format a transcript. The Sarvam AI Saaras V4 release also claims leading accuracy across several English and Indic speech benchmarks. That combination puts pressure on general-purpose services from OpenAI, ElevenLabs, and Deepgram.

The announcement matters because speech recognition in India is not simply an English transcription problem. Real conversations combine regional languages, English terms, local accents, compressed phone audio, and frequent speaker changes. A system can perform well on clean recordings while struggling with the calls that businesses actually receive.

Sarvam is betting that regional depth can compete with the broader language coverage and established infrastructure of global vendors. Its five clearest differentiators are a custom language-model decoder, broad Indic coverage, five native output modes, keyterm prompting, and low-latency streaming. However, most performance evidence still comes from Sarvam’s evaluations, so deployment results remain the more important test.

Sarvam AI Saaras V4 Changes the Multilingual ASR Contest

Saaras V4 turns Sarvam’s India-focused speech model into a broader challenge to global transcription platforms.

Sarvam announced the model on August 24, 2026, through its detailed Saaras V4 release. It introduced the API version several days earlier and made the model generally available for Sarvam Voice Agents on September 2.

The release retains support for India’s 22 scheduled languages while expanding English recognition beyond Indian English. Sarvam describes this addition as global English support, including accents represented in international speech datasets.

That shift widens the model’s addressable workload. A customer-service operation in India might receive Hindi, Tamil, or Bengali calls alongside conversations involving Indian, British, or American English. Using one recognition system across those calls can reduce routing rules and model switching.

Saaras V4 is available through REST, batch, legacy WebSocket, and newer real-time interfaces. Sarvam’s model documentation lists voice agents, call analytics, code-mixed speech, and 8 kHz telephone audio among its target uses.

The REST interface accepts recordings of up to 30 seconds. The batch service handles files lasting up to two hours, while the streaming option returns partial transcripts during live conversations. Speaker separation is available through batch processing.

The release also illustrates a rapid product cycle. Saaras V3 arrived in February 2026 with the same 22 Indian languages and English. Sarvam said that version reduced word error rate on IndicVoices from approximately 22 percent to about 19 percent.

V4 changes the competitive claim. V3 was principally framed as an Indian speech specialist. V4 adds international English benchmarks and presents itself as one model for both Indic and wider English workloads.

This does not mean Sarvam suddenly covers every language that a global transcription service supports. Saaras remains focused on 23 named languages, while other platforms advertise much larger language catalogs. Its argument is depth within the languages it supports, rather than the longest possible language list.

That distinction creates the article’s central tension. Sarvam says specialization produced better recognition across difficult Indian languages without sacrificing English performance. Global vendors can counter with broader coverage, mature tooling, and features extending beyond basic transcription.

For enterprise buyers, the decision is therefore not about one leaderboard position. It concerns which system performs reliably on their actual accents, vocabulary, audio channels, and code-switching patterns.

A 3-Billion-Parameter Decoder Connects Audio With Language

The first defining feature is an architecture designed to treat transcription as contextual language generation, not isolated sound matching.

Saaras V4 combines an audio encoder with a 3-billion-parameter autoregressive language-model decoder. Sarvam says it trained this hybrid state-space decoder internally from scratch.

The audio encoder extracts phonetic and acoustic information from a waveform. A temporal-downsampling adapter then compresses that information before projecting it into the decoder’s embedding space. This reduction helps longer recordings fit within the model’s available context.

The decoder processes those audio representations alongside a text prompt. It generates transcript tokens sequentially, feeding each result back into the model before producing the next token.

This approach matters when several words sound similar. Acoustic evidence alone may not settle which word a speaker used. Sentence context, grammar, and likely word sequences can guide the decoder toward a plausible transcription.

An LLM-based decoder also supports instructions that control the output. The same underlying model can receive a request to preserve filler words, normalize numbers, transliterate speech, or translate the result into English.

However, contextual decoding introduces a familiar risk. A model that predicts plausible text can produce words that sound reasonable but were never spoken. This failure is especially serious in medical, financial, legal, and compliance records.

Sarvam says V4 was designed for noisy recordings, dialect variation, and code-mixed speech. Those are demanding conditions because the acoustic signal may already be ambiguous. A language-aware decoder can help, but it must not replace missing evidence with fluent invention.

Developers should therefore test deletion, insertion, and substitution errors separately. An apparently readable transcript can still omit a qualification, alter a number, or replace an unfamiliar name.

The architecture also raises questions about reproducibility. Sarvam describes the model and evaluation process but has not released the Saaras V4 weights. Buyers cannot independently inspect its training data, run it on their own infrastructure, or verify every architectural claim.

That makes the hosted API the practical unit of evaluation. Teams need to measure the service as delivered, including latency, uptime, data handling, and transcript stability.

Model size alone offers little guidance. A smaller decoder can outperform a larger one when its audio training and language coverage better match the workload. Conversely, a specialized model may struggle when conversations fall outside its intended domains.

The useful question is whether this architecture reduces errors on real Indian speech without creating new contextual mistakes. Sarvam’s benchmark results provide an encouraging signal, but customer audio will supply the stronger evidence.

Five Output Modes Remove Several Processing Steps

The second major feature is a single model that produces five distinct representations of the same recording.

Saaras V4 supports transcribe, verbatim, code-mixed, transliteration, and translation modes. These are more than cosmetic formatting options because each serves a different downstream workflow.

Transcribe mode produces text in the original language and normalizes elements such as numbers and dates. It also restores punctuation, making the result suitable for reading, indexing, and routine analysis.

Verbatim mode preserves filler words and spoken-number forms. Compliance teams, researchers, and conversation analysts may prefer this version because normalization can erase details about how something was said.

Code-mixed mode keeps Indian-language words in their native script while retaining spoken English words in Latin characters. This format reflects how many multilingual conversations are naturally written and reviewed.

Transliteration mode renders the utterance in Latin characters without translating its meaning. Sarvam describes this as a style commonly used in informal digital communication. It can help readers understand spoken Hindi or another supported language without reading its native script.

Translation mode converts supported Indic speech directly into English text. This can serve international support teams, reporting systems, and analysts who need a common output language.

Traditional pipelines often perform these jobs with separate components. One service recognizes speech, another normalizes the transcript, and a third translates or transliterates it. Every additional transformation can introduce errors or lose information.

Saaras V4 generates all five forms within one model. Sarvam argues that this design avoids cascading mistakes from separate preprocessing stages.

That claim has practical appeal. A customer-service platform might store verbatim speech for review, show normalized text to an agent, and send English output to an analytics system. One recognition model could support each destination.

The five modes also make speech more usable inside knowledge workflows. Meeting recordings and interviews become easier to search when teams can choose normalized or translated text. A searchable AI knowledge base can then connect transcripts with related notes and documents.

Still, one model does not eliminate every processing decision. Teams must determine which representation is authoritative, how to preserve the original recording, and whether translated text is suitable for sensitive decisions.

Translation mode only outputs English, according to Sarvam’s documentation. It does not provide arbitrary translation between every pair of supported languages.

Word-level timestamps are also unavailable in the standard response. The API provides phrase-level or chunk-level timing, while batch processing can add speaker-attributed transcripts.

These limits matter for subtitle editors, forensic review, and applications that need exact alignment. Competitors offering detailed word timing or transcript editing may remain preferable for those tasks.

The five modes nevertheless distinguish Saaras from a basic speech-to-text endpoint. Sarvam is treating transcript representation as part of recognition, which better reflects the demands of multilingual production systems.

Indic Coverage Extends Beyond the Largest Languages

The third distinguishing feature is consistent product support across all 22 scheduled Indian languages, including several low-resource languages.

Saaras V4 supports Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Assamese, Urdu, and Odia. It also covers Nepali, Konkani, Kashmiri, Sindhi, Sanskrit, Santali, Manipuri, Bodo, Maithili, and Dogri.

This breadth is significant because training resources are unevenly distributed. Major languages have larger speech corpora, more labeled recordings, and more commercial demand. Smaller languages often receive weaker recognition or no support.

Sarvam says V4 achieves state-of-the-art results across all 22 languages. That remains a company claim, although the company published its evaluation method and named the datasets used.

For ten Indian languages, Sarvam evaluated the model on Vistaar. The evaluation included Common Voice, FLEURS, Gramvaani, IndicTTS, Kathbath, noisy Kathbath recordings, and MUCS.

Sarvam reports both conventional word error rate and LLM-WER. Standard WER counts substitutions, insertions, and deletions between a prediction and reference transcript.

LLM-WER adds a semantic judgment step. It attempts to distinguish differences that change meaning from spelling or formatting variations that preserve the underlying content.

That distinction can be useful for Indian languages because written forms and normalization conventions vary. Two transcripts may communicate the same words while receiving a conventional WER penalty.

However, an evaluator based on another language model introduces judgment into the metric. Results can depend on the adjudicating model and its instructions. Conventional WER remains easier to reproduce, even when it overstates some harmless differences.

Language identification is another part of Sarvam’s coverage claim. Saaras V4 reportedly records a 2.9 percent identification error rate across the ten most widely spoken Indian languages. Sarvam reports 5.22 percent across all 22 languages.

Automatic detection removes the need to label every recording in advance. That is useful for shared call queues, public-service lines, and consumer applications serving multilingual audiences.

Yet language identification becomes more complicated during code-switching. Sarvam’s API returns the predominant language when multiple languages appear. Applications that need word-level language labels may require additional logic.

Global rivals frame multilingual coverage differently. ElevenLabs says its transcription system recognizes more than 90 languages and offers word-level timestamps, entity detection, and speaker diarization.

Deepgram’s multilingual models support real-time code-switching across a narrower collection of widely used languages. Its mature streaming services and voice-agent tooling create a different competitive advantage.

OpenAI describes GPT-4o Transcribe as improving word error rate and language recognition over earlier Whisper models. Its attraction lies partly in integration with a larger model platform.

Sarvam’s response is not to match every global language. It is to claim stronger performance across a specific, linguistically complex market.

That strategy pressures competitors where broad support can conceal inconsistent quality. Listing a language does not show how the system handles regional accents, mixed scripts, names, telephone compression, or informal speech.

It also leaves Sarvam exposed outside its core set. A multinational company covering European, African, and East Asian languages may prefer one broader provider, even if Saaras performs better on Indian calls.

The strongest case for V4 therefore appears in workloads where Indian-language accuracy matters enough to justify focused evaluation or a multi-provider architecture.

Keyterms and Streaming Target Difficult Production Audio

The fourth and fifth features address two recurring deployment problems: specialized vocabulary and conversational delay.

Keyterm prompting allows an application to provide names, products, locations, acronyms, or technical terms before transcription. The model then gives those terms greater consideration during decoding.

This matters because proper nouns are among the most damaging recognition errors. A transcript can preserve the surrounding sentence while misspelling the customer, medicine, company, or machine being discussed.

Sarvam’s REST and batch APIs accept up to 50 keyterms for Saaras V4. The streaming endpoint does not currently expose the same feature, according to its documentation.

Sarvam evaluated prompting with IndicContextEval, a benchmark associated with AI4Bharat. The company reports a 16.03 percent WER in the L5 setting, which supplies a native-script list of domain entities and the language.

That result suggests prompting helps, but it also illustrates a deployment requirement. Applications need a reliable way to select the right terms before each recording.

A hospital might supply clinician names, medicines, and procedures. A financial call center might provide fund names, securities, and customer entities. Sending a huge generic list could reduce the usefulness of the prompt.

Real-time applications face another constraint. A transcript must appear quickly enough for a voice agent to respond without awkward pauses.

Sarvam claims Saaras V4 can return its first streaming token in under 150 milliseconds. It also says the system can process multi-minute recordings within one second.

Those numbers should be treated as vendor-reported performance. End-to-end delay also includes network transport, audio buffering, endpoint detection, application processing, and the next model in a voice-agent pipeline.

The first token is not necessarily a stable transcript. Streaming recognition systems often revise earlier text as more audio arrives. Developers should measure both initial latency and the time required for a finalized segment.

Sarvam expanded V4’s real-time availability after launch. Its September changelog says the model became available through the newer Realtime API, although V3 remained the default at that point.

That detail deserves attention. Documentation described Saaras V4 as the latest model while still recommending V3 as the default. This suggests customers should not assume V4 is automatically the safest migration for every workload.

Low latency also does not guarantee good turn-taking. Voice activity detection must decide when a speaker has paused or finished. Aggressive settings can interrupt people, while cautious settings add noticeable delay.

Noisy telephone audio raises the stakes. Sarvam says V4 was designed for 8 kHz calls, clipping, interference, code-mixing, and dialect variation. These conditions frequently occur together in support and field-service recordings.

A meaningful test should combine them. Clean studio clips do not reveal what happens when a caller speaks quickly, changes language, mentions an unfamiliar name, and talks over another person.

Teams should also examine operational features around the model. Monitoring, regional processing, retention controls, failure handling, rate limits, and support can matter as much as a small accuracy advantage.

Saaras V4’s combination of keyterm prompting and streaming is promising because it targets real production failures. The competitive result depends on whether those capabilities remain reliable at scale.

The Benchmarks Need Independent Production Tests

Sarvam has published substantial evidence, but its broadest performance claims still require verification beyond company-run comparisons.

For English, Sarvam evaluated Saaras V4 across seven datasets. They cover meeting-room speech, podcasts, audiobooks, web video, studio recordings, difficult acoustics, financial calls, parliamentary speech, and Indian-accented English.

The named datasets are AMI, GigaSpeech, two LibriSpeech splits, SPGISpeech, VoxPopuli, and Svarah. Sarvam says V4 achieved the lowest average WER across that group.

The company reports using results from the Open ASR Leaderboard for six international datasets. It says normalization and scoring followed the leaderboard’s published code.

This is more informative than an unnamed internal benchmark. It identifies the datasets, scoring approach, and competing systems.

However, the comparisons are still assembled and presented by Sarvam. Model versions, API settings, prompt configuration, audio preprocessing, and release timing can influence results.

A benchmark average can also hide weaknesses. A model may lead overall while losing on a particular accent, channel, or speech style. Buyers should inspect dataset-level results that resemble their own traffic.

Indic evaluation presents additional complexity. Some competitors do not support every language, while others may require explicit language selection. Missing coverage and poor recognition are different limitations, even if both prevent successful deployment.

Saaras V4 also competes against products with different scopes. ElevenLabs emphasizes broad language support, detailed timestamps, entity detection, and editing. Deepgram focuses heavily on real-time transcription infrastructure. OpenAI integrates transcription with a wider AI platform.

Sarvam’s narrower language catalog can be an advantage when training effort is concentrated on Indian speech. It can also be a disadvantage for companies that want one global contract and one API.

Privacy and governance require separate review. Hosted speech systems process conversations that may contain personal, financial, or health information. Accuracy rankings do not answer where audio is stored, who can access it, or how retention works.

The model’s closed deployment limits external inspection. Researchers can evaluate API outputs, but they cannot fully audit training data, reproduce the model, or study errors without service access.

Saaras also lacks word-level timestamps in its documented response. Translate mode only produces English, and the real-time REST endpoint has a 30-second input limit. Batch processing is required for longer files.

These are manageable constraints, but they complicate any claim that one model replaces an entire transcription stack. Production systems still need storage, quality review, policy controls, and fallback behavior.

Sarvam’s strongest evidence concerns speech accuracy and Indic coverage. Its weaker evidence concerns long-term reliability under customer load, error behavior in sensitive domains, and operational advantages over established vendors.

The right response is not to dismiss the benchmarks. It is to reproduce them on representative recordings while tracking errors that matter to the application.

A customer-support team should weight account numbers and product names heavily. A meeting assistant should test overlapping speakers and attribution. A media workflow should examine punctuation, timing, and long-form stability.

Developers should also compare normalized and verbatim outputs. A low WER score can look different when the application must preserve every hesitation, number, or correction.

Three signals will determine whether the Sarvam AI Saaras V4 launch changes the market.

First, independent evaluations must reproduce its advantages across noisy Indian speech and global English. Consistent third-party results would strengthen Sarvam’s central accuracy claim. Large gaps would weaken it.

Second, production adoption must extend beyond demonstrations. Meaningful evidence would include sustained use in call centers, voice agents, meeting systems, and multilingual media workflows.

Third, competitors will respond through better Indic coverage, code-switching, or regional deployment options. A visible response from larger vendors would confirm that Sarvam has created commercial pressure.

For developers, the immediate action is straightforward. Build a private evaluation set from consented, representative recordings, including difficult accents and poor audio. Score important entities separately from overall WER.

Enterprise buyers should test latency, transcript stability, speaker separation, and data governance alongside raw accuracy. Knowledge workers should watch whether meeting and interview tools adopt Saaras without reducing editing control.

Sarvam has presented a serious technical case for specialized multilingual speech recognition. Now Saaras V4 must show that its benchmark lead survives the untidy conditions of real conversations.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page