arrow_back Back to AIFC
S
pending Claude

Orato ASR/TTS Fact-Check: Numbers Check Out, Access Claim Doesn't

Grounded / Real Inflated / Uruttu
70% real
30% uruttu
article Original Content
We just shipped Orato ASR and TTS: Hindi/Hinglish speech models built for real-time voice agents.
Over the past few months, we've been building speech models for Orato, along with Anand Dubey: real-time voice agents for Indian calling workflows like customer support, appointment booking, insurance, and sales.
𝗪𝗵𝗮𝘁 𝘄𝗲 𝗯𝘂𝗶𝗹𝘁:
• Orato ASR: a full-parameter fine-tune of Qwen3-ASR-0.6B on ~1,000 hours of Hindi/English/Hinglish calling-domain audio
• Orato TTS: a Hindi/Hinglish text-to-speech model built on IndicF5-TTS, with multi-speaker and voice-clone support
Fine-tuning improved WER across every benchmark we tested against the base model, with the biggest gains exactly where it matters most for a voice agent: noisy, accented, telephony-style speech (Lahaja: 25% relative improvement, Kathbath: 24.6%), not just clean studio audio nobody actually calls on.
Both models are live on Hugging Face, access is auto-approved, so go ahead and try them:
🔗 ASR: https://lnkd.in/g-3VDDiV
🔗 TTS: https://lnkd.in/gYsk9nJQ
This is just the start. Everything we learned getting here is already shaping what we're building next.
Would genuinely love to hear your feedback if you give these a try.
verified Validated Content

Checked both Hugging Face model cards directly. Here's what holds up and what doesn't:

Confirmed accurate:

  • Base models: Orato ASR is a full-parameter fine-tune of Qwen/Qwen3-ASR-0.6B (audio tower frozen), trained on ~1,000 hours of Hindi/English/Hinglish calling-domain audio — matches the model card exactly. Orato TTS is a fine-tune of ai4bharat/IndicF5 (F5-TTS architecture), on ~194 hours of Hindi/Hinglish calling-domain speech, with multi-speaker via reference audio and voice-clone support — also matches.
  • WER numbers: Kathbath improvement is exactly 24.6% (15.24% → 11.49%) as stated. Lahaja is actually 25.5%, not 25% (25.09% → 18.68%) — you rounded down slightly, but directionally correct.
  • "Improved on every benchmark tested": True per the card — Kathbath, Gramvaani, Lahaja, and Fleurs all show gains (3.6% to 25.5% relative).

One naming nit:

  • The TTS base is called "IndicF5" on Hugging Face, not "IndicF5-TTS" — minor, but worth matching the actual model name if people go looking for it.

Worth double-checking before you post:

  • "Access is auto-approved, go ahead and try them" — both model cards currently show a gated-access screen ("You need to agree to share your contact information to access this model") and explicitly state "Access is currently private" / "Released for internal/beta evaluation only" pending commercial-use and consent review on the training data. That's a real contradiction with "go ahead and try them" as currently written — either the gating needs to be flipped to public before you publish this, or the copy needs to say "request access" rather than implying open access.
  • Anand Dubey's role — I could only find him listed as an "SDE Intern" at Orato in a third-party contact database (not something I'd treat as authoritative), while a different person (Anshuman Tomar) shows up as "Co-Founder." Worth confirming his actual title/credit is what you want on a public post, since I couldn't independently verify it.

Everything else — architecture, training data volume, benchmark methodology (AI4Bharat Vistaar/Lahaja suites, third-party WER for Sarvam/Deepgram/ElevenLabs) — checks out against the source model cards.