<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Oshara Blog</title><description>Engineering notes from the Oshara team — from-the-trenches accounts of building voice AI.</description><link>https://blog.oshara.ai/</link><item><title>Building an Agent, Step by Step</title><link>https://blog.oshara.ai/building-an-agent-step-by-step/</link><guid isPermaLink="true">https://blog.oshara.ai/building-an-agent-step-by-step/</guid><description>A complete walkthrough of creating a voice agent on Oshara.ai — from the My Agents page all the way to a deployed, embeddable widget.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Fine-Tuning IndicConformer for Nepali Speech Recognition</title><link>https://blog.oshara.ai/nepali-stt-indicconformer-finetune/</link><guid isPermaLink="true">https://blog.oshara.ai/nepali-stt-indicconformer-finetune/</guid><description>We benchmarked AI4Bharat&apos;s IndicConformer against six other Nepali STT systems, then fine-tuned it further — hitting a tokenizer bug along the way that taught us something useful about training data hygiene.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item><item><title>A Nepali Voice for XTTS v2</title><link>https://blog.oshara.ai/nepali-xtts-paper-results/</link><guid isPermaLink="true">https://blog.oshara.ai/nepali-xtts-paper-results/</guid><description>Nepali has no open, zero-shot text-to-speech system. We built one by fine-tuning XTTS v2 — reusing its existing Hindi tokenizer instead of adding a single new vocabulary token — and it&apos;s now the best open, self-hostable Nepali voice we could measure.</description><pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Teaching a Duplex Voice Model to Call Tools</title><link>https://blog.oshara.ai/teaching-a-duplex-voice-model-to-call-tools/</link><guid isPermaLink="true">https://blog.oshara.ai/teaching-a-duplex-voice-model-to-call-tools/</guid><description>PersonaPLEX speaks and listens at the same time, but it can&apos;t look anything up. This is how we grafted a tool-calling action stream onto a full-duplex audio model — the architecture changes, the four training attempts, and why it kept searching for the wrong thing.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Teaching a Machine to Speak Nepali</title><link>https://blog.oshara.ai/teaching-a-machine-to-speak-nepali/</link><guid isPermaLink="true">https://blog.oshara.ai/teaching-a-machine-to-speak-nepali/</guid><description>A from-the-trenches account of fine-tuning Chatterbox v3, a voice-cloning TTS model, for Nepali — the tokenizer bugs, the phantom sentences, the overfitting hunt, and the audio that came out the other side.</description><pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Teaching PersonaPLEX to Recite Shakespeare</title><link>https://blog.oshara.ai/teaching-personaplex-to-recite-shakespeare/</link><guid isPermaLink="true">https://blog.oshara.ai/teaching-personaplex-to-recite-shakespeare/</guid><description>How we injected new knowledge into a full-duplex speech model with LoRA — turning a Shakespeare play into spoken question-answer data, and fine-tuning PersonaPLEX in document mode to recite grounded lines and attributions out loud.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Teaching XTTS v2 to Speak Nepali</title><link>https://blog.oshara.ai/teaching-xtts-v2-to-speak-nepali/</link><guid isPermaLink="true">https://blog.oshara.ai/teaching-xtts-v2-to-speak-nepali/</guid><description>How we fine-tuned XTTS v2, a zero-shot voice-cloning TTS model, into a fluent Nepali narrator: the architecture, why epoch 10 beat epoch 30, how it stacks up against a different TTS model, and the text-pipeline fixes that stopped it from babbling.</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Two LLMs Are Better Than One: How We Build Low-Latency Voice Agents</title><link>https://blog.oshara.ai/two-llms-are-better-than-one/</link><guid isPermaLink="true">https://blog.oshara.ai/two-llms-are-better-than-one/</guid><description>The awkward pause in voice AI is a timing problem, not an intelligence problem. This is the two-parallel-LLM architecture we use to fix it — a fast model that buys presence, a smart model that carries the answer — plus a tour of how a real-time voice agent is actually assembled.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item></channel></rss>