manipuri text to speech that sounds human

create realistic manipuri speech that carries meitei words with clarity, tone, and local feel.

built to be versatile and human

built to be versatile and human

built to be versatile and human

voice agents, ivr, customer support, dubbing. hear silk across every use case.

manipuri (meitei)type ‹ to insert emotion tags
mulberry 1.6
manipuri (meitei)type ‹ to insert emotion tags
mulberry 1.6

lower latency for real-time speech


in live speech, latency is the first thing users feel. mulberry returns first audio in 162 ms on the published benchmark

lower is better
162
mulberry
188
cartesia sonic-3
264
elevenlabs turbo v2.5
288
elevenlabs flash v2.5
313
deepgram aura-2
337
rime mist-v3
450
rime arcana
1,232
elevenlabs multilingual
2,295
openai tts-1-hd

generate manipuri tts in three steps

generate manipuri tts in three steps

generate manipuri tts in three steps

add manipuri (meitei) text to speech with silk api and the official python sdk.

silk api

silk api

silk api

  1. Create your key

  1. Create your key

  1. Create your key

grab a key from api keys in the playground. name it whatever you like.

  1. choose a voice model

  1. choose a voice model

  1. choose a model

mulberry-1.6 covers all 22 languages. muga adds tone tags for hindi-english lines.

  1. call the api

  1. call the api

send the model and text, get a wav back. add audio_format for other formats.

from rumikai import Rumik

client = Rumik()  # reads RUMIK_API_KEY
audio = client.speech.create(
    text="Hello, what can I do for you?",
    model="mulberry-1.6",
    description="professional, Indian English accent, steady pace",
)
audio.save("speech.wav")
import requests

url = "[https://silk-api.rumik.ai/v1/tts](https://silk-api.rumik.ai/v1/tts)"

payload = {
    "model": "mulberry-1.6",
    "text": "Hello, what can I do for you?", 
    "description": "professional, British accent, steady pace", 
    "speaker": "aisha",
}
headers = {"Authorization": "Bearer <token>"}

response = requests.post(url, json=payload, headers=headers)

with open("speech.wav", "wb") as f:
    f.write(response.content)
import requests

url = "[https://silk-api.rumik.ai/v1/tts](https://silk-api.rumik.ai/v1/tts)"

payload = {
    "model": "mulberry-1.6",
    "text": "Hello, what can I do for you?", 
    "description": "professional, British accent, steady pace", 
    "speaker": "aisha",
}
headers = {"Authorization": "Bearer <token>"}

response = requests.post(url, json=payload, headers=headers)

with open("speech.wav", "wb") as f:
    f.write(response.content)

silk playground

silk playground

silk playground

  1. paste your text

    enter the english sentence you want to hear. keep punctuation, numbers, and mixed-language phrases exactly as you want them spoken.

  2. choose a model

    select the silk model for english, then adjust the available settings for the voice and delivery you need.

  3. generate the audio

    listen to the result, refine the text or settings when needed, and download the audio when the line is ready.

what makes manipuri text to speech hard?

manipuri speech has its own rules for pronunciation, rhythm, and written text. the voice has to carry the sentence as speech, with the right pauses, emphasis, and word shape.

  1. script and spelling

    turns mayek letters, apun iyek, and lonsum finals into clear speech without flattening the line.

  2. pronunciation details

    keeps lexical tones, final stops, and manipuri names clear while the voice holds its rhythm.

  3. numbers and time

    reads amounts, dates, and times the way people say them aloud.

  4. mixed-language lines

    keeps the same speaker through manipuri-english phrases, local names, and product words.

lower latency for real-time speech

lower latency for real-time speech

in live speech, latency is the first thing users feel. mulberry returns first audio in 162 ms on the published benchmark

lower is better
162
mulberry
188
cartesia sonic-3
264
elevenlabs turbo v2.5
288
elevenlabs flash v2.5
313
deepgram aura-2
337
rime mist-v3
450
rime arcana
1,232
elevenlabs multilingual
2,295
openai tts-1-hd

what makes manipuri text to speech hard?

what makes manipuri text to speech hard?

manipuri speech has its own rules for pronunciation, rhythm, and written text. the voice has to carry the sentence as speech, with the right pauses, emphasis, and word shape.

  1. script and spelling

    turns mayek letters, apun iyek, and lonsum finals into clear speech without flattening the line.

  2. pronunciation details

    keeps lexical tones, final stops, and manipuri names clear while the voice holds its rhythm.

  3. numbers and time

    reads amounts, dates, and times the way people say them aloud.

  4. mixed-language lines

    keeps the same speaker through manipuri-english phrases, local names, and product words.

silk keeps the voice intact even through language changes

silk keeps the voice intact even through language changes

silk keeps the voice intact even through language changes

narrator
description: a male 40s british voice, low pitch, gravelly timbre, slow pacing, neutral, formal register, like a dramatic narrator. text: the door creaked open. nobody was there. and yet, something watched.
0:00 / 0:00
podcast host
description: a female 30s hindi voice, normal pitch, smooth timbre, conversational pacing, energetic, casual register, like a podcast host. text: आज का episode थोड़ा अलग है। एक minute के लिए सीधा बैठ जाओ।
0:00 / 0:00
support
description: a female 30s indian voice, normal pitch, warm timbre, conversational pacing, neutral, neutral register, like a customer support agent. text: मैं आपकी help के लिए यहाँ हूँ। एक minute, मैं check करती हूँ।
0:00 / 0:00
streamer
description: a male 20s american voice, high pitch, smooth timbre, very fast pacing, excited, casual register, like a streamer reacting live. text: oh my god, did you see that play? that was insane!
0:00 / 0:00

generate tts in 22 languages

build voices for indian languages, and accents from the same speech stack

languages
englishhinditamiltelugubengalimarathi
showing 6 of 22 supported languages.

generate tts in 22 languages

generate tts in 22 languages

build voices for indian languages, and accents from the same speech stack

where teams use manipuri voice

where teams use manipuri voice

manipuri voice belongs in moments that need clarity, recall, and a human pace.

voice overs
turn scripts into audio for videos, product demos, presentations, ads, and social posts.
e-learning
create spoken lessons, onboarding modules, and mobile-first education content from written scripts.
customer support
create call flows, order updates, and support replies for real-time conversations.
products and apps
use voice in workflow tools, companion products, reading apps, and internal tools.
audiobooks and narration
turn chapters, explainers, and story scripts into narration people can listen to anywhere.
game dialogue
create character lines for games, stories, companion apps, and scripted scenes.

meet the teams already speaking through silk

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

100 + hours in 2 days

curvet ai

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snap tv

Image (9) (no background)

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

monk learning

meet the teams already speaking through silk

meet the teams already speaking through silk

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

100 + hours in 2 days

100 + hours in 2 days

curvet ai

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snap tv

Image (9) (no background)

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

monk learning

frequently asked questions

frequently asked questions

frequently asked questions

paste your manipuri (meitei) text into the rumik playground, choose a model, and generate the audio. you can use the playground for quick testing and then move the same script into silk api when you need it inside a product.
yes. rumik gives you free playground credits to test manipuri (meitei) tts with your own text before moving it into a product or workflow.
install the official python sdk with pip install rumik-ai and then store your api key. send manipuri (meitei) text through the rumikai client with a model and voice description.
silk mulberry 1.6 is the best fit for manipuri (meitei) tts when you need fast multilingual speech, voice descriptions, and low-latency speech.
yes. rumik supports manipuri (meitei)-english code-switching so a mixed line can keep the same speaker style across the switch.
yes. you can use rumik manipuri (meitei) tts in commercial products, apps, and content, subject to your plan and rumik's usage terms. check the current terms before shipping paid or high-volume usage.

start building with rumik's tts api

start building with rumik's tts api

join us

we are a close-knit group of researchers, engineers, and designers working on the hardest problems in ai.