text-to-speech api built for real-time conversations

text-to-speech api built for real-time conversations

built to be versatile and human

built to be versatile and human

built to be versatile and human

voice agents, ivr, customer support, dubbing. hear silk across every use case."

type to insert emotion tags
type to insert emotion tags

the models

the models

the models

mulberry

velvet & calm

• Starts speaking in 162 ms, faster than a blink.

• Creates the voice you describe instantly.

muga

warm & conversational

• 24 kHz studio-quality audio

• 6 built-in emotions (Happy, Sad, Angry, Excited, Whisper, Neutral)

spider

our most advanced model yet

mulberry

warm & conversational

• Starts speaking in 162 ms, faster than a blink.

• Creates the voice you describe instantly.

muga

velvet & calm

• 24 kHz studio-quality audio

• 6 built-in emotions (Happy, Sad, Angry, Excited, Whisper, Neutral)

spider

our most advanced model yet

spider

our most advanced model yet

mulberry

velvet & calm

• Starts speaking in 162 ms, faster than a blink.

• Creates the voice you describe instantly.

muga

warm & conversational

• 24 kHz studio-quality audio

• 6 built-in emotions (Happy, Sad, Angry, Excited, Whisper, Neutral)

spider

our most advanced model yet

built for natural, real-time voice applications

built for natural, real-time voice applications

built for natural, real-time voice applications

mulberry responds in 162ms, the fastest ttfb on the market, at ₹0.40/min (~$0.005/min).

mulberry responds in 162ms, the fastest ttfb on the market, at ₹0.40/min (~$0.005/min).

fast enough to respond

  1. fast enough to respond

speech starts coming back as little as 162ms. the conversation does not have to wait.

describe the voice. skip the catalogue.

  1. describe the voice. skip the catalogue.

write the age, accent, pitch, pace, timbre and character you need.

hindi. english. hinglish. same person.

  1. hindi. english. hinglish. same person.

change the language halfway through without changing the voice behind it.

pay for what gets spoken.

  1. pay for what gets spoken.

top up credits once. burn them down request by request.

implement silk in your project instantly

get a key. send a request. receive 24 khz audio.

from rumikai import Rumik

client = Rumik()

# muga: expressive, steer with an inline tone tag
audio = client.speech.create(text="[happy] Namaste! Kaise hain aap?", model="muga")
audio.save("hello.wav")

# mulberry: faster, steer with a description + preset speaker
audio = client.speech.create(
    text="Hi there, how can I help you today?",
    model="mulberry",
    description="a warm 30s female voice, conversational pacing",
    speaker="speaker_2",
)
audio.save("greeting.wav")
from rumikai import Rumik

client = Rumik()

# muga: expressive, steer with an 
inline tone tag
audio = client.speech.create
(text="[happy] Namaste! Kaise hain 
aap?", model="muga")
audio.save("hello.wav")

# mulberry: faster, steer with a 
description + preset speaker
audio = client.speech.create(
    text="Hi there, 
how can I help you today?",
    model="mulberry",
    description="a warm 30s female 
voice, conversational pacing",
    speaker="speaker_2",
)
audio.save("greeting.wav")
from rumikai import Rumik

client = Rumik()

# muga: expressive, steer with an inline tone tag
audio = client.speech.create(text="[happy] Namaste! 
Kaise hain aap?", model="muga")
audio.save("hello.wav")

# mulberry: faster, steer with a description + preset speaker
audio = client.speech.create(
    text="Hi there, how can I help you today?",
    model="mulberry",
    description="a warm 30s female voice, conversational pacing",
    speaker="speaker_2",
)
audio.save("greeting.wav")

silk keeps the voice intact even through language changes

silk keeps the voice intact even through language changes

silk keeps the voice intact even through language changes

narrator
description: a male 40s british voice, low pitch, gravelly timbre, slow pacing, neutral, formal register, like a dramatic narrator. text: the door creaked open. nobody was there. and yet, something watched.
0:00 / 0:00
podcast host
description: a female 30s hindi voice, normal pitch, smooth timbre, conversational pacing, energetic, casual register, like a podcast host. text: आज का episode थोड़ा अलग है। एक minute के लिए सीधा बैठ जाओ।
0:00 / 0:00
support
description: a female 30s indian voice, normal pitch, warm timbre, conversational pacing, neutral, neutral register, like a customer support agent. text: मैं आपकी help के लिए यहाँ हूँ। एक minute, मैं check करती हूँ।
0:00 / 0:00
streamer
description: a male 20s american voice, high pitch, smooth timbre, very fast pacing, excited, casual register, like a streamer reacting live. text: oh my god, did you see that play? that was insane!
0:00 / 0:00

meet the teams already speaking through silk

meet the teams already speaking through silk

meet the teams already speaking through silk

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

100 + hours in 2 days

100 + hours in 2 days

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

100 + hours in 2 days

curvet ai

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snap tv

Image (9) (no background)

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

monk learning

faqs

faqs

faqs

silk is rumik's expressive, multilingual text-to-speech api. send text, get natural, human-sounding speech back in real time.
silk mulberry starts at $0.005 per 1k characters (₹0.40/min in india). pay for what you need with credit packs — no subscriptions.
english, hindi, hinglish, tamil, telugu, bengali, marathi and more — with natural mid-sentence switching between them.
yes. websocket streaming with low time-to-first-byte, fast enough for live conversations.
yes. silk is built for voice agents, calls and conversational experiences, with pipecat and livekit integrations.
describe the voice you want — age, accent, pitch, tone and emotion — instead of picking from a fixed catalogue.
credits are deducted per token generated. track usage anytime from the playground.
generation pauses until you top up. your account and api keys stay active.
limits depend on your plan. contact the team if you need higher concurrency for production workloads.
yes, commercial usage is allowed on paid plans.
a few lines of code. rest and websocket apis with python and node examples in the docs.

start building with rumik's tts api

start building with rumik's tts api

join us

we are a close-knit group of researchers, engineers, and designers working on the hardest problems in ai.