realistic text to speech for every voice experience

realistic text to speech for every voice experience

create realistic ai voices for voice overs, apps, agents, lessons, and read-aloud experiences from one text to speech api

before you try silk
•what: text to speech api and ai voice generator for products, content, and live agents.
•voices: describe the speaker, accent, pace, emotion, and role.
•features: code-switching, real-time streaming, voice controls, and number and date handling.
•latency: mulberry starts speaking in 162 ms.
•pricing: silk mulberry starts at $0.005 per 1k characters. elevenlabs starts at $0.05.
•try: test voices in the playground before using the api.

built to be versatile and human

built to be versatile and human

built to be versatile and human

voice agents, ivr, customer support, dubbing. hear silk across every use case.

type ‹ to insert emotion tags
type ‹ to insert emotion tags

why realistic text to speech is hard?

why realistic text to speech is hard?

why realistic text to speech is hard?

text is written for eyes before it is spoken aloud. a script can carry dates, prices, abbreviations, product names, punctuation, and sentence fragments that a person would smooth out while reading.

the hard part is not pronunciation. the voice has to know on which word to stress, where to pause, and how much energy the line needs. the same words can sound like a warning, confirmation, a lesson, or a story depending on the delivery. longer audio makes the problem easier to understand. a short line can pass with a flat voice. a lesson, product walkthrough, or chapter needs steadier pacing and a voice that does not change after the first few sentences.

real use also brings mixed inputs. names from one language appear inside another. numbers need to be spoken the right way. borrowed words and brand names should not make the speaker sound like a different person. that is why realistic text to speech is less about reading words correctly and more about turning written text into speech someone can keep listening to.

text is written for eyes before it is spoken aloud. a script can carry dates, prices, abbreviations, product names, punctuation, and sentence fragments that a person would smooth out while reading.

the hard part is not pronunciation. the voice has to know on which word to stress, where to pause, and how much energy the line needs. the same words can sound like a warning, confirmation, a lesson, or a story depending on the delivery. longer audio makes the problem easier to understand. a short line can pass with a flat voice. a lesson, product walkthrough, or chapter needs steadier pacing and a voice that does not change after the first few sentences.

real use also brings mixed inputs. names from one language appear inside another. numbers need to be spoken the right way. borrowed words and brand names should not make the speaker sound like a different person. that is why realistic text to speech is less about reading words correctly and more about turning written text into speech someone can keep listening to.

text is written for eyes before it is spoken aloud. a script can carry dates, prices, abbreviations, product names, punctuation, and sentence fragments that a person would smooth out while reading.

the hard part is not pronunciation. the voice has to know on which word to stress, where to pause, and how much energy the line needs. the same words can sound like a warning, confirmation, a lesson, or a story depending on the delivery. longer audio makes the problem easier to understand. a short line can pass with a flat voice. a lesson, product walkthrough, or chapter needs steadier pacing and a voice that does not change after the first few sentences.

real use also brings mixed inputs. names from one language appear inside another. numbers need to be spoken the right way. borrowed words and brand names should not make the speaker sound like a different person. that is why realistic text to speech is less about reading words correctly and more about turning written text into speech someone can keep listening to.

how rumik silk makes generated speech sound natural

how rumik silk makes generated speech sound natural

how rumik silk makes generated speech sound natural

silk does not treat text as a flat string to read from left to right. it prepares the input, chooses the right model behavior, and uses voice controls before the audio is generated. that matters because the same sentence can fail in different ways. a number can be read wrong. a name can sound out of place. a voice can lose its tone halfway through a longer passage.

silk does not treat text as a flat string to read from left to right. it prepares the input, chooses the right model behavior, and uses voice controls before the audio is generated. that matters because the same sentence can fail in different ways. a number can be read wrong. a name can sound out of place. a voice can lose its tone halfway through a longer passage.

silk does not treat text as a flat string to read from left to right. it prepares the input, chooses the right model behavior, and uses voice controls before the audio is generated. that matters because the same sentence can fail in different ways. a number can be read wrong. a name can sound out of place. a voice can lose its tone halfway through a longer passage.

text normalization

written text needs cleanup before it can become speech. dates, amounts, ids, abbreviations, and punctuation have to be interpreted the way a person would say them aloud. for example, "Your refund of $1,000 was processed on August 8, 2026." silk can read the amount as one thousand dollars and the date as august eighth, two thousand twenty six. normalization runs by default, but it can be turned off when your input is already written exactly as it should be spoken.

written text needs cleanup before it can become speech. dates, amounts, ids, abbreviations, and punctuation have to be interpreted the way a person would say them aloud. for example, "Your refund of $1,000 was processed on August 8, 2026." silk can read the amount as one thousand dollars and the date as august eighth, two thousand twenty six. normalization runs by default, but it can be turned off when your input is already written exactly as it should be spoken.

written text needs cleanup before it can become speech. dates, amounts, ids, abbreviations, and punctuation have to be interpreted the way a person would say them aloud. for example, "Your refund of $1,000 was processed on August 8, 2026." silk can read the amount as one thousand dollars and the date as august eighth, two thousand twenty six. normalization runs by default, but it can be turned off when your input is already written exactly as it should be spoken.

speech tokens

silk generates speech through model-predicted audio units rather than stitched recordings. the model first turns the text and voice direction into compact speech tokens. a neural audio codec then decodes those tokens into the final waveform. that is what lets silk produce speech progressively. in streaming, the first audio can start before the model has finished generating the full line.

silk generates speech through model-predicted audio units rather than stitched recordings. the model first turns the text and voice direction into compact speech tokens. a neural audio codec then decodes those tokens into the final waveform. that is what lets silk produce speech progressively. in streaming, the first audio can start before the model has finished generating the full line.

silk generates speech through model-predicted audio units rather than stitched recordings. the model first turns the text and voice direction into compact speech tokens. a neural audio codec then decodes those tokens into the final waveform. that is what lets silk produce speech progressively. in streaming, the first audio can start before the model has finished generating the full line.

model choice

silk is a model family, not one voice with one behavior.

  1. silk mulberry 1.6 is the description-driven model. it can be used for generated speech in products, agents, content, and learning platforms.

  2. silk muga 1 is the expressive model. it uses tone tags and inline events for stronger control over emotion and vocal behavior inside the line.

silk is a model family, not one voice with one behavior.

  1. silk mulberry 1.6 is the description-driven model. it can be used for generated speech in products, agents, content, and learning platforms.

  2. silk muga 1 is the expressive model. it uses tone tags and inline events for stronger control over emotion and vocal behavior inside the line.

silk is a model family, not one voice with one behavior.

  1. silk mulberry 1.6 is the description-driven model. it can be used for generated speech in products, agents, content, and learning platforms.

  2. silk muga 1 is the expressive model. it uses tone tags and inline events for stronger control over emotion and vocal behavior inside the line.

voice descriptions in silk mulberry

mulberry lets you steer the delivery with a short description. the description controls style, accent, and pace. for example, "professional, steady pace" if you leave part of the description out, silk fills it with a default. the style becomes professional, the accent follows the script of the text, and the pace defaults to fast.

mulberry 1.6 also supports four named speakers: ira, aisha, siya, and zoya. if you do not send a speaker, silk uses ira. write the text in the language's own script, then use the description when you want to name the accent or adjust how the line is delivered.

mulberry lets you steer the delivery with a short description. the description controls style, accent, and pace. for example, "professional, steady pace" if you leave part of the description out, silk fills it with a default. the style becomes professional, the accent follows the script of the text, and the pace defaults to fast.

mulberry 1.6 also supports four named speakers: ira, aisha, siya, and zoya. if you do not send a speaker, silk uses ira. write the text in the language's own script, then use the description when you want to name the accent or adjust how the line is delivered.

mulberry lets you steer the delivery with a short description. the description controls style, accent, and pace. for example, "professional, steady pace" if you leave part of the description out, silk fills it with a default. the style becomes professional, the accent follows the script of the text, and the pace defaults to fast.

mulberry 1.6 also supports four named speakers: ira, aisha, siya, and zoya. if you do not send a speaker, silk uses ira. write the text in the language's own script, then use the description when you want to name the accent or adjust how the line is delivered.

tone tags and inline events in silk muga

silk muga gives control inside the script itself. a tone tag such as [happy], [sad], or [neutral] changes how the line is delivered. an inline event such as <laugh> places a vocal action at a specific point in the sentence. that is useful when the audio needs to sound performed rather than read. dialogue, companion voices, story narration, and character lines need more control than a plain voice description can carry.

silk muga gives control inside the script itself. a tone tag such as [happy], [sad], or [neutral] changes how the line is delivered. an inline event such as <laugh> places a vocal action at a specific point in the sentence. that is useful when the audio needs to sound performed rather than read. dialogue, companion voices, story narration, and character lines need more control than a plain voice description can carry.

silk muga gives control inside the script itself. a tone tag such as [happy], [sad], or [neutral] changes how the line is delivered. an inline event such as <laugh> places a vocal action at a specific point in the sentence. that is useful when the audio needs to sound performed rather than read. dialogue, companion voices, story narration, and character lines need more control than a plain voice description can carry.

real-time streaming

silk can stream generated speech over websocket. silk mulberry starts speaking in 162 ms on the published latency benchmark, so live agents and interactive products can begin audio before the full response has finished generating.

silk can stream generated speech over websocket. silk mulberry starts speaking in 162 ms on the published latency benchmark, so live agents and interactive products can begin audio before the full response has finished generating.

silk can stream generated speech over websocket. silk mulberry starts speaking in 162 ms on the published latency benchmark, so live agents and interactive products can begin audio before the full response has finished generating.

how to generate text to speech with rumik

how to generate text to speech with rumik

how to generate text to speech with rumik

build it into your app with the silk api in three steps, or skip the code and grab your audio from the playground

silk api

silk api

silk api

  1. Create an API key

  1. Create your key

  1. Create an API key

grab a key from api keys in the playground. name it whatever you like.

  1. choose a voice model

  1. choose a model

  1. choose a voice model

mulberry for voice descriptions and real-time streaming. muga adds tone tags and inline events.

  1. call the api

  1. call the api

send the model and text, get a wav back. add audio_format for other formats.

from rumikai import Rumik

client = Rumik()  # reads RUMIK_API_KEY
audio = client.speech.create(
    text="Hello, what can I do for you?",
    model="mulberry-1.6",
    description="professional, Indian English accent, steady pace",
)
audio.save("speech.wav")
import requests

url = "[https://silk-api.rumik.ai/v1/tts](https://silk-api.rumik.ai/v1/tts)"

payload = {
    "model": "mulberry-1.6",
    "text": "Hello, what can I do for you?", 
    "description": "professional, British accent, steady pace", 
    "speaker": "aisha",
}
headers = {"Authorization": "Bearer <token>"}

response = requests.post(url, json=payload, headers=headers)

with open("speech.wav", "wb") as f:
    f.write(response.content)
import requests

url = "[https://silk-api.rumik.ai/v1/tts](https://silk-api.rumik.ai/v1/tts)"

payload = {
    "model": "mulberry-1.6",
    "text": "Hello, what can I do for you?", 
    "description": "professional, British accent, steady pace", 
    "speaker": "aisha",
}
headers = {"Authorization": "Bearer <token>"}

response = requests.post(url, json=payload, headers=headers)

with open("speech.wav", "wb") as f:
    f.write(response.content)

silk playground

silk playground

silk playground

  1. paste your text

    enter the english sentence you want to hear. keep punctuation, numbers, and mixed-language phrases exactly as you want them spoken.

  2. choose a model

    select the silk model for english, then adjust the available settings for the voice and delivery you need.

  3. generate the audio

    listen to the result, refine the text or settings when needed, and download the audio when the line is ready.

one voice model family for content, products, and agents

one voice model family for content, products, and agents

one voice model family for content, products, and agents

silk gives teams one speech stack for narration, product audio, live replies, expressive lines, and research workflows

mulberry

• starts speaking in 162 ms, faster than a blink.

• creates the voice you describe instantly.

muga

• 24 kHz studio-quality audio

• 6 built-in emotions (happy, sad, angry, excited, whisper, neutral)

spider

our most advanced model yet

mulberry

warm & conversational

• Starts speaking in 162 ms, faster than a blink.

• Creates the voice you describe instantly.

muga

velvet & calm

• 24 kHz studio-quality audio

• 6 built-in emotions (Happy, Sad, Angry, Excited, Whisper, Neutral)

spider

our most advanced model yet

spider

our most advanced model yet

mulberry

velvet & calm

• Starts speaking in 162 ms, faster than a blink.

• Creates the voice you describe instantly.

muga

warm & conversational

• 24 kHz studio-quality audio

• 6 built-in emotions (Happy, Sad, Angry, Excited, Whisper, Neutral)

spider

our most advanced model yet

generate speech in 22 languages with accent control

generate speech in 22 languages with accent control

generate speech in 22 languages with accent control

create speech for multilingual scripts, accent-directed voices, and different delivery styles from the same api

built for more than one kind of generated speech

built for more than one kind of generated speech

built for more than one kind of generated speech

voice for the places your audience already listens, from screens and calls to courses and stories

voice overs
turn scripts into audio for videos, product demos, presentations, ads, and social posts.
e-learning
create spoken lessons, onboarding modules, and mobile-first education content from written scripts.
customer support
create call flows, order updates, and support replies for real-time conversations.
products and apps
use voice in workflow tools, companion products, reading apps, and internal tools.
audiobooks and narration
turn chapters, explainers, and story scripts into narration people can listen to anywhere.
game dialogue
create character lines for games, stories, companion apps, and scripted scenes.

low-latency speech for real-time voice

low-latency speech for real-time voice

low-latency speech for real-time voice

silk mulberry starts speaking in 162 ms so real-time voice products can feel responsive from the first reply

lower is better
162
mulberry
188
cartesia sonic-3
264
elevenlabs turbo v2.5
288
elevenlabs flash v2.5
313
deepgram aura-2
337
rime mist-v3
450
rime arcana
1,232
elevenlabs multilingual
2,295
openai tts-1-hd

how rumik compares with other text to speech apis

how rumik compares with other text to speech apis

how rumik compares with other text to speech apis

compare rumik with elevenlabs, deepgram, murf, and cartesia on price, streaming, voice control, and latency

text to speech api pricing comparison: rumik silk vs elevenlabs, deepgram, murf and cartesia
featurerumikelevenlabsdeepgrammurfcartesia
entry tts price$0.005/ 1k chars$0.05/ 1k chars$0.015/ 1k charsfrom $0.01/ 1k charsfrom $5/ mo credits
billing modelcharacters + minutescharacterscharacterscharacterscredits
pay as you goyesyesyesyescredit plans
unlimited tts plan$199/ monononono
real-time streamingyes162 msyesyesyesyes
featurerumikelevenlabs
entry tts price$0.005/ 1k chars$0.05/ 1k chars
billing modelcharacters + minutescharacters
pay as you goyesyes
unlimited tts plan$199/ mono
real-time streamingyes162 msyes

prices from each provider's public pricing page, in usd.

meet the teams already speaking through silk

meet the teams already speaking through silk

meet the teams already speaking through silk

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

100 + hours in 2 days

100 + hours in 2 days

curvet put mulberry and muga directly inside its ai workflow canvas. within two days, teams generated voices across education, design, crm and enterprise workflows.

100 + hours in 2 days

curvet ai

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

snap tv

Image (9) (no background)

snaptv uses silk to give a voice to bite-sized lessons made for how india learns, quickly, on mobile, in simple hindi and easy english.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

jee concepts are difficult enough. monk learning uses silk to turn dense explanations into clear, natural voice for aspirants preparing every day.

monk learning

frequently asked questions

frequently asked questions

frequently asked questions

text to speech converts written text into spoken audio. a tts system reads the words, decides how the sentence should sound, and returns audio as a file or stream.
an ai voice generator creates synthetic speech from text. in silk, you can describe the speaker, accent, pace, tone, and role instead of only choosing a fixed preset voice.
yes. the silk playground lets you test tts in the browser before you build with the api. you can check the voice style, pacing, and delivery with your own text.
yes, long-form audio needs steady pacing and consistent delivery. silk is built for passages such as lessons, narration, product walkthroughs, and read-aloud experiences.
rumik converts text to speech through the silk playground and the silk api. the playground is better for quick audio tests, while the api is better when generated speech needs to run inside a product or workflow.
yes. silk works from python, javascript, or any backend that can make an http request. that makes it easy to add generated speech to web apps, agents, internal tools, and media pipelines.
yes. use streaming text to speech when generated replies need to be heard during a live conversation. silk mulberry starts speaking in 162 ms on the published latency benchmark.
silk returns wav by default and can also output mp3, opus, pcm, mulaw, or alaw through the audio format setting.
silk mulberry starts speaking in 162 ms on the published benchmark. that matters for calls, agents, and interactive products where the first reply is important.
use silk mulberry for fast speech, streaming, and voice descriptions. use silk muga when the script needs tone tags, inline events, or more expressive delivery.
yes. rumik-oss 1 is an open-source speech model for research and non-commercial use. it is separate from the hosted silk api models used for production traffic.

start building with rumik's tts api

start building with rumik's tts api