introducing rumik oss 1

introducing rumik oss 1

published

published

sept 8, 2026

sept 8, 2026

by

by

rumik research team

rumik research team

we are releasing our first open source model of the silk family. to our knowledge this is the first open source indic text to speech model with expressive emotions.

many researchers assume that if they use muP (a method by Yang et al, 2022, for transferring hyperparameters across model sizes), they are safe from LR-tuning headaches.

you get the option to control descriptions for the voice parameterized by pace, accent and tone.

you get the option to control descriptions for the voice parameterized by pace, accent and tone.

rumik-oss supports 22 languages (native transcriptions), inline emotion tags with a global description as well.

listen to some examples below or if you want to test/hack it out yourself :

  • space

  • weights

rumik-oss supports 22 languages (native transcriptions), inline emotion tags with a global description as well.

listen to some examples below or if you want to test/hack it out yourself :

  • space

  • weights

sad
ira
hindi
hindi accent
slow
00:00
माँ,आजफिरतुम्हारेलिएचायबनादी।<sigh>कपठंडाहोगया,परतुम्हाराइंतज़ारनहीं।
sad
aisha
english
indian english accent
slow
00:00
Istillplayyourlastmessage,Dad.Youonlysaid,callmewhenyougethome.Igothome.Ijustnevergottotellyou.
excited
siya
telegu
telegu accent
fast
00:00
అమ్మా,నాకుఉద్యోగంవచ్చింది!నిజంగావచ్చింది!నువ్వునాకోసంచేసినప్రతిత్యాగంనాకుగుర్తుంది.<laugh>రోజునీకునచ్చినస్వీట్లునేనేకొనిస్తాను!
happy
aisha
bengali
bengali accent
steady
00:00
এতদিনপরেতোমাকেদেখেকীযেভালোলাগছে!তোমারপছন্দেরসবরান্নাকরেছি।আজআরকোথাওযেতেদেবনা,সবাইমিলেঅনেকগল্পকরব।
angry
zoya
tamil
tamil accent
fast
00:00
நீவருவேன்னுசொன்னதால்தான்இவ்வளவுநேரம்காத்திருந்தேன்!ஒருபோன்கூடபண்ணமுடியலையா?ஒவ்வொருமுறையும்மன்னிப்புகேட்டாமட்டும்எல்லாம்சரியாகிடாது!
professional
aisha
telegu + english
telegu accent
steady
00:00
మనకొత్తdemoసిద్ధంగాఉంది.Youcanchooseavoice,changethepace,andtrydifferentlanguages.ముందుగాతెలుగులోఒకవాక్యంవిందాం,thenwecanswitchtoEnglish.
happy
ira
english + hindi
hindi accent
steady
00:00
आजकीmeetingमेंसबकोहमाराideaपसंदआया।Iwassonervous,लेकिनतुमनेकहाथाना,बसदिलसेबोलो।<chuckle>अबcoffeeमेरीतरफ़से,औरcakeतुम्हारीतरफ़से!
professional
siya
kannada
kannada accent
00:00
ನಮ್ಮಹೊಸಗ್ರಂಥಾಲಯಕ್ಕೆಸ್ವಾಗತ.ಇಲ್ಲಿಕನ್ನಡಮತ್ತುಇಂಗ್ಲಿಷ್ಪುಸ್ತಕಗಳುಲಭ್ಯವಿವೆ.ಸದಸ್ಯರಾಗಲುನಿಮ್ಮಹೆಸರುಮತ್ತುವಿಳಾಸನೀಡಿ.ಓದಲುಶಾಂತವಾದಸ್ಥಳವೂಇದೆ.
excites
zoya
punjabi
punjabi accent
fast
00:00
ਮਾਂ,ਮੇਰਾਦਾਖ਼ਲਾਹੋਗਿਆ!ਸੱਚੀਂ,ਚਿੱਠੀਗਈਹੈ!ਜਿਹੜਾਸੁਪਨਾਅਸੀਂਇਕੱਠੇਵੇਖਿਆਸੀ,ਉਹਅੱਜਪੂਰਾਹੋਗਿਆ।ਹੁਣਸਭਨੂੰਫ਼ੋਨਕਰ!

benchmarks

we evaluate expressive delivery, vocalization adherence, and transcription accuracy separately. indicEmo measures how well a requested delivery is sustained through code-switched speech; NoVA tests whether an inline vocalization occurs at the requested position. WER and CER assess the correspondence between the synthesized speech and its input text.

indicEmo

we introduce indicEmo to evaluate expressive delivery in code-switched speech. its 100 prompts combine two to five languages drawn from english, hindi, telugu, tamil, kannada, bengali, and punjabi, with equal coverage of happy, sad, angry, excited, and professional delivery.

three automated judges rate anonymized recordings on a five-point rubric covering tone, dynamics, phrasing, and sustained expression. results use the 98-prompt intersection with valid ratings from all judges, aggregating the median rating per recording into category means. rumik-oss-1 scores 2.92/5 overall and 3.03/5 on the four emotion categories.

NoVA

NoVA evaluates adherence to inline laughter, chuckle, and sigh requests in english speech. each of its 149 prompts places a single vocalization at a specified word position. seven systems are evaluated over three synthesis runs using provider-specific prompting.

silk-asr, scribe v2, and a gemini verifier assess position-correct rendering. laugh and chuckle are pooled into one category; the final score averages laughter and sigh rendering rates across detectors, with unsupported categories contributing zero. rumik-oss-1 achieves a 0.884 rendering score.

(we also show WER and CER benchmarks in the post-training section)

pre-training

pre-training

we started with a very simple question, do you really need a million hours of audio data to pre-train a text to speech model from scratch ? the open labs do seem to say so.

we started with a very simple question, do you really need a million hours of audio data to pre-train a text to speech model from scratch ? the open labs do seem to say so.

welp, there was only one way to find out!

welp, there was only one way to find out!

text base choice. for saving us the hassle of multilingual text pre-training, we wanted to build upon a prior which has a good multilingual understanding such that our entire focus shifts to audio pre-training and post-training. we also wanted a simple design that :

  • is scalable and easy to serve

  • has a robust multilingual text distribution

  • follows the standard next token prediction objective

hence we start with tiny aya fire (becuase of its rich text multilingual prior for indic use cases) as the text base; and use mimi codec for discretizing audio (audio waveforms to speech tokens).

hence we start with tiny aya fire (becuase of its rich text multilingual prior for indic use cases) as the text base; and use mimi codec for discretizing audio (audio waveforms to speech tokens).

we keep mimi frozen and use 24khz audio, 12.5 frames per second and 8 codebooks. all 8 codes of one frame come before the next frame: 100 audio tokens per second. the model predicts these from text and conditions; mimi decodes them.

we keep mimi frozen and use 24khz audio, 12.5 frames per second and 8 codebooks. all 8 codes of one frame come before the next frame: 100 audio tokens per second. the model predicts these from text and conditions; mimi decodes them.

with a limited audio and compute budget, the first question was which public data would give us the best speech foundation.

speech pre-training. we first train the model for text-to-speech over 75,000 updates, using only emilia’s english subset, which contains roughly 40,000 hours of speech. we use a warmup–stable–decay learning rate schedule with a peak of 1e-4 and a global batch size of 96.

multilingual and voice adaptation. for the next 30,000 updates, we introduce indian-language speech from indicvoices, indicvoices-r, and rasa. indicvoices has 7,348 published hours. indicvoices-r provides 1,704 hours drawn from indicvoices and prepared specifically for tts, while rasa contributes another roughly 1,145 hours.

speaker adaptation and description following. we use the remaining quality proprietary audio to introduce natural-language control over style, accent, and pace; alternative categorical-affect control; and temporally localized non-verbal events. the supervision is applied through supervised fine-tuning and is interleaved with multilingual, named-voice, and ordinary text-conditioned speech replay to limit interference with pronunciation and speaker identity.

so the gist is that we first train the text base to predict english audio tokens, and progressively shift to multilingual data followed by clean and rich speaker's data, and finally add description following.

the training objective is standard next token prediction for speech tokens using cross entropy loss, just like it is the case for a decoder only transformer.

given textual context x and flattened acoustic sequence y, the model factorizes the conditional likelihood as :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

let m_i ∈ {0, 1} select valid acoustic targets after prefix and padding masks are applied. the autoregressive objective is :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

only acoustic positions contribute to the loss; textual and conditioning tokens are observed context rather than reconstruction targets. for binary frame-boundary target s_t and predicted termination probability ŝ_t, the complete supervised objective is :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

the transformer and termination head are optimized jointly, while the mimi encoder and decoder remain fixed throughout training.

post training

post training

we're gonna do RL post training yayyyy!

we're gonna do RL post training yayyyy!

hmm reasoning about the data or prompt curation; what can the number look like if we calculate diversity or at least the diverse combinations that we would need to address?

hmm reasoning about the data or prompt curation; what can the number look like if we calculate diversity or at least the diverse combinations that we would need to address?

okay so you have n languages, but then, you have m speakers, wait you also have description parameters : a, b, c; the search space is basically: m * n * a * b * c;

okay so you have n languages, but then, you have m speakers, wait you also have description parameters : a, b, c; the search space is basically: m * n * a * b * c;

hmm; not bad (manageable; maybe?)

but no the story doesn't end here; what if i wanna do mid sentence switching ? can we add different ASR judges per language ? can we add inline emotion tags too please ?

but but but how do you select the domains for the prompts ? (tongue twisters ? financial data ? telephony ? casual convos ?)

what about infra? handling training-inference mismatch? hyperprams? LR tuning? LoRA or full weight updates? optimizer sync to the rollout engine? phased GRPO with SFT update or only GRPO? and probably infinite many other things.

our reaction :

jokes aside; for post training we use GRPO to optimize for different rewards; this naturally fits the transformer regime for enhancing or redistributing probability masses in the distribution of decoder only models for our specific targets; the base being a transformer predicting discretized audio tokens (mimi) fits this naturally.

on GRPO or rather what it becomes based on our algorithmic choices:

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

we do not use any value network; the base generates the groups; so the simple pipeline is basically take a prompt, roll out G completions, and judge each against its siblings as per the rewards.

for a simple starting recipe, we use word error rate (wer), character error rate (cer) and a length penalty term.

word error rate. according to wikipedia, we use the standard wer definition.

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

length penalty. CER and WER are computed on a transcript, and a transcript has no duration. we define penalty as follows :

t_gen is how long the generated audio actually is, and t_exp is how long the text should take to say, which for the hand built sets is the character count divided by a per language characters per second rate:

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

so the penalty terms comes out to be :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

calculating advantages (A), mean and variance follow the standard recipe :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

the full objective with the PPO ratio and KL anchor is :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

we take exactly one gradient step per batch of rollouts; and with beta set to 0 the entire objective reduces to a plain policy gradient :

we take exactly one gradient step per batch of rollouts; and with beta set to 0 the entire objective reduces to a plain policy gradient :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

we normalise by the total unmaksed token count across the batch, rather than per sequence, the idea being that we want to stop long rollouts from domintaing the gradient by simply being longer in generation.

we normalise by the total unmaksed token count across the batch, rather than per sequence, the idea being that we want to stop long rollouts from domintaing the gradient by simply being longer in generation.

on brodcasting the rewards to the codec tokens

on brodcasting the rewards to the codec tokens

since mimi codes follow an RVQ structure, where codebook 0 follows captures most of the semantic content and the later codebooks carry the acoustic structure of the audio, broadcasting the right reward to the right codebook becomes very important.

since mimi codes follow an RVQ structure, where codebook 0 follows captures most of the semantic content and the later codebooks carry the acoustic structure of the audio, broadcasting the right reward to the right codebook becomes very important.

you would not want your major acoustic signal to go to codebook 0 entirely (which is responsible for major semantic structure), or you also would not want the semantic signal to be broadcasted to only the acoustic tokens.

you would not want your major acoustic signal to go to codebook 0 entirely (which is responsible for major semantic structure), or you also would not want the semantic signal to be broadcasted to only the acoustic tokens.

token 1 and token n of the same rollout receive an identical coefficient. a wrong code, can put the entire residual chain in the wrong place, and a wrong code, which is cosmetic, gets exactly the same learning signal.

the eight token frame structure (of the audio) becomes invisible to our post-training objective.

token 1 and token n of the same rollout receive an identical coefficient. a wrong code, can put the entire residual chain in the wrong place, and a wrong code, which is cosmetic, gets exactly the same learning signal.

the eight token frame structure (of the audio) becomes invisible to our post-training objective.

we later added KL monitoring and used the schulman k3 estimator, which is unbiased and non-negative, rather than using the raw log-ratio difference:

we later added KL monitoring and used the schulman k3 estimator, which is unbiased and non-negative, rather than using the raw log-ratio difference:

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

since q0 (code 0) is semantic and q1 ... q7 are acoustic residuals, a CER-derived reward pushing equally on all eight is spending major gradient signal (roughly seven eighths of the content gradient) on tokens that cannot carry content.

so instead we do reward routing as:

since q0 (code 0) is semantic and q1 ... q7 are acoustic residuals, a CER-derived reward pushing equally on all eight is spending major gradient signal (roughly seven eighths of the content gradient) on tokens that cannot carry content.

so instead we do reward routing as:

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

the stop token goes with the duration stream because it is the stop decision for the model.

the reward signal so far looks like :

the reward signal so far looks like :

w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t
w_{final} = w_{init} + \sum_{t=0}^{T} \Delta w_t

for the equation above things like timbre, prosody, speaker identity, emotions are free to move, and that means the target weights that encode these are free to move as well. this is an honest limitation and we plan to address and this in the upcoming tech report.

for the equation above things like timbre, prosody, speaker identity, emotions are free to move, and that means the target weights that encode these are free to move as well. this is an honest limitation and we plan to address and this in the upcoming tech report.

on prompt curation for GRPO

on prompt curation for GRPO

we first evaluate the base model, on our curated set weighted equally by language distribution. we use indic conformer as the ASR to judge our model against propreitary models for standard word error rate and character error rate.

we first evaluate the base model, on our curated set weighted equally by language distribution. we use indic conformer as the ASR to judge our model against propreitary models for standard word error rate and character error rate.

we find sets (by domains) where the base underperforms and synthetically generate more targeted sets for RL.

note that the our focus has been on improving consistency and speaker preservation for the released models, hence the RL prompt sets are targeted by domains (hard prompts, digits, financial data, conversational data etc.) and language distribution. we'll soon be releasing upgraded checkpoints where we've made emotion descriptions consistent across the model (to let you in on a secret, phased grpo rollouts with dpo updates work!).

what do the final set of rewards look like ?

what do the final set of rewards look like ?

we wanted (at least till now) to keep things simple, instead of making the axes for reward coverage overly complex to cover all the areas where we wanted to improve.

we wanted (at least till now) to keep things simple, instead of making the axes for reward coverage overly complex to cover all the areas where we wanted to improve.

so for indic scripts we later added GCER, which is the same quantity computed over grapheme clusters rather than unicode code points.

so for indic scripts we later added GCER, which is the same quantity computed over grapheme clusters rather than unicode code points.

GCER (like WER and CER) is also a levenshtein distance between a reference sequence and a hypothesis sequence, divided by the length of the reference.

GCER (like WER and CER) is also a levenshtein distance between a reference sequence and a hypothesis sequence, divided by the length of the reference.

levenshtein distance is the minimum number of single element insertions, deletions and substitutions that turn one sequence into the other. written as a recurrence over prefixes, with a_1..i and b_1..j :

levenshtein distance is the minimum number of single element insertions, deletions and substitutions that turn one sequence into the other. written as a recurrence over prefixes, with a_1..i and b_1..j :

LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}
LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}

the three rates differ only in what a "unit" is:

the three rates differ only in what a "unit" is:

LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}
LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}

putting every piece together, this is the whole function incorporating the final rewards :

putting every piece together, this is the whole function incorporating the final rewards :

LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}
LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}

where

where

LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}
LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}

with the individual scores collected from the sections above:

LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}
LR^*(D_2) \approx LR^*(D_1) \left( \frac{D_2}{D_1} \right)^{-0.32}

note that for building the softmax distribution over the rollouts we only build it over the audio tokens and drop the text tokens, although the model was pre-trained with an up projection over the entire vocabulary.

WER and CER benchmarks

WER and CER benchmarks

CER (character error rate) is the fraction of characters the model got wrong. WER (word error rate) is the same idea at the word level. both are computed by transcribing the generated audio with a speech recognition model (speech intelligibility measured with indicconformer asr.) and comparing that transcript to the text we asked it to say. Lower is better, and 0.0 means a perfect match.

CER (character error rate) is the fraction of characters the model got wrong. WER (word error rate) is the same idea at the word level. both are computed by transcribing the generated audio with a speech recognition model (speech intelligibility measured with indicconformer asr.) and comparing that transcript to the text we asked it to say. Lower is better, and 0.0 means a perfect match.

word error rate measured on selected languages against proprietary models.

word error rate measured on selected languages against proprietary models.

character error rate measured on selected languages against proprietary models.

go and play with the model weights on huggingface. and let us know what you think. we will very soon be following up with a detailed technical report and our training code (yes you heard that right!) as we want this to be truly open source. also stay tuned for updates on more post training for decoder only transformer models. happy hacking!


reach out to us at research@rumik.ai for more consulting about data, training recipes, or if you're building/experimenting on top of rumik-oss, we'd love to chat!