Once we find out how to crack SV2, I'll port them too. But for now, these are what I've made! The UTAU ones aren't the greatest, but they do work, and the SynthV ones are heavenly if I say so myself!
I put it through chatgpt Yep — I can see the problem from the actual voicebank you uploaded. The good news is the DiffSinger model itself looks internally consistent.
The important part of the error is:
Phoneme "-u" isn't supported by acoustic model
Your dsmain/phonemes.json contains the supported English ARPAbet phonemes (aa, ae, ah, ao, aw, etc.), but there is no -u phoneme.
So OpenUtau Mobile is feeding the model -u, while this Teto model only accepts its listed phonemes.
Why this is happening
Your bank has:
dsmain/phonemes.json → ARPAbet-style English phonemes dsdur/phonemes.json → same set default_phonemizer: OpenUtau.Core.DiffSinger.DiffSingerPhonemizer languages.json → English (en) envccv.yaml → an additional CV/CCV phoneme configuration
The model expects things like:
a → ah u → uw oo → uw
depending on the phonemizer.
But OpenUtau Mobile is apparently producing something like:
-u
The - is commonly used in UTAU-style phoneme notation to indicate a note/phoneme transition or special alias, but this particular DiffSinger acoustic model doesn't have such a symbol in its vocabulary.
Don't edit phonemes.json to add -u
That would not actually teach the ONNX acoustic model what -u means. The numerical IDs in phonemes.json correspond to the inputs the model was trained on. Adding an arbitrary entry could just cause a different failure or, worse, incorrect rendering.
Most likely fix
This looks like a phonemizer compatibility issue between this particular Teto port and the OpenUtau Mobile version you're using, rather than a broken ONNX model.
Try this first:
Make a new empty track. Select Teto SV1 EN. Make one very simple note, e.g. ah. Render it. If that works, enter an English word such as you. Check the generated phonemes before rendering.
If the phoneme display contains -u, that's the smoking gun.
If you can send me a screenshot of the notes + the phonemes OpenUtau generated immediately before rendering, I can tell you exactly what needs changing.
And because you uploaded the actual bank, I can also inspect its configuration and work out whether this is fixable by changing the phonemizer/config, rather than modifying the model itself.
For some reason when I downloaded the voicebank first it corrupted (prob cause of my file manager) so I got zarchiver to be my default file manager and it installed correctly
Maybe its on en vccv by accident? I use a mobile port of openutau for Android and it automatically has the en vccv phonemizer but I changed it to diffs en
This comment has been removed by the author.
ReplyDeleteWhen I use any diffsinger phonemizer a error pops up
ReplyDeletewhat error? if anything you could copy paste it into chatgpt and say "what happen :(" bc that's lowkey what i do lmao
DeleteI put it through chatgpt
DeleteYep — I can see the problem from the actual voicebank you uploaded. The good news is the DiffSinger model itself looks internally consistent.
The important part of the error is:
Phoneme "-u" isn't supported by acoustic model
Your dsmain/phonemes.json contains the supported English ARPAbet phonemes (aa, ae, ah, ao, aw, etc.), but there is no -u phoneme.
So OpenUtau Mobile is feeding the model -u, while this Teto model only accepts its listed phonemes.
Why this is happening
Your bank has:
dsmain/phonemes.json → ARPAbet-style English phonemes
dsdur/phonemes.json → same set
default_phonemizer: OpenUtau.Core.DiffSinger.DiffSingerPhonemizer
languages.json → English (en)
envccv.yaml → an additional CV/CCV phoneme configuration
The model expects things like:
a → ah
u → uw
oo → uw
depending on the phonemizer.
But OpenUtau Mobile is apparently producing something like:
-u
The - is commonly used in UTAU-style phoneme notation to indicate a note/phoneme transition or special alias, but this particular DiffSinger acoustic model doesn't have such a symbol in its vocabulary.
Don't edit phonemes.json to add -u
That would not actually teach the ONNX acoustic model what -u means. The numerical IDs in phonemes.json correspond to the inputs the model was trained on. Adding an arbitrary entry could just cause a different failure or, worse, incorrect rendering.
Most likely fix
This looks like a phonemizer compatibility issue between this particular Teto port and the OpenUtau Mobile version you're using, rather than a broken ONNX model.
Try this first:
Make a new empty track.
Select Teto SV1 EN.
Make one very simple note, e.g. ah.
Render it.
If that works, enter an English word such as you.
Check the generated phonemes before rendering.
If the phoneme display contains -u, that's the smoking gun.
If you can send me a screenshot of the notes + the phonemes OpenUtau generated immediately before rendering, I can tell you exactly what needs changing.
And because you uploaded the actual bank, I can also inspect its configuration and work out whether this is fixable by changing the phonemizer/config, rather than modifying the model itself.
Hm
DeleteHmmmm.
ReplyDelete(I am using a port of openutau for Android with utau vogen neutrino and diffsinger support
ReplyDeleteI tried other phonemes it still fails to render
DeleteI fucking realised that it didn't download correctly now it works really well tbh
ReplyDeletehow did you fix the download I'm curious
DeleteFor some reason when I downloaded the voicebank first it corrupted (prob cause of my file manager) so I got zarchiver to be my default file manager and it installed correctly
DeleteIt doesn't work. I'm using the "DIFFS EN" phonemizer
ReplyDeleteOn mobile or computer?
DeleteAlso has you installed nsf hifigan
DeleteThis comment has been removed by the author.
DeleteI'm on PC, it doesn't make an error like usual and...how do I explain it,..in the part where you see the phonemes it says error
DeleteI know that does happen sometimes but idk how to fix it, but what I do know is that the USTx must be in plain English, not VCCV or CVVC written
DeleteMaybe its on en vccv by accident? I use a mobile port of openutau for Android and it automatically has the en vccv phonemizer but I changed it to diffs en
ReplyDeleteWhen you put vccv en it will put a error saying there is no e.g -u phoneme
Delete