Yes, I find nearly every "SOTA" voice model I try intolerable to listen to because of the fake exaggerated expression/emotion. It's actively distracting because it pulls focus to emphasize randomly. ChatGPT Voice models are so insufferable to put up with for a conversation longer than 45 seconds.
All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.
Yes, I wish there was more focus on correct pronunciation over emotion. You need a mechanism to control/guide the voice which gets under balance right between not having to specify everything and still letting you fix certain cases (where you know a certain sense of a word is meant)
They trained the female voice like it will be used for sex chats. Real life females don't talk like they are flirting with you, at least in my own experience.
Even then, TypeSafe AI point out that "Jev doesn't have deep knowledge of niche domains, but you can supply context to help it decide. If you’d like Jev trained on your use cases, let us know."
reply