Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right?
Ive also used a text and audio input in order to get a text description or classification out.
I cannot for the life of me find a solution for Audio + text -> Audio
My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?
I was working with very simple elements, but I was surprised by some of the outputs.
I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference.
KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:
https://www.youtube.com/watch?v=WAeHgE94rVo
Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
This generates audio embeddings - much like CLIP does for visual inputs.
I think Google had one called riffusion (the first version was designed for specs)
The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)
What's your exact use case?