Neat. I threw a couple simple audio clips at it and it was able to at least reco...

		vunderba 3 months ago \| parent \| context \| favorite \| on: Qwen3-Omni: Native Omni AI model for text, image a... Neat. I threw a couple simple audio clips at it and it was able to at least recognize the instrumentation (piano, drums, etc). I haven't seen a lot of multimodal LLM focus around recognizing audio outside of speech, so I'd love to see a deep dive of what the SOTA is.