I've used it to build and/or run various machine learning models for text generation, speech recognition, image generation, depth estimation, etc. in the browser, in support of an agentic system I've been building out.
Lots of future possibilities as well once support is more ubiquitous!
I'll try to find time to write about it, but in the meantime if you just want to try something that works, Xenova published some tools and examples about two months ago which I'm sure will give you a good start.
Speaker diarization is quite difficult as you know, especially in loud or crowded environments, and the model is only part of the story. A lot of tooling needs to be built out for things like natural interruption, speaker memory, context-switching, etc. in order to create a believable experience.
I'm curious, what for?