Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> and use WebGPU all the time

I'm curious, what for?



I've used it to build and/or run various machine learning models for text generation, speech recognition, image generation, depth estimation, etc. in the browser, in support of an agentic system I've been building out.

Lots of future possibilities as well once support is more ubiquitous!


Your ideas are intriguing to me and I wish to subscribe to your newsletter.


I appreciate that, anything in particular catch your interest?


I am most interested in speech recognition including diarization.


I'll try to find time to write about it, but in the meantime if you just want to try something that works, Xenova published some tools and examples about two months ago which I'm sure will give you a good start.

https://github.com/xenova/whisper-web/tree/experimental-webg...

https://huggingface.co/spaces/Xenova/whisper-speaker-diariza...

https://huggingface.co/onnx-community/pyannote-segmentation-...

Speaker diarization is quite difficult as you know, especially in loud or crowded environments, and the model is only part of the story. A lot of tooling needs to be built out for things like natural interruption, speaker memory, context-switching, etc. in order to create a believable experience.


Thanks for the links! Most of the fined tuned diarization models are locked up behind SaaS paywalls, not open weights usable directly.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: