Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
KVarN: Native vLLM backend for KV-cache quantization by Huawei (github.com/huawei-csl)
143 points by theanonymousone 53 days ago | past | 16 comments
Sinkhorn: Make LLMs even smaller through quantisation while maintaining accuracy (github.com/huawei-csl)
4 points by ilitirit 9 months ago | past | 1 comment

Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: