Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is not how thinking works. Claude uses tens of thousands of thinking tokens to get to 100 words. Kimi is no different.


Additionally, GP can use the word ‘concise’ (or similar) in their prompt if they want more concise output from a model.


That doesn’t mean “use fewer thinking tokens”. It might use more as it mulls over how to make its response concise.


effort level is a separate parameter, and is probably what they want




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: