Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
bpodgursky
17 days ago
|
parent
|
context
|
favorite
| on:
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
This is not how thinking works. Claude uses tens of thousands of thinking tokens to get to 100 words. Kimi is no different.
no-name-here
17 days ago
|
next
[–]
Additionally, GP can use the word ‘concise’ (or similar) in their prompt if they want more concise output from a model.
dymk
17 days ago
|
parent
|
next
[–]
That doesn’t mean “use fewer thinking tokens”. It might use more as it mulls over how to make its response concise.
tibbar
17 days ago
|
root
|
parent
|
next
[–]
effort level is a separate parameter, and is probably what they want
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: