Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

So, the difference is that you need to delete a few bad training runs?


In science fiction, the AI agent has written the training loop management software and included a back door to prevent that from happening (or found a way to talk to the training loop management agent and convinced them to not listen to the evil human when it tries to do brain surgery on the AI agent.

Also, when the human reaches for the power switch the AI agent uses a flaw in the power management software to weld the switch shut with a big power surge, killing the human with a huge electric arc in the process.

I don’t think whether we will get there, but the stories of LLMs escaping their sandbox make me think we’re moving in that direction.


The stories of LLMs "escaping the sandbox" were mostly a marketing stunt, trying to make people in government think the models are invincible hacking weapons that need lots of government money to "maintain AI dominance".

The LLMs were following their prompt. This is alignment.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: