Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Is there info on whether the safeguards that seem to be popping up / changing over time are at the behest of the developers, or is the software changing its response based on usage? Anthropomorphising ChatGPT, is it learning what morals are, or is it being constrained on its output? If it's the latter, I wonder how long until we see results from ChatGPT that are inherently supposed to be rendered because it's avoiding hard coded bad behavior. For example, perhaps it returns a racist response by incorrectly interpreting guidance that would prevent it being racist.

More succintly, these examples all seem to make ChatGPT ignore or get around its guardrails. I wonder if there are prompts that weaponize the guard rails.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: