If I were OpenAI, I’d do it so that people will have to find increasingly creative exploits, which we can then also patch (and keep patched for future models).
Long term they’re really worried about AI alignment and are probably using this to understand how AI can be “tricked” into doing things it shouldn’t.
Long term they’re really worried about AI alignment and are probably using this to understand how AI can be “tricked” into doing things it shouldn’t.