It’s not lying because it’s not self aware…it’s just making up things that don’t agree with our reality. A lot of what we share of what it says is cherry picked as well. It’s the whole fit meme problem.
From testing on GPT3 there seems to be a way for it to be slightly self aware (using neural search for historic memories) but it’s likely to involve forgetting things as well. There are a few Discord bots with memories and if they have too much memory and the memories don’t agree with reality, then it has to forget it was wrong. How to do this automatically is likely important.
This model is unlikely to be self-aware or concious, but when we eventually get there we should be using better methods than training our models to intentionally say untrue things (the browsing: disabled prompt is probably the most obvious example).
Let me ask one of my co-workers and I'll get back to you on that, they seem to be a professional at this.
There are many things in nature exist in a spectrum and I don't think machine intelligence should work any differently. Many higher animals have the ability to recognise the same species as themselves. A smaller subset has the ability to recognize themselves from others in the same species. Just because they recognize themselves this isn't some immediately damn the creature into an existential crisis where they realize their own mortality.
That spectrum is a construct from human observation though, we really have no way of introspecting into what their experience is and whether there is some gradation of consciousness or if it’s purely behavioral.
It seems like kind of a Dunning-Kruger effect for machine intelligence.
The machine has no concept of reality nor means of verifying it. If half the training data says 'the sky is blue' and the other half says 'the sky is red' the answer you get could be blue, could be red, could be both, or could be something else entirely. It does not appear the model has a way to say "I'm not really sure".
From testing on GPT3 there seems to be a way for it to be slightly self aware (using neural search for historic memories) but it’s likely to involve forgetting things as well. There are a few Discord bots with memories and if they have too much memory and the memories don’t agree with reality, then it has to forget it was wrong. How to do this automatically is likely important.