Probably, they're still using relatively simple queries. Maybe the memory is larger, but with 4096 tokens, GPT-3 is no slouch either. RLHF has tuned it so the conversational output of GPT is more inline with human expectation, it's also less likely to make glaring errors. The cost of their approach is it sometimes get stuck rephrasing things and wraps everything in boilerplate that can derail subsequent generation.