Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations

ArXi:2605.02624v1 Announce Type: new There is growing interest in exploring user simulation as an alternative to gathering and scoring real user-chatbot interactions for AI chatbot evaluation. For this purpose, it is important to ensure the realism of the simulation, i.e., the extent to which simulated dialogues reflect real dialogues users have with chatbots. Most existing methods evaluating simulation realism produce coarse quality signal and remain solely at the level of individual dialogues.