Why not just use ChatGPT roleplay?
ChatGPT produces conversations. vLeader produces consequences — governed by fixed rules, traceable to the choice that caused them, and usable in Assurance-of-Learning reporting. A language model can give different, sometimes fabricated feedback to every student, so the scores aren't comparable across a cohort. Hand-coded rules are the same rules for everyone. Exact dialogue lines vary between runs; the causal logic does not.