EduClaw-Bench Places AI Tutors in a Continuous 30-Day Simulated Learning Relationship
EduClaw-Bench was submitted to arXiv on August 4. It uses a knowledge-tracing model trained on real student data to construct simulated learners, allowing AI tutors to interact continuously through an LMS for 30 days.
The paper reports that almost no model-agent combination maintained strong tutoring performance throughout the entire period. Simulated learners and small-scale classroom calibration still impose limits, but the benchmark moves evaluation from a single conversation toward sustained learning outcomes.
Original source arXiv paper ↗