StudentBench: AI and human tutoring yield equivalent GRE learning gains
Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner, Jonas Mueller
arXiv:2609.28470v1Today’s pick is StudentBench, a paper about a question that matters to both education and AI: can large language models actually teach people as well as humans do? The authors built a large public benchmark from more than 175,000 student-AI messages and ran controlled studies with over 2,000 participants on GRE-style quantitative and verbal questions. They compared AI tutoring, expert human tutoring, and no tutoring, then measured how much each approach improved learning. The surprising result is that the best AI tutors produced learning gains statistically equivalent to expert human tutors, and in several domains they even did better on average. The paper also breaks tutoring into pieces like lesson planning, practice generation, conversational style, cost, and engagement, showing where AI is already strong and where humans still have an edge. Why it matters is simple: this moves the conversation from benchmark scores to real educational outcomes, and suggests AI tutoring could scale high-quality help at dramatically lower cost.
Also spotted that day
Previous daily papers