Eye-tracking Study Reveals Where Human and AI Language Processing Converge—and Diverge
Anyone who has stumbled over a sentence and had to reread it has experienced one of the human brain’s most remarkable features— the ability to recognize that it misunderstood what it just read and reconstruct the sentence’s meaning.
Researchers at New York University and the University of Massachusetts Amherst have shown that this moment of rereading marks a significant difference between how people and today’s artificial intelligence or large language models (LLMs) process language.
Their findings, published in the Proceedings of the National Academy of Sciences reveal where humans and today’s LLMs process language similarly, and where they don’t. The work could help researchers better understand human language processing and guide future efforts to build AI systems that more closely resemble it.
Brian Dillon, UMass Amherst linguistics professor and director of the Computational Sentence Processing Lab in the Integrative Learning Center, co-authored the paper with six researchers, including NYU’s Tal Linzen, associate professor of linguistics and data science, and lead author William Timkey, a doctoral student in psycholinguistics.
Conducted from 2022 to 2025, the study began just as artificial intelligence and specifically chatbots, like ChatGPT, were emerging. LLMs excel at next-word prediction, which aligns with some theories of how the brain works and comprehends sentences.
“LLMs seem to capture some of the properties of language as we understand it—they can generate text fluently and they appear to react in a way that suggests they have some understanding of what’s going on,” Dillon explains.
Researchers used eye-tracking technology to analyze 368 adult readers, most of them NYU and UMass Amherst students, focusing on how long they spent reading and rereading each word in a carefully designed sentence. These included a diverse set of syntactically challenging sentences, including “garden path” sentences that initially lead readers toward one interpretation (or path) then suddenly force them to reconsider.
Dillon said two UMass Amherst researchers, Lyn Frazier and Keith Rayner, published a foundational research study in the 1980s that explained the “garden path effect,” the pause that arises when reading a garden path sentence. For example, consider the sentence: “Since Jay always jogs a mile seems like a short distance.” The beginning leads readers to think the sentence is about Jay jogging a mile, but then they get stuck at “seems,” and discover there is a second subject of the sentence; “a mile.”
Dillon adds, “A simple comma after ‘jogs,’ probably would have saved the reader time in processing that sentence.”
NYU and UMass Amherst researchers then compared participants’ eye movements with predictions generated by more than 400 AI language models. They found that humans and AI track language pace similarly during the earliest moments of reading, assumably by next-word prediction. However, as reading continues and becomes more complex, their processing differs from that of AI.
“Basically, the eyes can react in one of two ways when there’s difficulty. One is that the eyes might linger longer on a word. This would be very short; you wouldn’t notice it,” Dillon said. “And sometimes the eyes move backwards in a text; you go forward, you hit a word that you have difficulty with, and the eyes instinctively jump back. We call that a regression.”
Dillon explained that this regressive eye movement allowed them to examine the difficulty with unexpected words. “What we found was that the language models underpredicted the difficulty associated with them.”
Timkey said while their findings could help computational linguists build language models that better reflect human language processing, the study is equally valuable because it reveals something about “how” people understand language, therefore gaining a clearer picture of the cognitive processes of human communication.
“If we can develop models that explain how long we take to read, including these moments when we go back and reread, then that model would give us a very good characterization of what’s actually happening in the human mind as it constructs the meaning of a sentence,” Timkey said.
Beyond improving AI, Timkey said the findings, for example, could lead to more fine-grained methods for assessing literacy and help educators with tailoring instruction to a student’s specific language processing needs or challenges.
“I am excited to find ways of applying our findings to these important practical challenges,” Timkey said.
The study is part of a larger collaborative project, “Collaborative Research: Semantic Focusing: Controlling LM Interpretations for Human-Model Alignment” which has received approximately $1.2 million in combined total funding, including last year's award of $432,656 to UMass Amherst from the National Science Foundation.
More
Linguistics’ Brian Dillon Receives NSF Grant to Explore AI and Human Language Processing
Dillon, professor of linguistics in the College of Humanities and Fine Arts, has been awarded a four-year, $432,656 research grant from the National Science Foundation to investigate how artificial intelligence systems and humans differ in the way they process language.