Research

AI-generated text

Jennifer Neville on What AI Failures Reveal and How Evaluation Shapes Progress

Jennifer Neville, partner research manager at Microsoft Research and Purdue professor, discusses how careful evaluation uncovers surprising failures in current AI systems and guides improvements.

Jennifer Neville on What AI Failures Reveal and How Evaluation Shapes Progress

Jennifer Neville is a partner research manager at Microsoft Research and the Samuel D. Conte Chair Professor of Computer Science and Statistics at Purdue University. Her work examines how machine learning and AI behave in interactive domains and on structured data, with special attention to how training data and evaluation shape system behavior. Neville has authored more than 130 papers with over 10,000 citations and received honors including a National Science Foundation CAREER Award, inclusion on IEEE’s “10 to Watch” in AI list, and best-paper awards from the International Conference on Data Mining and the International Conference on Learning Representations.

Research focus: evaluation as a gateway to improvement

Neville leads the AI Interaction and Learning team, which aims to push AI behavior boundaries in realistic work settings. The group argues that common ML benchmarks—often single-turn and simplified—miss critical aspects of how users actually work. Their approach starts with designing practical evaluations that reflect multiturn behavior, collaborative settings, and long-horizon tasks; identifying where current models fail in those evaluations then guides algorithmic and modeling improvements.

Surprising failures in multi-turn and long-horizon workflows

The team has published studies showing that models optimized for single-turn benchmarks can degrade significantly when tasks are specified across multiple turns. In the 2025 publication “LLMs Get Lost In Multi-Turn Conversation,” they simulated users revealing task details incrementally and found notable performance drops compared with single-turn performance. That result resonated with real users who reported similar experiences.

In related work on agentic systems and document-centric, long-horizon tasks, the researchers observed that repeated edits and chained operations allow errors to accumulate, sometimes causing subtle loss of semantic content or further confusion in subsequent steps. A 2026 paper titled “LLMs Corrupt Your Documents When You Delegate” describes these failure modes, and a May 2026 blog post offers further notes on delegation and long-horizon reliability.

Using product data responsibly to inform research

Neville emphasizes the value of analyzing large-scale product signals to learn how users succeed or fail in the wild, but clarifies that the team does this in a privacy-preserving, restricted manner. They extract high-level, anonymized patterns of failures and successes rather than directly inspecting individual user interactions. Those aggregated insights motivate experiments and algorithmic directions without exposing private content.

Neville also notes that richer user feedback (beyond binary like/dislike) — for example, short textual descriptions of what went wrong — can be particularly useful to improve models downstream.

Practical guidance for users

Neville’s practical recommendations for working with current AI systems include:

  • Verify outputs; do not assume models are correct across the board.
  • If a multiturn interaction produces confused behavior, consider restarting the chat and providing a fully specified single-turn instruction — models often perform better with a clear single turn.
  • When something goes wrong, give specific feedback describing the error and what you expected, because that detailed signal can be used to improve systems.

She stresses that AI tools are useful today but are best employed in workflows with human oversight and verification rather than full delegation.

Where the science might head next

Neville discusses two complementary directions: improving the surrounding infrastructure (agentic wrappers, tool use, retrieval, monitoring) to mitigate transformer limitations, and exploring alternative model architectures that better support long-term state and complex computation. She expects rapid progress given the scale of current research activity and computing resources, and foresees work focusing on longer-term interactions and models that have a more tangible understanding of the environments in which they operate.

A measure of success

For Neville, success would look like workplaces where humans and AI systems collaborate effectively as teams to accomplish tasks and innovate. Her long-term research thread concerns getting AI to work reliably within complex systems and multi-human, multi-AI interactions.

Personal takeaways

In the interview she shared two practical lessons: a formative piece of advice to be willing to ask questions (others often appreciate the courage to ask), and the repeated lesson to "look at the data" — many unexpected failures are revealed by inspecting the underlying data distribution or dataset errors.

Related publications mentioned

  • LLMs Get Lost In Multi-Turn Conversation — Publication | May 2025
  • LLMs Corrupt Your Documents When You Delegate — Publication | April 2026
  • Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability — Microsoft Research blog | May 2026